Question answering processing method and system, device and storage medium
By introducing hierarchical disassembly of the first agent and the second agent in the question-and-answer system, the problem of insufficient multi-table joint query capabilities in the field of intelligent transportation is solved, and more efficient and accurate user problem responses are achieved, and the application scenario of LLM is expanded.
Patent Information
- Application Number
- PCT/IB2024/062702
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-01-03
- Filing Date
- 2024-12-16
- Publication Date
- 2025-07-10
AI Technical Summary
The existing question-and-answer system based on large language model (LLM) has weak multi-table joint query capabilities in the field of intelligent transportation, and due to the limitations of LLM's capabilities, it is difficult to effectively deal with complex user problems.
By introducing hierarchical disassembly of the first agent and the second agent, the first agent is responsible for planning the overall solution process of user problems, and the second agent is responsible for the task execution of specific functions, using multiple agents to work together to reduce the thinking complexity of LLM and expand the application scenario.
It improves the ability of the Q&A system to handle complex problems, reduces the probability of errors, expands the application boundaries of LLM, can respond to user questions more accurately, and covers a wider range of application scenarios.
Smart Images

Figure IB2024062702_10072025_PF_FP_ABST
Abstract
Description
[0001] Question and Answer Processing Method, Device, Storage Medium, and System This disclosure claims priority to Chinese Patent Application No. 202410016581.1, filed with the Patent Office of China on January 3, 2024, entitled "Question and Answer Processing Method, Device, Storage Medium, and System," the entire contents of which are incorporated herein by reference. Technical Field This disclosure relates to the field of artificial intelligence technology, and more particularly to a question and answer processing method, device, storage medium, and system. Background: With the continuous development of Large Language Model (LLM) technology and the significant improvement in various aspects of model performance, various LLM-based applications are receiving widespread attention. For example, in the field of intelligent transportation, users may want to understand traffic conditions such as daily congestion in their city. An existing approach to user interaction in the intelligent transportation field, based on LLM, utilizes LLM's code generation capabilities. After receiving a user's input, LLM automatically generates corresponding SQL (Structured Query Language) query instructions. By executing the SQL script, relevant data can be queried within the transportation service platform. This natural language-to-SQL query-to-SQL approach has limited application scenarios and is only suitable for SQL-based data query tasks. Furthermore, due to the current limitations of LLM, its multi-table join query capabilities are currently relatively weak. SUMMARY OF THE INVENTION The present disclosure provides a question-and-answer processing method, device, storage medium, and system for accurately and automatically responding to user questions through the collaboration of different intelligent agents.In a first aspect, an embodiment of the present disclosure provides a question-answering processing system, comprising: a first agent, a second agent, and an application server, wherein the first agent and the second agent share a same large language model; the first agent is configured to receive a user question, incorporate the user question into a first prompt word template corresponding to the first agent, obtain, based on first toolset information included in the first prompt word template and the large language model, multiple pieces of solution task information sequentially generated by the large language model and requiring execution by the second agent, as well as summary task information requiring execution by the first agent, and sequentially send the multiple pieces of solution task information to the second agent; the second agent is configured to incorporate received current solution task information into a second prompt word template corresponding to the second agent, obtain, based on second toolset information included in the second prompt word template and the large language model, task execution information corresponding to the current solution task information generated by the large language model, obtain, based on the task execution information, a task execution result for the current solution task information from the application server corresponding to the user question, and send the task execution result to the first agent. So that the first intelligent agent generates the next solution task information through the large language model after obtaining the task execution result of the current solution task information; the first intelligent agent is used to summarize the task execution results of the multiple solution task information based on the summarized task information to determine the reply information corresponding to the user question.In a second aspect, embodiments of the present disclosure provide a question-and-answer processing method, applied to a first agent, wherein the first agent interacts with a second agent to complete the question-and-answer processing method, and the first and second agents share a common large language model. The method includes: receiving a user question; splicing the user question into a first prompt word template corresponding to the first agent, and obtaining, based on first toolset information included in the first prompt word template and the large language model, multiple solution task information sequentially generated by the large language model and requiring execution by the second agent, as well as summary task information requiring execution by the first agent; sequentially sending the multiple solution task information to the second agent, so that the second agent obtains task execution information corresponding to the current solution task information, generated by the large language model, based on the second prompt word template corresponding to the second agent, the received current solution task information, and the large language model; obtaining, based on the task execution information, a task execution result for the current solution task information from an application server corresponding to the user question; and sending the task execution result to the first agent. So that after obtaining the task execution result of the current solution task information, the first intelligent agent generates the next solution task information through the large language model, and the second prompt word template includes second toolset information; obtains the task execution results of the multiple solution task information sent by the second intelligent agent; based on the summarized task information, summarizes the task execution results of the multiple solution task information to determine the reply information corresponding to the user question.In a third aspect, an embodiment of the present disclosure provides a question-and-answer processing apparatus, applied to a first agent, wherein the first agent interacts with a second agent to complete question-and-answer processing, and the first and second agents share the same large language model. The apparatus includes: a receiving module for receiving a user question; a generating module for splicing the user question into a first prompt word template corresponding to the first agent, and obtaining, based on first toolset information included in the first prompt word template and the large language model, multiple solution task information sequentially generated by the large language model and requiring execution by the second agent, as well as summary task information requiring execution by the first agent; a sending module for sequentially sending the multiple solution task information to the second agent, so that the second agent obtains task execution information corresponding to the current solution task information generated by the large language model based on the second prompt word template corresponding to the second agent, the received current solution task information, and the large language model; obtains a task execution result for the current solution task information from an application server corresponding to the user question based on the task execution information, and sends the task execution result to the first agent. So that the first agent generates the next solution task information through the large language model after obtaining the task execution result of the current solution task information, and the second prompt word template includes second toolset information; an acquisition module is used to obtain the task execution results of the multiple solution task information sent by the second agent; and a summary module is used to summarize the task execution results of the multiple solution task information based on the summarized task information to determine the reply information corresponding to the user question.In a fourth aspect, embodiments of the present disclosure provide a question-and-answer processing method for a second agent interacting with a first agent to complete the question-and-answer processing method, wherein the first agent and the second agent share a same large language model. The method comprises: receiving current solution task information corresponding to a user question sent by the first agent; wherein the first agent, based on the user question, a first prompt word template corresponding to the first agent, and the large language model, obtains, in sequence by the large language model, multiple solution task information that needs to be executed by the second agent and summary task information that needs to be executed by the first agent, wherein the first prompt word template includes first toolset information, and the current solution task information is the solution task information currently sent to the second agent from the multiple solution task information; splicing the current solution task information into a second prompt word template corresponding to the second agent, thereby obtaining task execution information corresponding to the current solution task information generated by the large language model based on the second toolset information included in the second prompt word template and the large language model; and obtaining, based on the task execution information, a task execution result for the current solution task information from an application server corresponding to the user question; The task execution result of the current task information to be solved is sent to the first agent, so that the first agent summarizes the task execution results of the multiple task information to determine the reply information corresponding to the user question based on the summarized task information.In a fifth aspect, embodiments of the present disclosure provide a question-and-answer processing apparatus for a second agent that interacts with a first agent to complete question-and-answer processing, wherein the first and second agents share the same large language model. The apparatus comprises: a receiving module configured to receive current solution task information corresponding to a user question, sent by the first agent; wherein the first agent, based on the user question, a first prompt word template corresponding to the first agent, and the large language model, obtains, sequentially generated by the large language model, multiple solution task information requiring execution by the second agent and summary task information requiring execution by the first agent; the first prompt word template includes first toolset information; and the current solution task information is solution task information currently being sent to the second agent from the multiple solution task information; a processing module configured to splice the current solution task information into a second prompt word template corresponding to the second agent, obtain task execution information corresponding to the current solution task information, generated by the large language model based on the second toolset information included in the second prompt word template and the large language model, and obtain, based on the task execution information, a task execution result for the current solution task information from an application server corresponding to the user question; A sending module is configured to send the task execution result of the current task information to the first agent, so that the first agent, based on the summarized task information, aggregates the task execution results of the multiple task information to determine a response corresponding to the user question. In a sixth aspect, embodiments of the present disclosure provide an electronic device, comprising: a memory, a processor, and a communication interface; wherein the memory stores executable code, and when the processor executes the executable code, the processor implements at least the question-and-answer processing method described in the second or fourth aspect. In a seventh aspect, embodiments of the present disclosure provide a non-transitory machine-readable storage medium, wherein the non-transitory machine-readable storage medium stores executable code, and when the processor of the electronic device executes the executable code, the processor implements at least the question-and-answer processing method described in the second or fourth aspect. In an eighth aspect, embodiments of the present disclosure provide a computer program product, comprising a computer program, which, when executed by the processor, implements at least the question-and-answer processing method described in the second or fourth aspect. The embodiment of the present disclosure provides a question-answering processing system, which includes a first agent as a controlling agent, a second agent as a controlled agent, and an application server (such as an intelligent traffic server), and the first agent and the second agent share the same LLM.The first agent's primary function is to plan the step-by-step process for solving the user's problem. Specifically, it decomposes the user's problem into its own solving tasks. Each decomposed solution task is then dispatched to the second agent for execution, obtaining the corresponding task execution results. Ultimately, the first agent summarizes the execution results of each task to generate a response message for the user. Each time the first agent generates a solution task message and dispatches it to the second agent for execution to obtain the corresponding task execution results, it generates the next solution task message. This process continues iteratively until the user's problem is solved and a response message is generated. Specifically, after receiving the user's problem, the first agent can incorporate the user's problem into a first prompt word template corresponding to the first agent, generating corresponding prompt words for input into the LLM. The first prompt word template contains information about a first set of tools the first agent can use when solving the user's problem. This includes, for example, tools for invoking the second agent to complete a specific solution task and tools for summarizing the execution results of multiple solution tasks. Based on the prompt words sequentially input by the first agent, the LLM generates multiple task information corresponding to the user's question, which requires the second agent to execute, as well as summary task information required by the first agent. The first agent then sends the multiple task information to the second agent. The second agent then incorporates the received current task information into the second prompt word template corresponding to the second agent. Based on the second toolset information contained in the second prompt word template, it generates corresponding prompt words for input to the LLM. Based on the task execution information output by the LLM, the first agent obtains the task execution result of the current task information from the application server corresponding to the user's question and sends this task execution result to the first agent. After obtaining the task execution result for the current task information, the first agent can use the LLM to generate the next task information for the current task information in the multiple task information based on this task execution result and the first prompt word template. After the first agent decomposes the user problem-solving process into multiple problem-solving task information and controls the second agent to complete the execution of each problem-solving task information and obtains the corresponding task execution results, the first agent summarizes the task execution results of multiple problem-solving task information based on the summary task information generated by LLM to determine the response information corresponding to the user problem.The question-and-answer processing solution provided by the disclosed embodiments incorporates a first agent for controlling the overall user question-solving process and a second agent for performing specific functions, such as data collection. This approach hierarchically decomposes agents according to their roles, with different agents handling different types of tasks. The first agent is responsible for planning the overall user question-solving process. Based on the user question and the results obtained after each step of the task, it determines whether to call the second agent to complete the next task or to summarize and respond to existing results. The second agent is responsible for the actual task execution logic. This reduces the complexity of each step in the LLM's thinking, reduces the potential for errors, and improves the agent's problem-solving capabilities and scalability. The introduction of the second agent enables integration with different functional tools (such as APIs) within the application server. Compared to solutions based on natural language-to-SQL query conversion, this expands the capabilities of the LLM and covers a wider range of application scenarios. Furthermore, since the LLM has a limit on the length of input text, excessively long input text can reduce its decision-making capabilities. The agent needs to send information such as descriptions of available tools and input parameters to the LLM in text form. For example, the first and second prompt word templates mentioned above both contain information about the corresponding toolset. When a large number of tools are required, the text input to the LLM can be very lengthy, potentially even exceeding the maximum length allowed. However, in the embodiments of the present disclosure, by dividing agents into different categories, the tools are categorized. Each agent uses fewer tools, so the text input of the corresponding toolset information by each agent to the LLM is much shorter. This reduces the length of the LLM input text and ensures reliable operation of the LLM. To more clearly illustrate the technical solutions in the embodiments of the present disclosure, the following briefly describes the figures used in describing the embodiments. Obviously, the figures described below represent some embodiments of the present disclosure. Those skilled in the art can derive other figures based on these figures without inventive effort.Figure 1 is a schematic diagram of a hardware execution environment for a question-and-answer processing method provided in an embodiment of the present disclosure; Figure 2 is a schematic diagram of a cloud computing environment for a question-and-answer processing method provided in an embodiment of the present disclosure; Figure 3 is a schematic diagram of an application of a question-and-answer processing method provided in an embodiment of the present disclosure; Figure 4 is a schematic diagram of the components of a question-and-answer processing system provided in an embodiment of the present disclosure; Figure 5 is a flow chart of a question-and-answer processing method provided in an embodiment of the present disclosure; Figure 6 is a flow chart of a question-and-answer processing method provided in an embodiment of the present disclosure; Figure 7 is a schematic diagram of the structure of a question-and-answer processing device provided in an embodiment of the present disclosure; Figure 8 is a schematic diagram of the structure of another question-and-answer processing device provided in an embodiment of the present disclosure; and Figure 9 is a schematic diagram of the structure of an electronic device provided in an embodiment of the present disclosure. To further clarify the objectives, technical solutions, and advantages of the embodiments of the present disclosure, the technical solutions in the embodiments of the present disclosure will be described clearly and completely below in conjunction with the accompanying drawings. It should be understood that the described embodiments represent only a portion of the embodiments of the present disclosure, and are not intended to be exhaustive. All other embodiments derived by persons of ordinary skill in the art based on the embodiments of the present disclosure without inventive effort are intended to fall within the scope of protection of the present disclosure. It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, storage, and display) involved in the embodiments of this disclosure are all authorized by the user or fully authorized by all parties. The collection, use, and processing of relevant data must comply with the relevant laws, regulations, and standards of relevant countries and regions, and corresponding operation portals are provided for users to choose to authorize or reject. The following detailed description of some embodiments of this disclosure is provided in conjunction with the accompanying drawings. The following embodiments and features may be combined unless there is a conflict between the embodiments. Furthermore, the sequence of steps in the following method embodiments is provided as an example and is not a strict limitation. The following terms and concepts involved in the embodiments of this disclosure are explained below: Large Language Model (LLM): Large Language Model I (LLM). Prompt: The prompt of the LLM is the input for interacting with the LLM, typically containing instructions, context, input and output formats, and other information. Token: The smallest unit of text processed by the LLM.Artificial Intelligence Agent (AI Agent) is an agent system that controls LLM to solve problems. It is an intelligent entity capable of perceiving the environment, making decisions, and executing actions. AI agents can be physical entities (such as robots) or virtual entities (such as computer programs). They achieve specific tasks or goals by perceiving information in the environment, making decisions, and executing actions. Application Programming Interface (API): In this article, it refers to the channel / interface through which AI agents interact with service platforms such as transportation. With the continuous development of LLM technology and the significant improvement in various aspects of model performance, various LLM-based applications are attracting widespread attention. Among them, LLM-driven AI agents (AI-Agents) are a key research direction. Through natural language dialogue, AI agents enable LLMs to perform task decomposition, call external tools, and analyze results, thereby achieving more general problem solving. Currently, many application areas use single-agent solutions: a single agent completes task decomposition, tool invocation, and result summary. However, this approach has the disadvantage that, when multiple complex tasks of different categories need to be invoked, the agent's attention is focused on processing the results of tool invocations, and it is easy to neglect the overall problem-solving situation. This may lead to deviations from the initial goal or neglect of previous solution results. In other words, it cannot effectively complete multi-step continuous solutions. In light of this, the disclosed embodiments provide a multi-agent collaborative solution for solving user problems. Specifically, by introducing a first agent (the master agent) responsible for controlling the overall problem-solving process and a second agent (the functional agent) responsible for completing specific functions, such as data collection, the agents are hierarchically broken down according to their roles, with different agents handling different types of tasks. The first agent is responsible for planning the overall problem-solving process. Based on the user question and the results of each task, it determines whether to call the second agent to complete the next task or to summarize and respond to existing results. The second agent is responsible for the actual task execution logic. This reduces the complexity of each LLM step, reduces the potential for errors, and improves the problem-solving capabilities and scalability of the agents. The question-answering solution provided by the disclosed embodiments is applicable to various application scenarios, including but not limited to intelligent transportation, database query, and e-commerce.In various application scenarios, users may generate questions that they wish the Q&A processing system to automatically answer. This system, by connecting to a corresponding backend service system (such as an intelligent transportation server), uses multiple agents within the system to gradually solve the user's question, gradually obtaining the required data from the backend service system and ultimately summarizing the answer to the user's question. Automatically answering user questions is essentially a Q&A task, and the gradual solution of the user's question is the gradual resolution of this Q&A task. This gradual resolution primarily involves the master agent splitting the Q&A task into multiple subtasks (hereinafter referred to as resolution tasks). Functional agents execute the subtasks to obtain corresponding results, and the master agent summarizes the results of each subtask to determine the answer to the user's question. The following describes the Q&A processing solution provided by the present disclosure. FIG1 is a schematic diagram of a hardware execution environment for a question-and-answer processing method provided in an embodiment of the present disclosure. As shown in FIG1 , the hardware execution environment of the question-and-answer processing method may be composed of a user device 101, a question-and-answer server 102, and an application server 103. The user device 101 is in communication with the question-and-answer server 102, and the question-and-answer server 102 is in communication with the application server 103. The user device 101 may be a standalone terminal device, such as a smartphone, tablet computer, or personal computer (PC), or may be two or more terminal devices used in conjunction, such as an extended reality device 101 a and a smartphone 101 b. The extended reality device 101 a may be a virtual reality device, an augmented reality device, or the like. The question-and-answer server 102 may be a server of an application provider or a cloud server of a cloud service provider. The application server 103 may be a server of an application provider or a cloud server of a cloud service provider. From a physical location perspective, the question-and-answer server 102 and the application server 103 can be located in the same or different physical hosts. The question-and-answer server 102 refers to a server that executes the question-and-answer processing solution provided in the embodiments of the present disclosure, while the application server 103 refers to a server that provides certain application services, such as an intelligent transportation server, an e-commerce server, or a video server, and needs to interact with the question-and-answer server 102 to obtain data related to user questions.Optionally, when user device 101 comprises an extended reality (XR) device 101a and a smartphone 101b, the Q&A processing method may be executed as follows: a user enters a question in natural language into smartphone 101b, smartphone 101b sends the question to Q&A server 102, Q&A server 102 executes the Q&A processing method provided in the embodiments of the present disclosure, interacts with application server 103, ultimately obtains a response corresponding to the user's question, and sends the response to smartphone 101b. Smartphone 101b then sends the response to XR device 101a for display. In actual applications, Q&A server 102 may be an independent physical server or physical server cluster maintained by an application provider, or a cloud server (referred to as a computing node) maintained by a cloud service provider. In the cloud computing environment shown in Figure 2, several distributed computing nodes (cloud servers) (illustrated as 201-1, 201-2, and so on) can be included. Each computing node has processing resources such as computing and storage. In a cloud computing environment, multiple computing nodes can be organized to provide a service. Of course, a single computing node can also provide one or more services, such as Service A, Service B, Service C, and Service D illustrated in Figure 2. The cloud computing environment can provide these services by providing a service interface 202, which client devices call to access the corresponding service. Service interface 202 includes a software development kit (SDK) and an application programming interface (API).
[0002] App Interface (API) and other forms. These services are deployed based on various virtualization technologies supported by cloud computing environments, such as virtual machines and containers. Taking container-based virtualization as an example, several containers corresponding to a service can be assembled into a container group (pod). For example, service B, shown in Figure 2, can be configured with one or more pods, each of which can include a proxy and one or more containers. The one or more containers in the pod are used to process requests related to one or more corresponding functions of the service, and the proxy in the pod is used to control network functions related to the service, such as routing and load balancing. During operation, executing a request from a user device may require invoking one or more services in the cloud computing environment, and executing one or more functions of one service may require invoking one or more functions of another service. As shown in Figure 2, after receiving a request from a user device, service A can invoke service B, and service B can request service D to execute one or more functions. In the aforementioned cloud computing environment, the present disclosure provides an application diagram of a question-and-answer processing method, as shown in Figure 3. In Figure 3, the cloud computing environment provides the following three services: a first agent, a second agent, and an LLM. These three services can reside on the same computing node or on different computing nodes. The first agent can call the second agent and the LLM, and the second agent can call the LLM. Furthermore, the second agent can interact with the application server 103, and the first agent can interact with the user device 101. After the first agent receives a user question sent by a user through the user device 101, the first agent, the second agent, and the LLM ultimately determine a response to the user question based on the question-and-answer processing solution provided by the embodiments of the present disclosure. The first agent then feeds this response back to the user device 101. Figure 4 illustrates the operating principle of a question-and-answer processing system provided by the embodiments of the present disclosure. As shown in Figure 4, the question-and-answer processing system includes a first agent, a second agent, an LLM, and an application server. The first and second agents share the same LLM. The first agent, as the master agent, plans and controls the overall process of solving the user question.As shown in Figure 4, in an optional embodiment, the first agent may include a background knowledge acquisition module. In this case, the first agent first receives a task from user input, namely, a user question. The background knowledge acquisition module then acquires background knowledge related to the user question from an external knowledge base. Based on this, the agent interacts with the LLM through a "think / act / feedback" mechanism to gradually solve the user's problem. After completing the task, the first agent invokes the summary module to summarize the execution results of all intermediate steps in the current task, generating and outputting the final response information. As described in the prior art, the main functional components of an agent include memory, thinking (i.e., planning), action, feedback (reflection), and tool use, corresponding to the "think / act / feedback" mechanism of the first and second agents in the disclosed embodiments. Essentially, the "think / act / feedback" mechanism is a prompt word template for solving a task, and problem solving is achieved through repeated iterations of the "think / act / feedback" cycle. Specifically, for the first agent, the mechanism's main functions at each step are as follows: Thinking: The LLM considers the user question and completed solution steps, generating a summary. Action: The LLM generates the tool to be invoked for the next action, along with the tool's input parameters. Feedback: The backend (referring to the first agent) invokes the corresponding tool or other agent based on the "action" output by the LLM, and obtains the execution result after completion. The execution process of the first agent's "think / act / feedback" mechanism can be simply described as follows: The first agent receives the user question, concatenates (i.e., inserts) the user question into the first prompt word template corresponding to the first agent. Based on the first toolset information contained in the first prompt word template and the LLM, the first agent obtains the multiple solution task information generated by the LLM, which requires the second agent to execute, as well as the summary task information required by the first agent. The first agent then sequentially sends the multiple solution task information to the second agent. Based on the summary task information, the first agent aggregates (i.e., summarizes) the task execution results of the multiple solution task information to determine the response information corresponding to the user question. The execution results of the multiple task information solutions are fed back to the first agent by the second agent. The above execution process is briefly illustrated with reference to FIG4 .As shown in Figure 4 , the execution process of the first agent's "think / action / feedback" mechanism includes the following: First, the first agent obtains relevant information from multiple dimensions (such as the user question, background knowledge, historical conversation records obtained from the memory module, and current thought records, as shown in Figure 4 ). This information is then combined into a first prompt word template to generate the first prompt word generated for the current round of the "think / action / feedback" mechanism. This prompt word template includes information about the first tool set available to the first agent. In other words, the LLM can determine from the first prompt word which tools it can use to generate "action" content. Figure 4 illustrates three tools available to the first agent: the "summary" tool represented by the summary module, the "askuser" tool represented by the interaction module, and the "data collection" tool, which invokes a second agent to collect data. Here, it is assumed that the second agent provides data collection functionality. If the LLM-generated thoughts and actions indicate the invocation of a second agent, the generated results are referred to as "solution task information." If they indicate a summary, the generated results are referred to as "summary task information." If they indicate interaction with the user, the generated results are referred to as "interaction task information." It should be noted that the "feedback" content generated by the LLM actually only includes a feedback identifier and does not contain the actual feedback result. For example, the feedback identifier is -feedback. As shown in Figure 4, when the first agent generates the first prompt word, it can use historical conversation records and current thought records obtained from the memory module. Historical conversation records are questions and corresponding responses preceding the current user question in multiple rounds of conversation with the same user. The current thought record refers to the data set generated by the first agent after each round of the "think / action / feedback" mechanism as it gradually solves the user's problem. This data set forms the final thought, action, and feedback from that round, serving as the current thought record for the next round of the mechanism. At this point, the feedback data received by the first agent from the previous round is already filled in after the feedback indicator. The functions of the first agent have been briefly introduced above. Now, the functions of the second agent will be explained. Compared to the first agent, which functions as the master agent, the second agent is a controlled agent, providing a specific function. To meet the needs of different application scenarios, one or more agents with different functions can be configured. Figure 4 illustrates the second agent that provides data collection.In some application scenarios, for example, agents providing data analysis and prediction capabilities, agents providing data visualization capabilities, and so on, may also be included. Different functional agents utilize different tool sets. Assuming the second agent is an agent providing data collection (referred to as a data agent), the second agent is responsible for receiving data collection-related task information issued by the first agent, retrieving the corresponding data (the raw data retrieved from the application server, as shown in Figure 4) from an external application server (e.g., an application server containing a database), and feeding it back to the first agent. The data acquisition process is carried out through a "think / act / feedback" mechanism. The "think / act / feedback" mechanism of the second agent is similar to that of the first agent. In summary, it also generates information containing "think, act, and feedback" content through the LLM based on a second prompt word template containing the second tool set information. Based on the action content indicated in the prompt word template, the second agent invokes the corresponding tool to retrieve the corresponding data from the application server. The main differences between the first and second prompt templates are: first, the toolsets used are different; second, the information required to be incorporated into the prompt templates is different. The first toolset of the first agent, for example, includes tools such as summary, data, and askuser, used to control the macro-level process of solving user problems. The second toolset of the second agent may include various API information for interfacing with application servers, such as different APIs for querying different data contents. Each API information piece may include information such as API function descriptions, input parameter descriptions and formats, and return result descriptions. As mentioned above, the information incorporated into the first prompt template may include user questions, historical conversation records, background knowledge, and the first agent's current thought process. The information incorporated into the second prompt template may include the task information (including thoughts, actions, and feedback) currently being sent by the first agent to the second agent, as well as the second agent's current thought process. Among them, the current thinking record corresponding to the second intelligent agent refers to: assuming that the second intelligent agent also needs to gradually solve a certain solution task information issued by the first intelligent agent in the process of executing the solution task information (similar to the process of the first intelligent agent gradually solving the user problem), then in the solution process, after each round of the "thinking / action / feedback" mechanism is executed, the thinking, action and feedback content finally obtained in this round are formed into a data group as the current thinking record for the next round of execution of the mechanism. At this time, the feedback data content obtained by the first intelligent agent has been filled in after the feedback identifier of the previous round.Based on this, it can be seen that the first and second agents may each have their own "current thought records," which are stored independently. In summary, the second agent's operation process includes: splicing the current task information received from the first agent into the second prompt word template corresponding to the second agent; obtaining task execution information corresponding to the current task information generated by the LLM based on the second toolset information and LLM contained in the second prompt word template; obtaining the task execution result of the current task information from the application server corresponding to the user's question based on the task execution information; and sending this task execution result to the first agent. After obtaining the task execution result of the current task information, the first agent generates the next task information through the LLM. The task execution information describes the operation triggered by the application on the application server to obtain the corresponding task execution result. The task execution information includes the thought and action content, as well as a feedback identifier. At this point, the feedback content is empty. After the second agent performs the corresponding action based on the action content, the task execution result is obtained and filled into the feedback identifier, forming the feedback content. The memory module illustrated in Figure 4 includes three submodules: historical conversation records, current thought records (the current thought records corresponding to the first agent and the current thought records corresponding to the second agent), and raw data (data obtained by the second agent from the application server). The historical conversation record module stores historical question-and-answer data between the user and the first agent, including each "user question" and "returned reply information," typically from previous rounds of multi-round conversations. The current thought record module stores the intermediate execution results generated by the "think / execute / feedback" mechanism in the current round of conversation between the first and second agents. The raw data module stores the raw results obtained by the second agent from the application server during the query process, which is used for front-end rendering and display. This means that the final output to the user, in addition to the reply information, may also include a visual display of the raw data. The summary module in Figure 4 summarizes the agent's solution process and provides the final reply information. Its input includes the user question, background knowledge, and the intermediate execution results generated by the first agent in each round of the "think / execute / feedback" mechanism. The summary module acquires and combines this information, allowing the LLM to provide the final response information. It is understood that after the response information corresponding to the current user question is determined, the question-answer pair can be stored in the historical conversation record in the memory module to be used as the historical conversation record for the user's next round of user questions.In summary, the question-and-answer processing solution provided by the embodiments of the present disclosure introduces a first agent for controlling the overall user question-solving process and a second agent for performing specific functions, such as data collection. This breaks down the agents into hierarchical roles, with different agents handling different types of tasks. The first agent is responsible for planning the overall user question-solving process. Based on the user question and the results obtained after each task, it determines whether to call the second agent to complete the next task or to summarize and respond to existing results. The second agent is responsible for the actual task-solving logic. This reduces the complexity of each step in the LLM's thinking, reduces the potential for errors, and improves the agent's problem-solving capabilities and scalability. The introduction of the second agent enables integration with different functional tools (such as APIs) within the application server. Compared to solutions based on natural language-to-SQL query conversion, this expands the capabilities of the LLM and covers a wider range of application scenarios. Furthermore, since the LLM has a limit on the length of input text, excessively long input text can reduce its decision-making capabilities. The agent needs to send information such as descriptions of available tools and input parameters to the LLM in text form. For example, the first and second prompt word templates mentioned above both contain information about the corresponding toolset. When a large number of tools are required, the text input to the LLM can be very lengthy, potentially even exceeding the maximum length allowed. However, in the disclosed embodiment, by dividing agents into different categories, the tools are categorized. Each agent uses fewer tools, so the text input of the corresponding toolset information to the LLM by each agent is much shorter. This reduces the length of the LLM input text and ensures reliable operation of the LLM. To facilitate understanding, the following example illustrates the multi-round execution process of the first agent's "think / execute / feedback" mechanism. Assuming the user's question is "Query today's congestion index for Yuhang District and Xihu District," the step-by-step solution process is as follows:
[0003] 1. Thinking (generated by LLM): The user needs to query the congestion index of Yuhang District and Xihu District today. First, query the congestion index of Yuhang District;
[0004] 2. Action (generated by LLM): [data] Query the congestion index of Yuhang District today;
[0005] 3. Feedback (the first agent parses the previous "action" text, invokes the tool based on the data, and calls the second agent to complete the query): Today's congestion index in Yuhang District is 1.3;
[0006] 4. Thinking (generated by LLM. Note that the prompt words entered into the LLM at this time will include the [Thinking, Action, Feedback] data set generated in the previous round): The congestion index for Yuhang District has been obtained. Now query the congestion index for Xihu District.
[0007] 5. Action (generated by LLM): [data] Query today's congestion index in Xihu District;
[0008] 6. Feedback (obtained after the first agent parses the previous "action" text and calls the second agent to complete the query): Today's congestion index in Xihu District is 1.41;
[0009] 7. Thinking (generated by LLM. Note that the prompt words entered into LLM at this time will include the [Thinking, Action, Feedback] data set generated in the previous round): The congestion index of Yuhang District and Xihu District has been obtained, and the results can now be returned;
[0010] 8. Action (generated by LLM): [summary] Today's congestion index in Yuhang District is 1.3, and today's congestion index in Xihu District is 1.41;
[0011] 9. Feedback: The first agent recognizes the "summary" field, invokes the summary module, obtains a response, and returns it to the user. In the above example, the format of the "Action" information is: Action: [Tool ID] Action Input. The tool ID, also known as the tool name, indicates which tool is to be invoked for this action, while the "Action Input" primarily provides the tool's input parameters. For example, in the above example, the natural language description "Query today's congestion index in Xihu District" serves as the input parameter for the tool data, indicating that the second agent should be invoked to perform the task "Query today's congestion index in Xihu District." In practical applications, the task information input to the second agent may include, but is not limited to, two components: a thought and an action. Referring to Figure 4, the overall execution flow of the question-and-answer processing method provided by the present embodiment is described below. The steps may include the following: Step 1: User Input and Background Knowledge Extraction. The user enters a character string as a question in natural language text. After receiving the question string, the first agent first parses it. If the string is a command indicating clearing history, such as the command " / Clear", the user's historical conversation records are cleared from the memory module, and the final result "History cleared successfully" is directly returned, ending the current round of question-and-answering. If the string is not a clearing command, the normal question-solving process begins. The user's question may be a relatively clear query, such as: "What is the congestion index in Hangzhou today?" or "Query the travel volume from Xihu District to Yuhang District." It may also be an open-ended analytical question, such as: "What is the traffic situation in Hangzhou today?" In the disclosed embodiment, historical conversation records and background knowledge are used to enable the LLM to more accurately understand the user's current question intent and ultimately provide a response that conforms to common language conventions. In this step, the user's question string can be input into the background knowledge acquisition module of the first agent, which then searches for background knowledge related to the user's question from an external knowledge base. The process flow is as follows: Step 1.1: User Intent Identification. This step is optional and is used to rewrite the user's input question into a more standardized format. The user question entered is represented as S'. The first agent sends Su to the LLM, requesting it to parse the domain and keywords of the user question and refine the question. For example, if the user question Su = "Is Hangzhou congested today?", the LLM returns the recognized domain as "traffic" and the keywords as "Hangzhou" and "congestion situation". It then refines the question to obtain s'u = "Query the congestion situation in Hangzhou today."Thus, LLM rewrites the user's more colloquial question expression into a more standardized form. Step 1.2: Extract relevant tools. Based on the domain identified in step 1.1, determine the functional agent (i.e., the second agent) corresponding to that domain and the corresponding toolset information for the functional agent. For example, when the identified domain is "transportation," the second agent determined is a functional agent applicable to the transportation field, and the toolset determined is relevant tools for connecting to the intelligent transportation service system, such as various APIs. It is understandable that if only functional agents applicable to a specific domain and the corresponding toolset are provided, then step 1.2 is unnecessary. Step 1.3: Extract information from the external knowledge base. Assume that the external knowledge base K contains several (e.g., m) background knowledge paragraphs: K = . The title and content of a knowledge paragraph in K are, for example: ([k] = "Analyzing the city's congestion situation", [V] = "Analyzing the congestion situation in a city, proceeding from the following three steps: 1. Querying the city's congestion index; 2. Querying the congestion index of each area in the city; 3. Querying the city's most congested roads."). The first agent can include a vector representation model. After each paragraph title in K is processed by the vector representation model Os, the resulting high-dimensional vector group is X. k = [x kl , x k2 , ... , x kra ] oHere, s′ represents the vector representation of the knowledge passage title. Assume that a similarity threshold is set. The first agent inputs the above s′ into the vector representation model, obtaining the corresponding vector representation: Xs%. Next, the cosine similarity between Xs′ and each element in Xk is calculated. If the cosine similarity with Xkj is greater than the threshold, the corresponding knowledge passage content Vj is determined as the background knowledge relevant to the user's question obtained from the query. Alternatively, if the cosine similarity with Xkj is greater than the threshold, and the cosine similarity with Xkj is greater than the cosine similarity with the titles of other knowledge passages, the corresponding knowledge passage content Vj is determined as the background knowledge relevant to the user's question obtained from the query. Alternatively, the first n (n is a set value, such as 3) knowledge passages with cosine similarities greater than the threshold are selected as the background knowledge relevant to the user's question. If there are no knowledge passages with a cosine similarity greater than the threshold, the returned background knowledge is an empty string. The use of this background knowledge information can help the large language model better understand the current user question, thereby facilitating more accurate responses. Step 2: First Agent Task Execution. After completing background knowledge extraction, the first agent begins parsing and executing the user's problem-solving task. The specific steps are as follows: Step 2.1: First Agent Initialization. The first agent clears the user's current thought record corresponding to the first agent from the memory module, as well as the original data record corresponding to the user. Step 2.2: LLM Prompt Term Generation. The first agent combines the user's question, background knowledge, the user's historical conversation records, and the first agent's current thought record into the first prompt term template corresponding to the first agent to generate the first prompt term. Since the first agent's current thought record has been cleared during the initialization process in Step 2.1 during the first round of execution, an empty string is added to the first prompt term template. Background knowledge and historical conversation records are optional. For ease of understanding, the following example of a first prompt term template is provided using an intelligent transportation scenario: You are an expert in the transportation field. Please answer the user's question, referencing specific data. First, the following background knowledge may be relevant to the question:
[0012]
Background knowledge of splicing
[0013] [Joining Historical Conversation Records] You need to answer in steps. Each step consists of three parts: "Thinking," "Action," and "Feedback." The format is as follows: Thinking: Summarize historical steps and results. Action: [Tool Name] Action input information. Feedback: Contains the execution results of the action. The optional tools for each action step are as follows:
[0014] [summary] : When all relevant information about the question has been obtained, use this tool to summarize and end this round of Q&A.
[0015] [data]: A data acquisition tool that can be called to query specific data. The input requires a complete description of the query object, quantity, metrics, and other requirements in Chinese. [askuser]: When a tool fails or the instructions are unclear, use this tool to request further user feedback. Let's get started! Question: [Concatenate user questions] Current thought record: [Concatenate current thought record] In the example of the first prompt word template above, the content within [] will be updated, concatenating the corresponding user question, background knowledge, historical conversation records, and the first agent's current thought record, while the rest of the content remains unchanged. The first prompt word template also provides the first toolset information for the example above. This shows that the first prompt word template is used to indicate the format of the task information required to be generated at each step as the LLM solves the user question. Each task information includes thought information, action information, and a feedback indicator. Thought information describes the task to be completed, action information describes the tool selected from the first toolset information and its input parameters to execute the task, and the feedback indicator indicates where to enter the task's execution result. The LLM does not output this execution result. Furthermore, in the example above of "Querying today's congestion index for Yuhang District and Xihu District," when the first agent invokes the LLM in the first round, the "Current Thought Record" added to the first prompt template is empty, so the generated first prompt does not contain the relevant content of this field. When the first agent invokes the LLM in the second round, the "Current Thought Record" added to the first prompt template is "Thought: User needs to query today's congestion index for Yuhang District and Xihu District, first query the congestion index for Yuhang District; Action: [data] Query today's congestion index for Yuhang District; Feedback: Today's congestion index for Yuhang District is 1.3." Therefore, the generated first prompt will contain the relevant content of this field. Step 2.3: The LLM generates an action plan. The first agent sends the first prompt generated in step 2.2 to the LLM. The LLM outputs "thought information" and "action information" in natural language as the action to be performed by the first agent in this round. For example, in the example above of "querying today's congestion index in Yuhang District and Xihu District," when the first agent calls the LLM in the first round, the LLM outputs the following task information: Thought: The user needs to query today's congestion index in Yuhang District and Xihu District. First, query the congestion index in Yuhang District; Action: [data] Query today's congestion index in Yuhang District; Feedback: Step 2.4: The first agent executes the action.In practical applications, the first agent can parse the content of the task-solving information continuously output by the LLM in real time. When the parsed output contains the word "feedback," the LLM can be controlled to pause its operation. The first agent parses the LLM's output in step 2.3 to obtain the "thinking information" and "action information" for this round. If the result returned in step 2.3 contains format errors or missing fields, resulting in a parsing failure, the first agent uses the corresponding error information as the "feedback" result for this round and skips to step 2.5. If the "action information" returned in step 2.3 calls the "summary" tool, the first agent skips to step 3. If the "action information" returned in step 2.3 calls the "askuser" tool, the first agent enters the interactive module process, which uses the "action input information" contained in the "action information" as the content of the question to be posed to the user. In the above example, the action input information is: Query today's congestion index in Yuhang District. The interaction module receives the user's response. The first agent uses the user's response as the "feedback" result for this round and proceeds to step 2.5. If the tool called in the "action information" returned in step 2.3 is "data, i.e., calling the second agent," the task execution process for the second agent begins. The first agent uses the "action input information" generated in this round (for example, querying today's congestion index in Yuhang District) as input to the second agent's task solution. Alternatively, the first agent uses both the "thinking information" and "action information" generated in this round as input to the second agent's task solution, and proceeds to step 2.4.a1. oThe above describes the execution process of a round of the "think / action / feedback" mechanism for the first agent. Next, we will describe the execution process of a round of the "think / action / feedback" mechanism for the second agent. Step 2.4.a1: Initialize the second agent. The second agent clears the current thought record corresponding to the second agent in the memory module. Step 2.4.a2: Generate LLM prompts. The second agent concatenates the second agent's input, the current thought record corresponding to the second agent, and the list of available APIs into the second prompt template corresponding to the second agent to generate the second prompt. Optionally, the list of available APIs, serving as the second toolset information, can be directly included in the second prompt template, eliminating the need for concatenation. The second agent's input is the solution task information output by the first agent. Similar to the current thought record corresponding to the first agent, when the second agent is first invoked, the current thought record corresponding to the second agent is empty. In fact, the current thinking record corresponding to the second agent and the current thinking record corresponding to the first agent are both used to store information about the intermediate execution results generated when the corresponding agents execute the "thinking / action / feedback" mechanism in each round. The second prompt word template is relatively similar to the first prompt word template. The main differences are as follows: First, the content that needs to be spliced is different. The second prompt word template can include the following splicing fields: [Splicing current thinking record], [Splicing solution task information] Second, the toolset information is different. The second toolset information included in the second prompt word template can be various API information required for interaction with the application server, such as API 1 for querying the congestion index, API 2 for querying the travel volume, and API 3 for querying traffic accidents. OIn addition, the second toolset may also include a special API (return final answer) that is used to provide feedback to the first agent regarding the task execution results after the second agent completes processing the currently received task information. Third, the action information format is different. The action information format in the first prompt template is: [Tool Name] Action Input Information, where the action input information is in natural language text format. In the second prompt template, the action input information format can be JSON, for example, {"Parameter 1": "Value 1"; "Parameter 2": "Value 2"}. For example, Parameter 1 = Region, Value 1 = Xihu District. Step 2.4.a3: The LLM generates an action plan. The second agent sends the second prompt generated in step 2.4.a2 to the LLM. The LLM outputs "Thinking Information" and "Action Information" in natural language format as the actions to be performed by the second agent in this round. Step 2.4.a4: The second agent executes the action. The second agent parses the "thinking information" and "action information" output in step 2.4.a3. If the parsing fails, the second agent uses the corresponding error information as the "feedback" result for this round and jumps to step 2.4.a5. If the tool called in the "action information" returned in step 2.4.a3 is "API x, return final answer," it indicates that the second agent has completed the current task. The "action input information" contained in the "action information" is sent to the first agent as the second agent's output, and the process jumps to step 2.5. At this point, the "action input information" is the task execution result corresponding to the current task. If the tool called in the "action information" returned in step 2.4.a3 is the name of another API in the second tool set, then after checking to determine that the parameter format in the "action input information" is the correct parameter format for the API, the corresponding API is called with the parameters given in the "action input information". The result returned by the API call is used as the "feedback" result of this round, and the raw data obtained from the API query is stored in the memory module, and the process jumps to step 2.4.a5. If the call parameters are incorrect or other abnormal information occurs, the error information is used as the "feedback" result of this round, and the process jumps to step 2.4.a5. oStep 2.4.a5: Save the intermediate step results of the second agent. The second agent's "thinking information," "action information," and "feedback information" for this round are used as a set of data for this round and added to the current thinking record corresponding to the second agent. If the number of data sets in the current thinking record corresponding to the second agent is greater than the set threshold, it is considered that the second agent still cannot find the correct result after multiple cycles. This error information is sent to the first agent as the final output of the second agent, and the process jumps to step 2.5; otherwise, the process jumps to step 2.4.a2 to execute the next round of iteration. It should be noted that jumping to step 2.4.a2 at this time often corresponds to a situation where the second agent needs to execute multiple rounds of the "thinking / action / feedback" mechanism to complete the currently received solution task information. If the action "calling API x" is executed, the process will not jump to step 2.4.a2 again. oStep 2.5: The first agent saves the intermediate step results. The first agent adds the "thought information," "action information," and "feedback information" of this round as a set of data to its corresponding current thought record. If the number of data sets in its current thought record exceeds the set threshold, the first agent is deemed unable to find the correct answer after multiple iterations and proceeds to step 3. Otherwise, it proceeds to step 2.2 and performs the next iteration. Step 3: Task Summary and Return Result: When the first agent reaches a certain round and the tool name parsed from the corresponding "action information" is "summary," it determines that the task of solving the user's problem has been completed. At this point, the summary module is called to summarize and analyze the execution results of each intermediate step, that is, the multiple task execution results sent to the second agent, to generate the final response information for the user. Step 3.1: Summary prompt word generation. The first agent adds the user question, the background knowledge extracted in step 1.3, and the execution results of each intermediate step stored in step 2.5—that is, the task execution results corresponding to each solution task information stored in the first agent's current thinking record—to a third prompt word template to generate a third prompt word that guides the LLM to summarize. Step 3.2: Generate the final result. The first agent sends the third prompt word generated in step 3.1 to the LLM, which then outputs a response message to the user. Optionally, the first agent can also retrieve the raw data obtained from the application server from the memory module and feed this raw data back to the user, or visualize this raw data in some form of graphical representation for output to the user. As can be seen from the above description, the LLM in the disclosed embodiments primarily plays a planning and decision-making role. At each step in problem solving, the LLM is provided with information including the user question, relevant background knowledge, available tools, and completed solution steps. The LLM makes decisions based on this information and generates the next step of thinking and action. Below is an example of a third prompt template, which might consist of the following information: You are a robot based on a large language model, organizing query results and answering user questions. Possibly relevant background knowledge (ignore if irrelevant to the original question): [Concatenate background knowledge] Query results (which may come from different channels and methods): [Concatenate feedback from each step] Original question: [Concatenate user question] Now, based on these results, provide a detailed answer to the original question. Be professional and rigorous, do not deviate from the question, and ignore irrelevant results.If the query result does not contain relevant knowledge, the model's own knowledge is used to answer the question. Please answer in Chinese. The answer to the original question is as follows: Below are two examples of the execution process of the second agent's "think / act / feedback" mechanism. In the first, assume that the task information received from the first agent is simplified as: "Query the congestion index of Yuhang District today." The task execution information output by the LLM based on the corresponding second prompt is as follows: Think: Query the congestion index of Yuhang District today; Action: [API 1] {Region: Yuhang District; Time: Today}; Feedback: Here, assume that AP 11 is the API used to query the congestion index. When the second agent calls API 11 based on the above action information to query the congestion index of Yuhang District from the application server, the feedback result is: The congestion index of Yuhang District today is 1.3. In the above example, since the task is relatively simple, a single API call suffices to query the corresponding data. Subsequently, the next round of the "think / act / feedback" mechanism results in the following task execution information: Think: Today's congestion index for Yuhang District has been queried, and the result can be returned; Action: [API x] {Today's congestion index for Yuhang District is 1.3}; Feedback: Today's congestion index for Yuhang District is 1.3. In the second scenario, assuming the task information received from the first agent is simplified to "Query today's congestion index and traffic accidents in Yuhang District," the second agent will need to perform three rounds of the "think / act / feedback" mechanism to obtain the final result, as follows:
[0016] 1. Question: Query the congestion index of Yuhang District today. 2. Action: [AP I 1] {Region: Yuhang District; Time: Today};
[0017] 3. Feedback: Today's congestion index in Yuhang District is 1.3;
[0018] 4. Consider: Now that we have the congestion index for Yuhang District, we need to query traffic accidents that occurred in Yuhang District today.
[0019] 5. Action: [AP I 2] {Region: Yuhang District; Time: Today};
[0020] 6. Feedback: Three traffic accidents occurred in Yuhang District today;
[0021] 7. Thinking: Now that we have the congestion index and traffic accidents for Yuhang District, we can return the results.
[0022] 8. Action: [AP I x] {The congestion index in Yuhang District today is 1.3, and three traffic accidents occurred in Yuhang District today};
[0023] 9. Feedback: Today's congestion index in Yuhang District is 1.3, and three traffic accidents occurred in Yuhang District today. Assume that API 2 is the API used to query traffic accidents, and API x is the API used to return the final answer. In summary, the question-and-answer processing solution provided by the disclosed embodiments breaks down the user's question step by step. The first agent is responsible for planning the overall solution process for the user's question. Based on the user's question and the current available data, it determines whether to call the second agent to complete a sub-task or to call the summary module to summarize the execution results of existing sub-tasks and obtain a response. The second agent is responsible for the actual task execution logic, reducing the complexity of each step of the LLM's thinking and minimizing the potential for error. Furthermore, by introducing an external knowledge base, when a user enters a question, background knowledge related to the user's question is obtained and inserted into the prompt word. This background knowledge can include relevant concepts or expert problem-solving ideas, thereby enhancing the LLM's ability to analyze and solve problems. The second agent can design its own processing flow based on its needs. It interacts independently with the LLM, has its own prompts, and can share the first agent's historical thought process (the task information sent to the second agent can include the first agent's thought process information). This can help mitigate the impact of the LLM's token restrictions on complex tasks. Figure 5 is a flowchart of a question-and-answer processing method provided in an embodiment of the present disclosure. This method can be executed by the first agent described above. As shown in Figure 5, the method includes the following steps:
[0024] 501. Receive user questions.
[0025] 502. The user question is spliced into the first prompt word template corresponding to the first agent. Based on the first toolset information and the LLM included in the first prompt word template, the LLM is used to obtain information about multiple solution tasks that require execution by the second agent and summary task information that requires execution by the first agent. Optionally, the first agent may also obtain background knowledge information from an external knowledge base whose similarity to the user question meets set conditions, and / or obtain historical conversation records corresponding to the user question, and splice the background knowledge information and / or historical conversation records into the first prompt word template. Historical conversation records are questions and corresponding responses preceding the user question in multiple rounds of conversation with the same user.
[0026] 503. Send multiple solution task information to the second agent in sequence.
[0027] 504. Obtain the task execution results of the multiple task information sent by the second agent. Based on the second prompt word template corresponding to the second agent, the received current task information, and the LLM, the second agent obtains the task execution information corresponding to the current task information generated by the LLM. Based on the task execution information, the second agent obtains the task execution result for the current task information from the application server corresponding to the user's question. The task execution result is then sent to the first agent. After obtaining the task execution result for the current task information, the first agent generates the next task information through the LLM. The second prompt word template includes the second toolset information. This indicates that the multiple task information generated by the first agent are not generated simultaneously, but rather incrementally in an iterative process. That is, the process of generating the above-mentioned multiple solution task information includes: after obtaining the task execution result of the current solution task information, splicing the current solution task information (such as thinking information and action information) and the task execution result of the current solution task information (such as the feedback result described above) into a first prompt word template to obtain a first prompt word, and inputting the first prompt word into the LLM to obtain the next solution task information generated by the LLM.
[0028] 505. Based on the aggregated task information, the task execution results of the multiple task information solutions are aggregated to determine the reply information corresponding to the user question. Optionally, the first agent can directly splice the task execution results of the multiple task information solutions together and output them to the user as reply information. Alternatively, the process of determining the reply information corresponding to the user question includes: splicing the user question, the task execution results of the multiple task information solutions, and background knowledge information obtained from an external knowledge base whose similarity with the user question meets the set conditions into a third prompt word template to obtain a third prompt word, and inputting the third prompt word into the LLM to obtain the reply information output by the LLM. The execution process of the first agent can refer to the relevant descriptions in the other embodiments mentioned above and will not be elaborated here. Figure 6 is a flowchart of a question and answer processing method provided in an embodiment of the present disclosure. The method can be executed by the above-mentioned second agent. As shown in Figure 6, the method includes the following steps:
[0029] 601. Receive current task information corresponding to a user question from a first agent. Based on the user question, a first prompt word template corresponding to the first agent, and the LLM, the first agent obtains multiple task information sequentially generated by the LLM that requires invoking a second agent for execution, as well as summary task information required for execution by the first agent. The first prompt word template includes first toolset information. The current task information is the task information currently being sent to the second agent from among the multiple task information.
[0030] 602. Splice the current solution task information into the second prompt word template corresponding to the second agent, and obtain task execution information corresponding to the current solution task information generated by the LLM based on the second toolset information and the LLM included in the second prompt word template.
[0031] 603. Obtain the task execution result of the current task information from the application server corresponding to the user problem according to the task execution information.
[0032] 604. Send the task execution results for the current task information to the first agent, so that the first agent, based on the aggregated task information, aggregates the task execution results for multiple task information solutions to determine a response corresponding to the user's question. The execution process for the second agent can refer to the relevant descriptions in the aforementioned other embodiments and will not be elaborated here. The following describes in detail the question-and-answer processing devices according to one or more embodiments of the present disclosure. Those skilled in the art will appreciate that these devices can be constructed using commercially available hardware components and configured according to the steps taught in this solution. Figure 7 is a schematic structural diagram of a question-and-answer processing device provided by an embodiment of the present disclosure. As shown in Figure 7, the device is applied to a first agent. The first agent interacts with a second agent to complete question-and-answer processing. The first and second agents share the same large language model. The device includes: a first receiving module 11, a generating module 12, a first sending module 13, an acquiring module 14, and a summarizing module 15. The first receiving module 11 is configured to receive user questions. A generation module 12 is configured to incorporate the user question into a first prompt word template corresponding to the first agent, and to obtain, based on the first toolset information included in the first prompt word template and the large language model, multiple solution task information sequentially generated by the large language model and requiring execution by the second agent, as well as summary task information required to be executed by the first agent. A first sending module 13 is configured to sequentially send the multiple solution task information to the second agent. An acquisition module 14 is configured to obtain task execution results for the multiple solution task information sent by the second agent. The second agent, based on the second prompt word template corresponding to the second agent, the received current solution task information, and the large language model, obtains task execution information corresponding to the current solution task information generated by the large language model. Based on the task execution information, the second agent obtains the task execution result for the current solution task information from the application server corresponding to the user question, and sends the task execution result to the first agent. After obtaining the task execution result for the current solution task information, the first agent generates the next solution task information using the large language model. The second prompt word template includes the second toolset information. Summarizing module 15 is configured to summarize the task execution results of the multiple task-solving information based on the summarized task information to determine a response corresponding to the user's question. The apparatus shown in FIG7 can execute the steps provided by the first agent in the aforementioned embodiment. The detailed execution process and technical effects are described in the aforementioned embodiment and will not be repeated here.FIG8 is a schematic diagram of the structure of another question-and-answer processing device provided by an embodiment of the present disclosure. As shown in FIG8 , the device is applied to a second agent that interacts with a first agent to complete question-and-answer processing. The first and second agents share the same large language model. The device includes: a second receiving module 21, a processing module 22, and a second sending module 23. The second receiving module 21 is configured to receive current task information corresponding to a user question, sent by the first agent. The first agent, based on the user question, a first prompt word template corresponding to the first agent, and the large language model, obtains multiple task information sequentially generated by the large language model that requires invoking the second agent for execution, as well as summary task information required for execution by the first agent. The first prompt word template includes first toolset information. The current task information is the task information currently being sent to the second agent from the multiple task information. The processing module 22 is configured to incorporate the current task information into the second prompt word template corresponding to the second agent, obtain task execution information corresponding to the current task information generated by the large language model based on the second toolset information contained in the second prompt word template and the large language model, and obtain the task execution result of the current task information from the application server corresponding to the user question based on the task execution information. The second sending module 23 is configured to send the task execution result of the current task information to the first agent, so that the first agent, based on the summarized task information, summarizes the task execution results of the multiple task information to determine the response information corresponding to the user question. The device shown in FIG8 can execute the steps provided by the second agent in the aforementioned embodiment. The detailed execution process and technical effects are described in the aforementioned embodiment and will not be repeated here. In one possible design, the structure of the question-answering processing device shown in FIG7-8 can be implemented as an electronic device. As shown in FIG9, the electronic device may include: a processor 31, a memory 32, and a communication interface 33. Memory 32 stores executable code. When executed by processor 31, processor 31 can at least implement the question-and-answer processing method provided in the aforementioned embodiments. In an optional embodiment, the electronic device used to execute the question-and-answer processing method provided in the embodiments of the present disclosure can be any user terminal, such as a mobile phone, a laptop, or a PC, or an Extended Reality (XR) device. XR is a general term for various forms of virtual reality, augmented reality, and so on.In addition, embodiments of the present disclosure provide a non-transitory machine-readable storage medium storing executable code. When the executable code is executed by a processor of an electronic device, the processor is enabled to implement at least the question-and-answer processing method provided in the aforementioned embodiments. Embodiments of the present disclosure provide a computer program product, including a computer program. When executed by a processor, the computer program is enabled to implement at least the question-and-answer processing method provided in the aforementioned embodiments. The apparatus embodiments described above are merely illustrative, and the network elements described as separate components may or may not be physically separate. Some or all of these modules may be selected based on actual needs to achieve the objectives of the present embodiments. Persons of ordinary skill in the art can understand and implement these embodiments without inventive effort. Through the above description of the embodiments, persons of ordinary skill in the art can clearly understand that each embodiment can be implemented by adding a necessary general-purpose hardware platform, or alternatively, by a combination of hardware and software. Based on this understanding, the essence of the above-mentioned technical solutions, or the portion that contributes to the prior art, can be embodied in the form of a computer product. The present disclosure can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to magnetic disk storage, CD-ROM (Compact Disk Read-Only Memory), optical storage, etc.) containing computer-usable program code. Finally, it should be noted that the above embodiments are merely illustrative of the technical solutions of the present disclosure and are not intended to limit them. Although the present disclosure has been described in detail with reference to the above-mentioned embodiments, persons of ordinary skill in the art will understand that the technical solutions described in the above-mentioned embodiments may be modified or some of the technical features thereof may be replaced by equivalents. Such modifications or replacements do not deviate from the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the various embodiments of the present disclosure.
Claims
Claims 1. A question and answer processing system, wherein, Including: A first intelligent agent, a second intelligent agent, and an application server, where the first intelligent agent and the second intelligent agent share the same large language model; The first intelligent agent is configured to receive a user question, splice the user question into a first prompt template, and based on the first toolset information included in the first prompt template and the large language model, obtain multiple solution task information and summary task information sequentially generated by the large language model, and sequentially send the multiple solution task information to the second intelligent agent; The second intelligent agent is configured to splice the currently received solution task information in the multiple solution task information into a second prompt template, and based on the second toolset information included in the second prompt template and the large language model, obtain task execution information corresponding to the currently received solution task information, and obtain the task execution result of the currently received solution task information from the application server corresponding to the user question according to the task execution information, and send the task execution result to the first intelligent agent, so that the first intelligent agent generates the next solution task information through the large language model after obtaining the task execution result of the currently received solution task information; The first intelligent agent is configured to summarize the task execution results of the multiple solution task information based on the summary task information to determine the reply information corresponding to the user question.
2. The system according to claim 1, wherein, The first intelligent agent is further configured to obtain background knowledge information whose similarity with the user question meets a set condition from an external knowledge base, and / or obtain a historical conversation record corresponding to the user question; splice the background knowledge information and / or the historical conversation record into the first prompt template, and the historical conversation record is the questions and corresponding reply information before the user question in multiple rounds of conversations of the same user.
3. The system according to claim 1 or 2, wherein In the process of generating the next solution task information through the large language model after obtaining the task execution result of the currently received solution task information, the first intelligent agent is configured to: splice the currently received solution task information and the task execution result of the currently received solution task information into the first prompt template to obtain a first prompt, and input the first prompt into the large language model to obtain the next solution task information.
4. The system according to any one of claims 1 to 3, wherein The first prompt template is used to prompt the format of the task information that needs to be generated when the large language model gradually solves the user question; each task information in the multiple solution task information and the summary task information generated by the large language model includes first thinking information, first action information, and a first feedback identifier, where the first thinking information is used to describe the task that needs to be completed, the first action information is used to describe the tool to be selected from the first toolset information for executing the task and the input parameters of the tool, and the first feedback identifier is used to indicate the position where the execution result of the task is filled.
5. The system according to any one of claims 1 to 4, wherein The second intelligent agent is further configured to receive the background knowledge information and / or historical conversation record corresponding to the user question sent by the first intelligent agent, and splice the background knowledge information and / or the historical conversation record into the second prompt template. The historical conversation record is 23 the questions and corresponding reply information before the user question in multiple rounds of conversations of the same user. The background knowledge information is the background knowledge information in the external knowledge base whose similarity with the user question meets the set conditions.
6. The system according to any one of claims 1 to 5, wherein The second prompt template is used to prompt the large language model for the format of the task execution information required for the current solution task information. The task execution information includes second thinking information, second action information, and a second feedback identifier. The second thinking information is used to describe the solution task to be completed. The second action information is used to describe the tools to be selected from the second tool set information and the input parameters of the tools for executing the solution task. The second feedback identifier is used to indicate the filling position of the execution result of the solution task.
7. The system according to claim 6, wherein, The number of task execution information corresponding to the current solution task information is multiple. The second intelligent agent is further configured to: after obtaining the corresponding first task execution result from the application server based on the first task execution information, splice the first task execution information and the first task execution result into the second prompt template, and input the obtained second prompt into the large language model to obtain the second task execution information generated by the large language model.
8. The system according to any one of claims 1 to 7, wherein In the process that the first intelligent agent summarizes the task execution results of the multiple solution task information to determine the reply information corresponding to the user question, the first intelligent agent is configured to: splice the user question, the task execution results of the multiple solution task information, and the background knowledge information whose similarity with the user question meets the set conditions obtained from the external knowledge base into the third prompt template to obtain a third prompt, and input the third prompt into the large language model to obtain the reply information output by the large language model.
9. A question-and-answer processing method, wherein, Applied to the first intelligent agent, the first intelligent agent interacts with the second intelligent agent to complete the question-and-answer processing method. The first intelligent agent and the second intelligent agent share the same large language model. The method includes: receiving a user question; splicing the user question into a first prompt template to obtain multiple solution task information and summary task information sequentially generated by the large language model based on the first toolset information included in the first prompt template and the large language model; sequentially sending the multiple solution task information to the second intelligent agent, so that the second intelligent agent obtains task execution information corresponding to the current solution task information based on a second prompt template, the received current solution task information, and the large language model, obtains the task execution result of the current solution task information from the application server corresponding to the user question according to the task execution information, and sends the task execution result to the first intelligent agent, and the second prompt template includes second toolset information; obtaining the task execution results of the multiple solution task information sent by the second intelligent agent; based on the summary task information, summarizing the task execution results of the multiple solution task information to determine the reply information corresponding to the user question.
10. The method according to claim 9, wherein, The method further includes: obtaining background knowledge information whose similarity with the user question meets a set condition from an external knowledge base, and / or obtaining a historical conversation record corresponding to the user question; splicing the background knowledge information and / or the historical conversation record into the first prompt template, and the historical conversation record is the questions and corresponding reply information before the user question in multiple rounds of conversations of the same user.
11. The method according to claim 9 or 10, wherein The generation process of the multiple solution task information includes: after obtaining the task execution result of the current solution task information, splicing the current solution task information and the task execution result of the current solution task information into the first prompt template to obtain a first prompt. Inputting the first prompt into the large language model to obtain the next solution task information.
12. The method according to any one of claims 9 to 11, wherein The summarizing the task execution results of the multiple solution task information to determine the reply information corresponding to the user question includes: splicing the user question, the task execution results of the multiple solution task information, and the background knowledge information whose similarity with the user question meets a set condition obtained from an external knowledge base into a third prompt template to obtain a third prompt; inputting the third prompt into the large language model to obtain the reply information output by the large language model.
13. A question and answer processing method, wherein A second intelligent agent applied to interact with a first intelligent agent to complete the question-and-answer processing method. The first intelligent agent and the second intelligent agent share the same large language model. The method includes: receiving current solution task information corresponding to a user question sent by the first intelligent agent; wherein, the first intelligent agent, based on the user question, a first prompt word template, and the large language model, obtains multiple solution task information and summary task information sequentially generated by the large language model. The first prompt word template contains first tool set information, and the current solution task information is the solution task information currently sent to the second intelligent agent among the multiple solution task information; splicing the current solution task information into a second prompt word template to obtain task execution information corresponding to the current solution task information generated by the large language model based on the second tool set information contained in the second prompt word template and the large language model; obtaining a task execution result of the current solution task information from an application server corresponding to the user question according to the task execution information; sending the task execution result of the current solution task information to the first intelligent agent so that the first intelligent agent, based on the summary task information, summarizes the task execution results of the multiple solution task information to determine a reply information corresponding to the user question.
14. An electronic device, wherein, Including: A memory, a processor, and a communication interface; wherein, executable code is stored on the memory. When the executable code is executed by the processor, the processor executes the question-and-answer processing method according to any one of claims 9 to 12 or claim 13.
15. A non-transitory machine-readable storage medium, wherein, On the non-transitory machine-readable storage medium There is stored executable code, which, when executed by a processor of an electronic device, causes the processor to execute the question-and-answer processing method according to any one of claims 9 to 12 or claim 13.
16. A computer program product includes a computer program, and when the computer program is executed by a processor, it implements the question-and-answer processing method according to any one of claims 9 to 12 or claim 13. 26
Citation Information
Patent Citations
Medical service method, apparatus and device based on LLM intelligent agent architecture, and medium
CN117112759A
Traffic data analysis complex task intelligent disassembly and completion method based on large language model
CN117194624A
Content generation method and system based on large language model and user report generation method
CN117807979A
Cited By
Knowledge base question and answer method and system based on large language model agent
CN120492598A
Modularized man-machine collaborative decision-making method and device based on multiple agents
CN120705416A
A multi-agent based modular human-machine collaborative decision-making method and device
CN120705416B
Task execution method, device and system, electronic device and storage medium
CN120723411A
Large-model intelligent workflow interaction system and method based on graph process cues
CN120821757A