Question and answer processing method, device, storage medium and system
By introducing the first agent and the second agent in the intelligent traffic question and answer system, the role of the agent is dismantled in a hierarchical manner, the problem of insufficient multi-table joint query capabilities of the existing LLM is solved, the solution capabilities and scalability of the question and answer system are improved, the application scenarios are expanded, and the input text length limit of the LLM is reduced.
Patent Information
- Application Number
- CN202410016581.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-01-03
- Publication Date
- 2025-07-04
AI Technical Summary
The existing intelligent traffic question and answer system based on large language model (LLM) has weak multi-table joint query capabilities, and is limited by the ability of LLM, making it difficult to effectively deal with complex user problems.
The first agent and the second agent are introduced, and the role of the agent is dismantled in a hierarchical manner. The first agent is responsible for planning the overall solution process of user problems, and the second agent is responsible for task execution of specific functions. It uses the same large language model (LLM) to decompose and summarize the results to reduce the complexity of LLM's thinking.
It improves the problem-solving ability and scalability of the Q&A system, expands the application boundaries of LLM, covers a wider range of application scenarios, reduces the possibility of errors, and reduces the impact of the length limit of LLM input text.
Smart Images

Figure CN120256852A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of artificial intelligence technology, and in particular, to a question and answer processing method, device, storage medium, and system. Background Art
[0002] With the continuous development of large language model (LLM) technology and the significant improvement of various aspects of the model's performance, various applications based on LLM are also receiving extensive attention.
[0003] Taking the field of intelligent transportation as an example, in the field of intelligent transportation, users have the need to understand traffic conditions such as the congestion situation in the city on the current day. One existing way to interact with users based on LLM in the field of intelligent transportation is as follows: Utilizing the code generation ability of LLM, after receiving the question input by the user, LLM can automatically generate the corresponding SQL query instruction, and then by executing the SQL script, the relevant data can be queried in the traffic service platform. This scheme based on natural language to SQL query has relatively limited application scenarios, is only applicable to data query tasks based on SQL statements, and is restricted by the current capabilities of LLM. Currently, the multi-table joint query ability is relatively weak. Summary of the Invention
[0004] Embodiments of the present invention provide a question and answer processing method, device, storage medium, and system, which can accurately automatically reply to user questions through the cooperation of different intelligent agents.
[0005] In a first aspect, an embodiment of the present invention provides a question and answer processing system, including: a first intelligent agent, a second intelligent agent, and an application server, where the first intelligent agent and the second intelligent agent share the same large language model;
[0006] The first intelligent agent is configured to receive a user question, splice the user question into a first prompt word template corresponding to the first intelligent agent, and based on the first tool set information included in the first prompt word template and the large language model, obtain multiple solution task information that needs to be called by the second intelligent agent for execution and summary task information that needs to be executed by the first intelligent agent generated by the large language model in sequence, and sequentially send the multiple solution task information to the second intelligent agent;
[0007] The second intelligent agent is used to splice the received current solution task information into the second prompt word template corresponding to the second intelligent agent, and based on the second tool set information included in the second prompt word template and the large language model, obtain the task execution information corresponding to the current solution task information generated by the large language model. According to the task execution information, obtain the task execution result of the current solution task information from the application server corresponding to the user question, and send the task execution result to the first intelligent agent, so that the first intelligent agent can generate the next solution task information through the large language model after obtaining the task execution result of the current solution task information;
[0008] The first intelligent agent is used to summarize the task execution results of the multiple solution task information based on the summary task information to determine the reply information corresponding to the user question.
[0009] In a second aspect, an embodiment of the present invention provides a question and answer processing method, which is applied to a first intelligent agent. The first intelligent agent interacts with a second intelligent agent to complete the question and answer processing method. The first intelligent agent and the second intelligent agent share the same large language model. The method includes:
[0010] Receive a user question;
[0011] Splice the user question into the first prompt word template corresponding to the first intelligent agent, so as to obtain, based on the first tool set information included in the first prompt word template and the large language model, multiple solution task information that needs to be called by the second intelligent agent to execute and summary task information that needs to be executed by the first intelligent agent generated by the large language model in sequence;
[0012] Send the multiple solution task information to the second intelligent agent in sequence, so that the second intelligent agent can obtain, based on the second prompt word template corresponding to the second intelligent agent, the received current solution task information and the large language model, the task execution information corresponding to the current solution task information generated by the large language model. According to the task execution information, obtain the task execution result of the current solution task information from the application server corresponding to the user question, and send the task execution result to the first intelligent agent, so that the first intelligent agent can generate the next solution task information through the large language model after obtaining the task execution result of the current solution task information. The second prompt word template includes second tool set information;
[0013] Obtain the task execution results of the multiple solution task information sent by the second intelligent agent;
[0014] Based on the summarized task information, summarize the task execution results of the multiple solution tasks to determine the response information corresponding to the user's question.
[0015] In a third aspect, an embodiment of the present invention provides a question and answer processing device, which is applied to a first intelligent agent. The first intelligent agent interacts with a second intelligent agent to complete question and answer processing. The first intelligent agent and the second intelligent agent share the same large language model. The device includes:
[0016] A receiving module, configured to receive a user's question;
[0017] A generating module, configured to splice the user's question into a first prompt word template corresponding to the first intelligent agent, and based on the first tool set information included in the first prompt word template and the large language model, obtain multiple solution task information that needs to be called by the second intelligent agent to execute and summarized task information that needs to be executed by the first intelligent agent generated by the large language model in sequence;
[0018] A sending module, configured to sequentially send the multiple solution task information to the second intelligent agent, so that the second intelligent agent, based on a second prompt word template corresponding to the second intelligent agent, the received current solution task information, and the large language model, obtains task execution information corresponding to the current solution task information generated by the large language model, obtains the task execution result of the current solution task information from the application server corresponding to the user's question according to the task execution information, and sends the task execution result to the first intelligent agent, so that the first intelligent agent generates the next solution task information through the large language model after obtaining the task execution result of the current solution task information. The second prompt word template includes second tool set information;
[0019] An obtaining module, configured to obtain the task execution results of the multiple solution task information sent by the second intelligent agent;
[0020] A summarizing module, configured to summarize the task execution results of the multiple solution task information based on the summarized task information to determine the response information corresponding to the user's question.
[0021] In a fourth aspect, an embodiment of the present invention provides a question and answer processing method, which is applied to a second intelligent agent that interacts with a first intelligent agent to complete the question and answer processing method. The first intelligent agent and the second intelligent agent share the same large language model. The method includes:
[0022] Receive the current solution task information corresponding to the user question sent by the first agent; wherein, the first agent, based on the user question, the first prompt word template corresponding to the first agent, and the large language model, obtains multiple solution task information that needs to be called by the second agent for execution and summary task information that needs to be executed by the first agent generated by the large language model in sequence, the first prompt word template contains first tool set information, and the current solution task information is the solution task information currently sent to the second agent among the multiple solution task information;
[0023] Splice the current solution task information into the second prompt word template corresponding to the second agent, so as to obtain task execution information corresponding to the current solution task information generated by the large language model based on the second tool set information contained in the second prompt word template and the large language model;
[0024] Obtain the task execution result of the current solution task information from the application server corresponding to the user question according to the task execution information;
[0025] Send the task execution result of the current solution task information to the first agent, so that the first agent, based on the summary task information, summarizes the task execution results of the multiple solution task information to determine the reply information corresponding to the user question.
[0026] In a fifth aspect, an embodiment of the present invention provides a question and answer processing device, which is applied to a second agent that interacts with a first agent to complete question and answer processing. The first agent and the second agent share the same large language model. The device includes:
[0027] A receiving module, configured to receive the current solution task information corresponding to the user question sent by the first agent; wherein, the first agent, based on the user question, the first prompt word template corresponding to the first agent, and the large language model, obtains multiple solution task information that needs to be called by the second agent for execution and summary task information that needs to be executed by the first agent generated by the large language model in sequence, the first prompt word template contains first tool set information, and the current solution task information is the solution task information currently sent to the second agent among the multiple solution task information;
[0028] A processing module, configured to splice the current solution task information into a second prompt template corresponding to the second intelligent agent, so as to obtain task execution information generated by the large language model corresponding to the current solution task information based on the second tool set information and the large language model included in the second prompt template, and obtain a task execution result of the current solution task information from the application server corresponding to the user question according to the task execution information;
[0029] A sending module, configured to send the task execution result of the current solution task information to the first intelligent agent, so that the first intelligent agent summarizes the task execution results of the multiple solution task information based on the summary task information to determine a reply information corresponding to the user question.
[0030] In a sixth aspect, an embodiment of the present invention provides an electronic device, including: a memory, a processor, and a communication interface; wherein, an executable code is stored on the memory, and when the executable code is executed by the processor, the processor can at least implement the question and answer processing method as described in the second aspect or the fourth aspect.
[0031] In a seventh aspect, an embodiment of the present invention provides a non-transitory machine-readable storage medium, on which an executable code is stored, and when the executable code is executed by a processor of an electronic device, the processor can at least implement the question and answer processing method as described in the second aspect or the fourth aspect.
[0032] An embodiment of the present invention provides a question and answer processing system, which includes a first intelligent agent as a master intelligent agent, a second intelligent agent as a controlled intelligent agent, and an application server (such as an intelligent transportation server). The first intelligent agent and the second intelligent agent share the same LLM. Among them, the function of the first intelligent agent is mainly to plan the step-by-step solution process of the user question, that is, to gradually decompose the solution tasks of the user question, and each decomposed solution task will be sent to the second intelligent agent for execution, that is, to obtain the task execution result of the corresponding solution task. Finally, the first intelligent agent summarizes the task execution results of each solution task to obtain the reply information to be output to the user. Among them, every time the first intelligent agent generates a solution task information and sends it to the second intelligent agent for execution to obtain the corresponding task execution result, it will generate the next solution task information, and so on, until the solution of the user question is completed and the reply information is summarized.
[0033] Specifically, after receiving the user's question, the first agent can splice the user's question into the first prompt template corresponding to the first agent to generate a corresponding prompt for input into the LLM. The first prompt template contains the first tool set information that the first agent can use in the process of processing the user's question, such as including tools for calling the second agent to complete a certain solving task and tools for summarizing the execution results of multiple solving tasks. Based on the prompts sequentially input by the first agent, the LLM sequentially generates multiple solving task information that needs to be executed by the second agent and summary task information that needs to be executed by the first agent corresponding to the user's question, and the first agent sequentially sends the multiple solving task information to the second agent.
[0034] The second agent splices the received current solving task information into the second prompt template corresponding to the second agent, generates corresponding prompts for input into the LLM based on the second tool set information included in the second prompt template, obtains the task execution result of the current solving task information from the application server corresponding to the user's question based on the task execution information output by the LLM, and sends the task execution result to the first agent. In this way, after obtaining the task execution result of the current solving task information, the first agent can generate the next solving task information of the current solving task information in the above multiple solving task information through the LLM based on the task execution result and the first prompt template.
[0035] After the first agent decomposes the solving process of the user's question into multiple solving task information and controls the second agent to complete the execution of each solving task information to obtain the corresponding task execution result, the first agent summarizes the task execution results of the multiple solving task information based on the summary task information generated by the LLM to determine the reply information corresponding to the user's question.
[0036] The above-mentioned question-and-answer processing solution provided by the embodiments of the present invention hierarchically disassembles agents according to their roles by introducing a first agent for controlling the overall solution process of user questions and a second agent for completing a specific function, such as a data collection function. Different agents handle different types of tasks: the first agent is responsible for planning the overall solution process of user questions, and based on the user questions and the execution results obtained after each step of the solution task is executed, it determines whether to continue to call the second agent to complete the next solution task, or to summarize and reply to the existing execution results. The actual execution logic of the solution task is the responsibility of the second agent. This reduces the complexity of each step of thinking of the LLM, reduces the possibility of errors, and improves the problem-solving ability and scalability of the agents. By introducing the second agent, it is possible to connect to tools of different functional types (such as APIs) in the application server. Compared with the solution based on natural language to SQL query, the ability boundary of the LLM is extended, and a wider range of application scenarios can be covered. In addition, since the LLM has a limit on the length of the input text, if the input text is too long, its decision-making ability will decline. And the agent needs to send information such as the description of available tools and input parameters to the LLM in text form. For example, the above first and second prompt templates both contain corresponding tool set information. When there are many tools to be used, the text input to the LLM will be very long, and may even exceed the upper limit of the allowed input length. In the embodiments of the present invention, by classifying different types of agents, it is equivalent to classifying the tools. The number of tools used by each agent is relatively small, so the text of the corresponding tool set information independently input to the LLM by each agent will be much shorter, thereby reducing the length of the text input to the LLM and ensuring the reliable operation of the LLM. BRIEF DESCRIPTION OF THE DRAWINGS
[0037] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0038] Figure 1 Schematic diagram of the hardware execution environment of a question-and-answer processing method provided by an embodiment of the present invention;
[0039] Figure 2 Schematic diagram of the cloud computing environment of a question-and-answer processing method provided by an embodiment of the present invention;
[0040] Figure 3 Schematic diagram of the application of a question-and-answer processing method provided by an embodiment of the present invention;
[0041] Figure 4Schematic diagram of the composition of a question and answer processing system provided by an embodiment of the present invention;
[0042] Figure 5 Flowchart of a question and answer processing method provided by an embodiment of the present invention;
[0043] Figure 6 Flowchart of a question and answer processing method provided by an embodiment of the present invention;
[0044] Figure 7 Schematic diagram of the structure of a question and answer processing device provided by an embodiment of the present invention;
[0045] Figure 8 Schematic diagram of the structure of another question and answer processing device provided by an embodiment of the present invention;
[0046] Figure 9 Schematic diagram of the structure of an electronic device provided by an embodiment of the present invention. Detailed implementation manners
[0047] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the scope of protection of the present invention.
[0048] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in the embodiments of the present invention are all information and data that have been authorized by the user or fully authorized by all parties. Moreover, the collection, use, and processing of relevant data need to comply with the relevant laws, regulations, and standards of relevant countries and regions, and corresponding operation entrances are provided for users to select authorization or rejection.
[0049] Some embodiments of the present invention will be described in detail below with reference to the accompanying drawings. Without conflict between the embodiments, the following embodiments and the features in the embodiments can be combined with each other. In addition, the step sequence in the following method embodiments is only an example and is not strictly limited.
[0050] First, the terms or concepts involved in the embodiments of the present invention will be explained:
[0051] Large language model: Large Language Model, abbreviated as LLM.
[0052] Prompt: The prompt of a large language model is the input for interacting with the large language model, usually containing information such as instructions, context, input / output format, etc.
[0053] Token: The smallest unit for a large language model to process text.
[0054] Agent: Artificial Intelligent Agent, abbreviated as AI Agent. An agent is a proxy system that controls an LLM to solve problems. It is an intelligent entity that can perceive the environment, make decisions, and execute actions. An AI agent can be a physical entity (such as a robot) or a virtual entity (such as a computer program), which realizes specific tasks or goals by perceiving information in the environment, making action decisions, and executing actions.
[0055] Application Programming Interface, abbreviated as API. In this article, it refers to the channel / interface for the agent to interact with service platforms such as transportation.
[0056] With the continuous development of LLM technology and the significant improvement of various aspects of the model's performance, various applications based on LLM are receiving wide attention. Among them, the LLM-driven agent (AI-Agent) is an important research direction. It enables the LLM to perform task decomposition, external tool invocation, and result analysis through natural language dialogue, thereby achieving more general problem-solving. Currently, single-agent-based solutions are used in many application fields: tasks are decomposed, tools are invoked, and results are summarized by a single agent. Its disadvantage is that in scenarios where multiple complex tasks of different categories need to be invoked, the agent's attention is focused on processing the results of tool invocation, and it is easy to ignore the overall problem-solving situation, which may lead to problems such as deviating from the initial goal or ignoring the previous solution results, that is, it cannot complete multi-step continuous problem-solving well.
[0057] In view of this, an embodiment of the present invention provides a solution for jointly solving user problems based on multiple agents. Specifically, by introducing a first agent (with the role of a master agent) for controlling the overall solution process of user problems and a second agent (with the role of a functional agent) for completing a specific function, such as data collection, the agents are hierarchically disassembled according to their roles. Different agents handle different types of tasks: the first agent is responsible for planning the overall solution process of user problems, and based on the user problems and the execution results obtained after each step of the solution task is executed, it determines whether to continue to call the second agent to complete the next solution task, or to summarize the existing execution results and reply. The actual execution logic of the solution task is the responsibility of the second agent. This reduces the complexity of each step of thinking of the LLM, reduces the possibility of errors, and improves the problem-solving ability and scalability of the agent.
[0058] The Q&A processing solution provided by the embodiment of the present invention can be applied to various application scenarios, including but not limited to intelligent transportation scenarios, database query scenarios, e-commerce scenarios, etc. In different application scenarios, users will generate user problems that they want the Q&A processing system to automatically reply to. The Q&A processing system docks with the corresponding background service system (such as an intelligent transportation server), and based on multiple agents included in the Q&A processing system, gradually solves the user problems to gradually obtain the required data from the background service system, and finally summarizes the reply information of the user problems. Among them, the automatic reply to the user problem is actually a Q&A task, and the gradual solution to the user problem is the gradual solution to this Q&A task. The gradual solution is mainly reflected in that the master agent disassembles the Q&A task into multiple subtasks (solution tasks in the following text), the functional agent executes the subtasks to obtain the corresponding execution results, and the master agent summarizes the execution results of each subtask to determine the reply information of the user problem.
[0059] The Q&A processing solution provided by the embodiment of the present invention will be introduced and described below.
[0060] Figure 1 It is a schematic diagram of the hardware execution environment of a Q&A processing method provided by an embodiment of the present invention. As Figure 1 shown, the hardware execution environment of this Q&A processing method can be composed of a user device 101, a Q&A server 102, and an application server 103. The client device 101 is communicatively connected to the Q&A server 102, and the Q&A server 102 is communicatively connected to the application server 103.
[0061] The user device 101 can be a certain type of terminal device for independent use, such as a smart phone, a tablet computer, a PC, etc., or can be two or more terminal devices for combined use, such as an extended reality device 101a and a smart phone 101b. The extended reality device 101a can be a virtual reality device, an augmented reality device, etc.
[0062] The Q&A server 102 can be a server of an application provider or a cloud server of a cloud service provider. The application server 103 can be a server of an application provider or a cloud server of a cloud service provider. From the perspective of physical location, the Q&A server 102 and the application server 103 can be located in the same or different physical hosts. The Q&A server 102 refers to the server that executes the Q&A processing solution provided by the embodiments of the present invention, and the application server 103 refers to the server that provides a certain application service with which the Q&A server 102 needs to interact to obtain data related to the user's question, such as an intelligent transportation server, an e-commerce server, a video server, etc.
[0063] Optionally, when the user device 101 consists of an extended reality (XR) device 101a and a smart phone 101b, the execution process of the above Q&A processing method can be as follows: The user inputs a user question to the smart phone 101b in natural language form. The smart phone 101b sends the user question to the Q&A server 102. The Q&A server 102 executes the Q&A processing method provided by the embodiments of the present invention, and finally obtains the reply information corresponding to the user question through interaction with the application server 103, sends the reply information to the smart phone 101b, and the smart phone 101b sends the reply information to the extended reality device 101a for display.
[0064] In practical applications, the above Q&A server 102 can be an independent physical server or a physical server cluster maintained by an application provider, or can be a cloud server maintained by a cloud service provider - called a computing node. In a Figure 2 cloud computing environment as shown, it can include a number of ([ Figure 2 such as those indicated by 201-1, 201-2,... in the figure) computing nodes (cloud servers) deployed distributively, and each computing node has processing resources such as computing and storage. In the cloud computing environment, a number of computing nodes can be organized to provide a certain service. Of course, a single computing node can also provide one or more services, such as Figure 2Services A, B, C, and D as shown. The way to provide these services in the cloud computing environment can be to externally provide a service interface 202, and the client device calls this service interface 202 to use the corresponding service. The service interface 202 includes forms such as a Software Development Kit (SDK) and an Application Programming Interface (API).
[0065] The above services are deployed according to various virtualization technologies supported by the cloud computing environment, such as virtualization technologies based on virtual machines and containers. Taking the virtualization technology based on containers as an example, several containers corresponding to one service can be assembled into a container group (pod). For example Figure 2 Service B as shown can be configured with one or more pods, and each pod can include a proxy and one or more containers. One or more containers in the pod are used to process requests related to one or more corresponding functions of the service, and the proxy in the pod is used to control network functions related to the service, such as routing, load balancing, etc.
[0066] During the operation process, when executing a request from a user device, it may be necessary to call one or more services in the cloud computing environment, and when executing one or more functions of a service, it may be necessary to call one or more functions of another service. As Figure 2 shown, after Service A receives a request sent by the user device, it can call Service B, and Service B can request Service D to execute one or more functions.
[0067] Under the above cloud computing environment, the embodiment of the present invention provides an application schematic diagram of a question and answer processing method as shown Figure 3 in.
[0068] In Figure 3 the cloud computing environment, there are the following three services provided: a first agent, a second agent, and an LLM. These three services can be located on the same computing node or different computing nodes. The first agent can call the second agent and the LLM, and the second agent can call the LLM. In addition, the second agent can also interact with the application server 103, and the first agent can interact with the user device 101. After the first agent receives a user question sent by the user through the user device 101, the first agent, the second agent, and the LLM finally determine the reply information of the user question based on the question and answer processing solution provided by the embodiment of the present invention, and the first agent feeds back this reply information to the user device 101.
[0069] Figure 4 is a schematic diagram of the working principle of a question and answer processing system provided by the embodiment of the present invention, as shown Figure 4As shown, the question-and-answer processing system includes: a first agent, a second agent, an LLM, and an application server. The first agent and the second agent share the same LLM.
[0070] The first agent, as the main control agent, is responsible for planning and controlling the overall solution process of the user question-solving task. As Figure 4 shown, in an optional embodiment, the first agent may include a background knowledge acquisition module. At this time, the first agent first receives a task from the user input, that is, receives the user question, and then obtains background knowledge related to the user question from the external knowledge base through the background knowledge acquisition module. On this basis, through the "Think / Act / Feedback" mechanism, it interacts with the LLM to gradually solve the user question-solving task. After completing the task solution, the first agent will call the summary module to summarize the execution results generated by all intermediate steps of the current task to obtain the final reply information and output it.
[0071] As introduced in the existing related technologies, the main functional components of the agent include memory, thinking (i.e., planning), acting, feedback (reflection), using tools, etc., which correspond to the "Think / Act / Feedback" mechanism of the first and second agents in the embodiments of the present invention.
[0072] Essentially, the "Think / Act / Feedback" mechanism is a prompt template for solving tasks, and it completes the problem-solving by repeating and iterating "Think / Execute / Feedback". Specifically, for the first agent, the main work of this mechanism at each step is as follows:
[0073] Think: The LLM combines the user question and the completed solution steps, thinks, and gives a summary;
[0074] Act: The LLM gives the tools to be called for the next action and the tool input parameters;
[0075] Feedback: The backend (referring to the first agent) calls the corresponding tool or other agents according to the "Act" output by the LLM, and obtains the execution result after completion.
[0076] The execution process of the "Think / Act / Feedback" mechanism of the first agent can be simply described as:
[0077] The first agent receives the user's question, splices the user's question (i.e., inserts it into the corresponding position in the template) into the first prompt template corresponding to the first agent, and based on the first toolset information and the LLM included in the first prompt template, obtains multiple solution task information that needs to be executed by the second agent and summary task information that needs to be executed by the first agent successively generated by the LLM, and sends the multiple solution task information to the second agent in sequence. The first agent summarizes (i.e., concludes) the task execution results of the multiple solution task information based on the summary task information to determine the response information corresponding to the user's question. Among them, the task execution results of the multiple solution task information are fed back to the first agent by the second agent.
[0078] Combined with Figure 4 to simply illustrate the above execution process. As Figure 4 shown, the execution process of the "Think / Act / Feedback" mechanism of the first agent includes: First, the first agent obtains relevant information in multiple dimensions (such as Figure 4 the user's question, background knowledge, historical conversation records obtained from the memory module, and current thinking records shown), and splices this information into the first prompt template to obtain the first prompt generated by the current round of execution of the "Think / Act / Feedback" mechanism, which includes three parts: thinking, acting, and feedback.
[0079] The first prompt template includes the first toolset information that the first agent can use, that is, the LLM can know what tools can be used when generating the "action" content from the first prompt. Figure 4 shown in
[0080] three tools that the first agent can use: the "summary" tool represented by the summary module, the "askuser" tool represented by the interaction module, and the "data" tool for calling the second agent to collect data. Here, it is assumed that the function provided by the second agent is the data collection function.
[0081] It should be noted that the "feedback" content generated by the LLM actually only includes the feedback identifier and does not include the real feedback result. The feedback identifier is, for example, ——Feedback:.
[0082] As Figure 4As shown in [figure], during the process of the first agent generating the first prompt, the historical conversation records and the current thinking records obtained from the memory module can be used.
[0083] Among them, the historical conversation records are the questions and corresponding reply information before the current user question in multiple rounds of conversations with the same user. Among them, the current thinking record refers to that during the process of the first agent gradually solving the user question, after each execution of the "thinking / action / feedback" mechanism, the thinking, action, and feedback content finally obtained in this round forms a data group, which is used as the current thinking record for the next round of execution of this mechanism. At this time, the feedback data content obtained by the first agent has been filled in after the feedback identifier of the previous round.
[0084] The functions of the first agent have been briefly introduced above. Next, the functions of the second agent will be introduced.
[0085] The second agent is a controlled agent relative to the first agent, which is the main control agent, and is a functional agent that provides a specific function. For different application scenario requirements, one or more different functional agents can be configured. Figure 4 Shown in [figure] is the second agent that provides the data collection function. In some application scenarios, for example, it can also include agents that provide data analysis and prediction functions, agents that provide data visualization display functions, and so on. The toolset information used by different functional agents is different.
[0086] Assume that the second agent is an agent that provides the data collection function (which can be called a data agent). Then the second agent is responsible for receiving the solution task information related to data collection sent by the first agent, and obtaining the corresponding data from an external application server (such as a certain application server including a database) (the original data obtained from the application server as shown in [figure]) and feeding it back to the first agent. The process of obtaining data is carried out through the "thinking / execution / feedback" mechanism. Figure 4 The "thinking / execution / feedback" mechanism of the second agent is similar to that of the first agent. Generally speaking, it is also based on the second prompt template containing the second toolset information, generates information containing "thinking, action, and feedback" content through the LLM, and based on the indication of the action content, calls the corresponding tool to obtain the corresponding data from the application server. The differences between the first prompt template and the second prompt template are mainly reflected in: First, the toolset is different; Second, the information that needs to be spliced into the prompt template is different.
[0087]
[0088] Among them, the first tool set of the first intelligent agent may include tools such as summary, data, and askuser exemplified above for controlling the macro-solving process of user questions. The second tool set of the second intelligent agent may include various API information for docking with the application server, such as different APIs for querying different data contents. Each API information may include information such as API function description, input parameter description and format, and return result description.
[0089] Among them, as described above, the information spliced into the first prompt template may include the user question, historical conversation records, background knowledge, and the current thinking record corresponding to the first intelligent agent. And the information spliced into the second prompt template may include the solution task information (including thinking, action, and feedback content) currently sent by the first intelligent agent to the second intelligent agent and the current thinking record corresponding to the second intelligent agent.
[0090] Among them, the current thinking record corresponding to the second intelligent agent means that: assuming that the second intelligent agent is also required to gradually solve a certain solution task information sent by the first intelligent agent during the execution process (similar to the process of the first intelligent agent gradually solving the user question), then during the solving process, after each round of "thinking / action / feedback" mechanism is executed, the thinking, action, and feedback content finally obtained in this round form a data group, which is used as the current thinking record for the next round of execution of this mechanism. At this time, the feedback data content obtained by the first intelligent agent has been filled in after the feedback identifier of the previous round.
[0091] Based on this, it can be known that both the first intelligent agent and the second intelligent agent may have their own "current thinking records" and store them independently.
[0092] Generally speaking, the working process of the second intelligent agent includes: splicing the current solution task information received from the first intelligent agent into the second prompt template corresponding to the second intelligent agent, obtaining the task execution information generated by the LLM corresponding to the current solution task information based on the second tool set information and the LLM included in the second prompt template, obtaining the task execution result of the current solution task information from the application server corresponding to the user question according to the task execution information, and sending the task execution result to the first intelligent agent, so that the first intelligent agent can generate the next solution task information through the LLM after obtaining the task execution result of the current solution task information. Among them, the task execution information is used to describe what operations the application triggers on the application server to obtain the corresponding task execution result.
[0093] Among them, the task execution information includes both thinking and action content and a feedback identifier. At this time, the feedback content is empty. After the second intelligent agent executes the corresponding action according to the action content and obtains the task execution result to fill it in after the feedback identifier, the feedback content is formed.
[0094] Figure 4 The memory module shown in [description] contains three types of sub-modules: historical conversation records, current thinking records (the current thinking records corresponding to the first intelligent agent and the current thinking records corresponding to the second intelligent agent), and raw data (the data obtained by the second intelligent agent from the application server). Among them, the historical conversation record module is used to store the historical Q&A data between the user and the first intelligent agent, including each "user question" and "returned reply information", which are generally the Q&A data of the previous rounds in a multi-round conversation scenario. The current thinking record module stores the intermediate execution results generated after the operation of the "thinking / executing / feedback" mechanism in the current round of conversation of the first intelligent agent and the second intelligent agent. The raw data module stores the raw results obtained by the second intelligent agent from the application server during the query process, which are used for front-end rendering and display. That is, in addition to the reply information, the visual display results of the raw data can also be included in the final output to the user.
[0095] Figure 4 The summary module in [description] is responsible for summarizing the solution process of the intelligent agent and giving the final reply information. Its inputs include the user question, background knowledge, and each intermediate execution result generated by the first intelligent agent under the "thinking / executing / feedback" mechanism in each round. After obtaining and combining these information, the summary module asks the LLM to give the final reply information. It can be understood that after the reply information corresponding to the current user question is determined, this Q&A pair can be stored in the historical conversation record in the memory module for use as the historical conversation for the user's next user question.
[0096] In summary, in the above Q&A processing solution provided by the embodiments of the present invention, by introducing a first intelligent agent for controlling the overall solution process of the user's question and a second intelligent agent for completing a specific function, such as a data collection function, the intelligent agents are hierarchically disassembled according to their roles, and different intelligent agents handle different types of tasks: the first intelligent agent is responsible for planning the overall solution process of the user's question, and based on the user's question and the execution results obtained after each step of the solution task is executed, it determines whether to continue to call the second intelligent agent to complete the next solution task, or to summarize and reply to the existing execution results. The actual execution logic of the solution task is the responsibility of the second intelligent agent. This reduces the complexity of each step of thinking of the LLM, reduces the possibility of errors, and improves the problem-solving ability and scalability of the intelligent agent. By introducing the second intelligent agent, it is possible to connect to tools of different functional types (such as APIs) in the application server. Compared with the solution based on natural language to SQL query, the ability boundary of the LLM is extended, and a wider range of application scenarios can be covered. In addition, since the LLM has a limit on the length of the input text, if the input text is too long, its decision-making ability will decline. And the intelligent agent needs to send information such as the description of the available tools and input parameters to the LLM in text form. For example, the above first and second prompt templates both contain corresponding tool set information. When there are many tools to be used, the text input to the LLM will be very long, and may even exceed the upper limit of the allowed input length. In the embodiments of the present invention, by classifying different types of intelligent agents, it is equivalent to classifying the tools. The number of tools used by each intelligent agent is relatively small, so the text of the corresponding tool set information independently input to the LLM by each intelligent agent will be much shorter, thereby reducing the length of the text input to the LLM and ensuring the reliable operation of the LLM.
[0097] For ease of understanding, the following example illustrates the multi-round execution process of the "thinking / execution / feedback" mechanism of the first intelligent agent.
[0098] Suppose the user's question is: "Query the congestion indices of Yuhang District and Xihu District today", and the step-by-step solution process is as follows:
[0099] 1. Thinking (generated by the LLM): The user needs to query the congestion indices of Yuhang District and Xihu District today. First, query the congestion index of Yuhang District;
[0100] 2. Action (generated by the LLM): [data] Query the congestion index of Yuhang District today;
[0101] 3. Feedback (the first intelligent agent parses the previous "Action" text, and based on the tool call information of [data], calls the second intelligent agent to complete the query and obtains): The congestion index of Yuhang District today is 1.3;
[0102] 4. Thought (generated by LLM, note that the prompt words input to the LLM at this time will include the [Thought, Action, Feedback] data group generated in the previous round): The congestion index of Yuhang District has been obtained. Now query the congestion index of Xihu District;
[0103] 5. Action (generated by LLM): [data] Query the congestion index of Xihu District today;
[0104] 6. Feedback (after the first agent parses the "Action" text in the previous step and calls the second agent to complete the query): The congestion index of Xihu District today is 1.41;
[0105] 7. Thought (generated by LLM, note that the prompt words input to the LLM at this time will include the [Thought, Action, Feedback] data group generated in the previous round): The congestion indices of Yuhang District and Xihu District have been obtained. Now the result can be returned;
[0106] 8. Action (generated by LLM): [summary] The congestion index of Yuhang District today is 1.3, and the congestion index of Xihu District today is 1.41;
[0107] 9. Feedback: The first agent recognizes the "summary" field, calls the summary module, obtains the reply information and returns it to the user.
[0108] In the above example, the format of the "Action" part of the information is:
[0109] Action: [Tool identifier] Action input. Among them, the tool identifier can also be called the tool name, indicating what tool this action needs to call, and the "Action input" mainly gives the input parameters of the tool. For example, in the above example, the natural language description "Query the congestion index of Xihu District today" is used as the input parameter of the tool data, indicating that the second agent needs to be called to execute the solution task of "Query the congestion index of Xihu District today". In practical applications, the solution task information input to the second agent can include, but is not limited to: the Thought + Action parts of the content.
[0110] Reference Figure 4 , the overall execution process of the question and answer processing method provided by the embodiments of the present invention is introduced below, which may include the following steps:
[0111] Step 1. User input and background knowledge extraction.
[0112] The user inputs a string as the user question in the form of natural language text. After receiving the string of the user question, the first intelligent agent first parses the string. If the string is an instruction to clear the history, such as the " / Clear" instruction, the historical conversation record of this user in the memory module is cleared, and at the same time, the final result "The history has been successfully cleared" is directly returned, and this round of Q&A ends. If the string is not such a clearing instruction, it enters the normal problem-solving process.
[0113] The user question may be a relatively clear query question, such as: What is the congestion index in Hangzhou today? Query the traffic volume from Xihu District to Yuhang District. It may also be an open-ended analysis question, such as: How is the traffic condition in Hangzhou today?
[0114] In the embodiments of the present invention, the historical conversation record and background knowledge are both for enabling the LLM to more accurately understand the current question intention of the user and finally give a reply information expression that conforms to common language habits.
[0115] In this step, the string of the user question can be input into the background knowledge acquisition module of the first intelligent agent to search for background knowledge related to this user question from the external knowledge base through the background knowledge acquisition module. The process flow is as follows:
[0116] Step 1.1: User intention recognition. This step is an optional step for rewriting the user question input by the user into a more standardized format.
[0117] Represent the user question input by the user as s u . The first intelligent agent sends s u to the LLM, requesting it to parse the domain and keywords of the user question and refine the question again. For example, if the user question s u = "Is it congested in Hangzhou today?", the LLM returns that the identified domain is "traffic", the keywords are "Hangzhou" and "congestion situation", and the question is refined to obtain s' u = "Query the congestion situation in Hangzhou today". Thus, the LLM rewrites the relatively colloquial question expression of the user into a relatively standardized expression form.
[0118] Step 1.2: Extraction of relevant tools.
[0119] According to the domain identified in step 1.1, determine the functional intelligent agent corresponding to this domain (i.e., the second intelligent agent) and the tool set information corresponding to the functional intelligent agent. For example, when the identified domain is "traffic", the determined second intelligent agent is a functional intelligent agent applicable to the traffic domain, and the determined tool set is related tools for docking the intelligent transportation service system, such as various APIs.
[0120] It is understandable that if only the functional agents applicable to a specific field and the corresponding tool sets are provided, then step 1.2 does not need to be executed.
[0121] Step 1.3, Extracting external knowledge base information.
[0122] Suppose the external knowledge base K contains several (such as m) background knowledge paragraphs: K = [(k1, v1), (k2, v2), …, (k m , v m )], where k i , v i , i ∈ [1, m] respectively represent the title and content of the knowledge paragraph organized in natural language form. For example: (k1 = "Analyzing the urban congestion situation", v1 = "Analyzing the congestion situation of a city starts from the following three steps: 1. Query the congestion index of the whole city; 2. Query the congestion index of each area in the city; 3. Query the several most congested roads in the city.").
[0123] The first agent may include a vector representation model θ s . After each paragraph title in K is processed by the vector representation model θ, the obtained high-dimensional vector group is Among them, represents the vector representation of the knowledge paragraph title k i . And assume that the set similarity threshold is β t .
[0124] The first agent inputs the above s′ u into the vector representation model θ s and obtains the corresponding vector representation: After that, calculate and the cosine similarity of each element in X k . If the cosine similarity with is greater than the threshold β t , then determine the corresponding knowledge paragraph content v j as the background knowledge related to the user's question obtained by query. Or, if the cosine similarity with is greater than the threshold βt, and the cosine similarity with is greater than the cosine similarity with the titles of other knowledge paragraphs, then determine the corresponding knowledge paragraph content v j as the background knowledge related to the user's question obtained by query. Or take the content of the first n (n is a set value, such as 3) knowledge paragraphs with cosine similarity greater than this threshold as the background knowledge related to the user's question. If there is no knowledge paragraph with cosine similarity greater than the threshold β t , then the returned background knowledge is an empty string.
[0125] The use of the above background knowledge information can help the large language model to more clearly define the current user question, so as to obtain a more accurate response result.
[0126] Step 2: First agent task execution.
[0127] After the background knowledge extraction is completed, the first agent starts to parse and execute the user question solving task. The specific steps are as follows:
[0128] Step 2.1: Initialization of the first agent. The first agent clears the current thinking record corresponding to the first agent of this user in the memory module and clears the original data record corresponding to this user.
[0129] Step 2.2: LLM prompt generation. The first agent splices the user question, background knowledge, the historical conversation record of this user, and the above-mentioned current thinking record corresponding to the first agent into the first prompt template corresponding to the first agent to generate the first prompt.
[0130] Among them, since during the first round of execution, after the initialization process in Step 2.1, the current thinking record corresponding to the first agent has been cleared, so the empty string is spliced into the first prompt template. Background knowledge and historical conversation records are optional.
[0131] For easy understanding, the following takes the intelligent transportation scenario as an example to give an example of a first prompt template:
[0132] You are an expert in the field of transportation. Please answer the user's question, and specific data needs to be referred to when answering.
[0133] First, the following is the background knowledge that may be related to the question:
[0134]
Splice background knowledge
[0135]
Splice historical conversation record
[0136] You need to answer step by step. Each step consists of three parts, including "Thought", "Action", and "Feedback", in the following format:
[0137] Thought: Summarize by referring to historical steps and results
[0138] Action: [Tool name] Action input information
[0139] Feedback: Include the execution result of the action
[0140] The optional tools for each action are as follows:
[0141] [summary]: When all the information related to the question has been obtained, use this tool to summarize and end this round of Q&A.
[0142] [data]: A data acquisition tool that can be called to query specified data. It is necessary to use Chinese in the input to fully describe requirements such as the object, quantity, and metrics to be queried.
[0143] [askuser]: When the tool fails to run or the instruction is unclear, use this tool to request further suggestions from the user.
[0144] Let's start!
[0145] Question: [Concatenate the user's question]
[0146] Current thinking record: [Concatenate the current thinking record]
[0147] In the above example of the first prompt template, the content in [ ] will be updated, that is, it will concatenate the corresponding user question, background knowledge, historical conversation record, and the current thinking record of the first agent in the input, while other content remains unchanged. And the first toolset information of the above example is given in the first prompt template.
[0148] Thus, the first prompt template is used to prompt the format of the task information required for each step when the LLM gradually solves the user's question. Among them, each task information includes thinking information, action information, and feedback identification. The thinking information is used to describe the task to be completed, the action information is used to describe the tool to be selected from the first toolset information and the input parameters of the tool for executing the task, and the feedback identification is used to indicate the filling position of the execution result of the task, where the LLM does not output the execution result.
[0149] In addition, taking the example of "query the congestion index of Yuhang District and Xihu District today" in the above text, when the first agent calls the LLM for the first time, the "current thinking record" concatenated into the first prompt template is empty, so the generated first prompt does not contain the relevant content of this field. When the first agent calls the LLM for the second time, the "current thinking record" concatenated into the first prompt template is "Thought: The user needs to query the congestion index of Yuhang District and Xihu District today. First, query the congestion index of Yuhang District; Action: [data] Query the congestion index of Yuhang District today; Feedback: The congestion index of Yuhang District today is 1.3", so the generated first prompt will contain the relevant content of this field as above at this time.
[0150] Step 2.3, The LLM generates an action plan. The first agent sends the first prompt generated in Step 2.2 to the LLM, and the LLM outputs "thinking information" and "action information" in natural language as the actions to be performed by the first agent in this round.
[0151] Taking the example of "querying the congestion index of Yuhang District and Xihu District today" mentioned above, when the first intelligent agent calls the LLM in the first round, the LLM outputs the following solution task information:
[0152] Thought: The user needs to query the congestion index of Yuhang District and Xihu District today. First, query the congestion index of Yuhang District;
[0153] Action: [data] Query the congestion index of Yuhang District today;
[0154] Feedback:
[0155] Step 2.4. The first intelligent agent executes the action.
[0156] In practical applications, the first intelligent agent can parse the content in the above solution task information continuously output by the LLM in real time. When it parses the word "Feedback", it can control the LLM to pause. The first intelligent agent parses the output result of the LLM in Step 2.3 to obtain the "Thought Information" and "Action Information" of this round.
[0157] If the result returned in Step 2.3 has a format error or a missing field, resulting in a parsing failure, the first intelligent agent will use the corresponding error information as the "Feedback" result of this round and jump to Step 2.5;
[0158] If the tool called in the "Action Information" returned in Step 2.3 is "summary, i.e., summarize", then jump to Step 3;
[0159] If the tool called in the "Action Information" returned in Step 2.3 is "askuser, i.e., ask the user", then enter the interactive module process. The interactive module uses the "Action Input Information" included in the "Action Information" as the content to ask the user. In the above example, the action input information is: Query the congestion index of Yuhang District today. The interactive module obtains the answer content input by the user. The first intelligent agent uses the user's answer content as the "Feedback" result of this round and jumps to Step 2.5;
[0160] If the tool called in the "Action Information" returned in Step 2.3 is "data, i.e., call the second intelligent agent", then enter the task execution process of the second intelligent agent. The first intelligent agent uses the "Action Input Information" generated in this round (such as querying the congestion index of Yuhang District today) as the solution task information input to the second intelligent agent, or uses the "Thought Information" and "Action Information" generated in this round as a whole as the solution task information input to the second intelligent agent, and jumps to Step 2.4.a1.
[0161] The above is the execution process of one round of the "Thought / Action / Feedback" mechanism of the first intelligent agent. Next, the execution process of one round of the "Thought / Action / Feedback" mechanism of the second intelligent agent will be introduced.
[0162] Step 2.4.a1, Initialization of the second agent. The second agent clears the current thinking record corresponding to the second agent in the memory module.
[0163] Step 2.4.a2, Generation of LLM prompt words. The second agent splices the input of the second agent, the current thinking record corresponding to the second agent, and the list of available APIs into the second prompt word template corresponding to the second agent to generate the second prompt word.
[0164] Optionally, the above list of available APIs as the second toolset information can be directly included in the second prompt word template, in which case there is no need for splicing. The input of the second agent is a certain solution task information output by the first agent. Similar to the current thinking record corresponding to the first agent, when the second agent is called for the first time, the current thinking record corresponding to the second agent is empty.
[0165] In fact, the current thinking record corresponding to the second agent and the current thinking record corresponding to the first agent are both used to store the information of the intermediate execution results generated when the corresponding agent executes the "thinking / acting / feedback" mechanism in each round.
[0166] The second prompt word template is quite similar to the first prompt word template. Now, the main differences are described:
[0167] First, the content to be spliced is different. The second prompt word template can include the following splicing fields:
Splice the current thinking record
Splice the solution task information
[0168] Second, the toolset information is different. The second toolset information included in the second prompt word template can be various API information required when interacting with the application server, such as API1 for querying the congestion index, API2 for querying the travel volume, and API3 for querying traffic accidents. In addition, a special APIx (returning the final answer) can also be included in the second toolset, which is used to feedback the task execution result to the first agent after the second agent completes the processing of the current received solution task information.
[0169] Third, the format of the action information is different. The format of the action information in the first prompt word template is: [Tool name] Action input information, where the action input information is in the form of natural language text. In the second prompt word template, the format of the action input information can adopt the json format, such as {"parameter1": "value1"; "parameter2": "value2"}. For example, parameter1 = region, value1 = Xihu District.
[0170] Step 2.4.a3. The LLM generates an action plan. The second agent sends the second prompt word generated in Step 2.4.a2 to the LLM, and the LLM outputs "thinking information" and "action information" in natural language form as the actions to be performed by the second agent in this round.
[0171] Step 2.4.a4. The second agent executes the action. The second agent parses the "thinking information" and "action information" output in Step 2.4.a3.
[0172] If the parsing fails, the second agent takes the corresponding error message as the "feedback" result of this round and jumps to Step 2.4.a5;
[0173] If the tool called in the "action information" returned by Step 2.4.a3 is "APIx returns the final answer", it means that the second agent has completed the execution of the current problem-solving task information. The "action input information" contained in this "action information" is sent to the first agent as the output of the second agent, and it jumps to Step 2.5; at this time, the "action input information" is the task execution result corresponding to the current problem-solving task information.
[0174] If the tool called in the "action information" returned by Step 2.4.a3 is some other API name in the second tool set, when it is checked and determined that the parameter format in the "action input information" is the correct parameter format for this API, the corresponding API is called, the call parameter is the parameter given in the "action input information", the result returned by the call to the API is taken as the "feedback" result of this round, and the original data queried by the API is stored in the memory module, and it jumps to Step 2.4.a5; if the call parameter is incorrect, or other exception information occurs, the error message is taken as the "feedback" result of this round, and it jumps to Step 2.4.a5.
[0175] Step 2.4.a5. Save the intermediate step results of the second agent.
[0176] The "thinking information", "action information" and "feedback information" of the second agent in this round are taken as a set of data in this round and added to the corresponding current thinking record of the second agent. If the number of data groups in the current thinking record corresponding to the second agent is greater than the set threshold, it is considered that the second agent still cannot find the correct result after multiple loops. The error message is taken as the final output of the second agent and sent to the first agent, and it jumps to Step 2.5; otherwise, it jumps to Step 2.4.a2 to execute the next round of iteration.
[0177] Here, it should be noted that at this time, it jumps to step 2.4.a2, which often corresponds to the situation where the second intelligent agent needs to execute multiple rounds of the "think / act / feedback" mechanism to complete the currently received problem-solving task information. If it reaches the action of "calling APIx", it will no longer jump to step 2.4.a2.
[0178] Step 2.5: Saving the intermediate step results of the first intelligent agent. The first intelligent agent takes the "thinking information", "action information", and "feedback information" of this round as a set of data for this round and adds them to the corresponding current thinking record of the first intelligent agent. If the number of data groups in its current thinking record is greater than the set threshold, it is considered that the first intelligent agent still cannot find the correct result after multiple loops, and it jumps to step 3; otherwise, it jumps to step 2.2 to execute the next round of iteration.
[0179] Step 3: Task summary and return result
[0180] When the first intelligent agent executes to a certain round, if the tool name parsed from the corresponding "action information" is: summary (i.e., summary), it is determined that the problem-solving task of the user's question has been completed. At this time, it will call the summary module to summarize and analyze the execution results of each intermediate step obtained, that is, the task execution results of multiple problem-solving task information sent to the second intelligent agent, to give the reply information finally output to the user.
[0181] Step 3.1: Generating the summary prompt. The first intelligent agent adds the user's question, the background knowledge extracted in step 1.3, and the execution results of each intermediate step stored in step 2.5, that is, the task execution results corresponding to each problem-solving task information stored in the current thinking record of the first intelligent agent, to the third prompt template to generate the third prompt for guiding the LLM to summarize.
[0182] Step 3.2: Generating the final result. The first intelligent agent sends the third prompt generated in step 3.1 to the LLM to obtain the reply information output by the LLM for output to the user. Additionally, optionally, the first intelligent agent can also obtain the original data obtained from the application server from the memory module, and feedback this original data to the user together, or perform visual display processing on the original data in a certain visual graph manner for output to the user.
[0183] Combined with the above introduction, it can be seen that the LLM in the embodiments of the present invention mainly plays the role of planning and decision-making. In each step of problem-solving, the information provided to the LLM includes: user questions, relevant background knowledge, available tools, and completed problem-solving steps; the LLM makes judgments based on this information and generates the next thinking and actions.
[0184] The following gives an example of a third prompt template, which can be composed of the following information:
[0185] You are a robot based on a large language model that can organize query results and answer users' questions.
[0186] Potentially relevant background knowledge (if the background knowledge is not relevant to the original question, ignore the background knowledge): [Spliced background knowledge]
[0187] Query results (possibly from different channels and methods): [Spliced feedback results for each step]
[0188] Original question: [Spliced user question]
[0189] Now, based on the above results, give a detailed answer to the original question, paying attention to professionalism and rigor, not deviating from the topic, and ignoring results that are not relevant to the question. If the query results do not contain relevant knowledge, use the knowledge of the model itself to answer. Please answer in Chinese.
[0190] The answer to the original question is as follows:.
[0191] The following gives two example scenarios of the execution process of a round of "Think / Act / Feedback" mechanism of the second agent.
[0192] First, assume that the solution task information received from the first agent is simplified as: "Query the congestion index of Yuhang District today", and the task execution information output by the LLM based on the corresponding second prompt is as follows:
[0193] Think: Query the congestion index of Yuhang District today;
[0194] Act: [API1]{Region: Yuhang District; Time: today};
[0195] Feedback:
[0196] Among them, assume that API1 is an API for querying the congestion index.
[0197] When the second agent queries the congestion index of Yuhang District today from the application server by calling API1 based on the above action information, the feedback result is obtained: The congestion index of Yuhang District today is 1.3.
[0198] In the above example scenario, since the solution task is relatively simple, calling the API once can complete the query of the corresponding data. After that, in the process of the next round of "Think / Act / Feedback" mechanism, the following task execution information will finally be formed:
[0199] Think: The congestion index of Yuhang District today has been queried, and the result can be returned;
[0200] Action: [APIx] {The congestion index in Yuhang District today is 1.3};
[0201] Feedback: The congestion index in Yuhang District today is 1.3.
[0202] Second, assume that the solution task information received from the first agent is simplified as: "Query the congestion index and traffic accidents that occurred in Yuhang District today". Then, at this time, the second agent will need to execute three rounds of "Think / Act / Feedback" mechanisms to obtain the final result. The process is as follows:
[0203] 1. Think: Query the congestion index in Yuhang District today;
[0204] 2. Action: [API1] {Region: Yuhang District; Time: Today};
[0205] 3. Feedback: The congestion index in Yuhang District today is 1.3;
[0206] 4. Think: The congestion index of Yuhang District has been obtained. Now, it is necessary to query the traffic accidents that occurred in Yuhang District today;
[0207] 5. Action: [API2] {Region: Yuhang District; Time: Today};
[0208] 6. Feedback: There were 3 traffic accidents in Yuhang District today;
[0209] 7. Think: The congestion index and traffic accidents that occurred in Yuhang District have been obtained. The result can be returned;
[0210] 8. Action: [APIx] {The congestion index in Yuhang District today is 1.3, and there were 3 traffic accidents in Yuhang District today};
[0211] 9. Feedback: The congestion index in Yuhang District today is 1.3, and there were 3 traffic accidents in Yuhang District today.
[0212] Among them, assume that API2 is the API for querying traffic accidents, and APIx is the API for returning the final answer.
[0213] In summary, in the Q&A processing solution provided by the embodiments of the present invention, by hierarchically disassembling the user's question, the first intelligent agent is responsible for planning the overall solution process of the user's question. According to the user's question and the current available data situation, it determines whether to continue to call the second intelligent agent to complete the sub-solution task, or call the summary module to summarize the execution results of the existing sub-solution tasks and obtain the reply information. The actual task execution logic is responsible for the second intelligent agent, reducing the complexity of each step of thinking of the LLM and reducing the possibility of errors. Moreover, by introducing an external knowledge base, when the user inputs a question, by obtaining background knowledge related to the user's question and inserting it into the prompt words, the background knowledge can be related concepts, or the expert thinking of solving the problem, etc., to improve the ability of the LLM to analyze and solve problems. The second intelligent agent can design the processing flow according to the requirements. It interacts independently with the LLM and has a separate prompt word. At the same time, it can share the historical thinking process record of the first intelligent agent (the solution task information sent to the second intelligent agent can include the thinking information of the first intelligent agent), which can help reduce the impact of the token limit of the LLM on complex tasks.
[0214] Figure 5 It is a flowchart of a Q&A processing method provided by an embodiment of the present invention. This method can be executed by the above-mentioned first intelligent agent. As Figure 5 shown, the method includes the following steps:
[0215] 501. Receive the user's question.
[0216] 502. Concatenate the user's question to the first prompt word template corresponding to the first intelligent agent, so as to obtain, based on the first tool set information and the LLM included in the first prompt word template, multiple solution task information that needs to be called by the second intelligent agent to execute and summary task information that needs to be executed by the first intelligent agent generated by the LLM in sequence.
[0217] Optionally, the first intelligent agent can also: obtain background knowledge information that meets the set conditions for similarity with the user's question from the external knowledge base, and / or obtain the historical conversation record corresponding to the user's question, and concatenate the background knowledge information and / or the historical conversation record to the first prompt word template. Among them, the historical conversation record is the question and the corresponding reply information before the user's question in the multi-round conversation of the same user.
[0218] 503. Send the multiple solution task information to the second intelligent agent in sequence.
[0219] 504. Obtain the task execution results of the multiple solution task information sent by the second intelligent agent.
[0220] Among them, the second intelligent agent obtains the task execution information generated by the LLM corresponding to the current solution task information based on the second prompt word template corresponding to the second intelligent agent, the received current solution task information, and the LLM, obtains the task execution result of the current solution task information from the application server corresponding to the user question according to the task execution information, and sends the task execution result to the first intelligent agent, so that the first intelligent agent generates the next solution task information through the LLM after obtaining the task execution result of the current solution task information. The second prompt word template contains second tool set information.
[0221] It can be seen from this that the multiple solution task information generated by the first intelligent agent is not generated simultaneously, but is gradually generated in the process of iterative loop. That is to say, the generation process of the above multiple solution task information includes:
[0222] After obtaining the task execution result of the current solution task information, splice the current solution task information (such as thinking information and action information) and the task execution result of the current solution task information (such as the feedback result above) into the first prompt word template to obtain the first prompt word, and input the first prompt word into the LLM to obtain the next solution task information generated by the LLM.
[0223] 505. Based on the summary task information, summarize the task execution results of multiple solution task information to determine the reply information corresponding to the user question.
[0224] Optionally, the first intelligent agent can directly splice the task execution results of multiple solution task information together and output them as reply information to the user.
[0225] Or, optionally, the process of determining the reply information corresponding to the user question includes:
[0226] Splice the user question, the task execution results of multiple solution task information, and the background knowledge information that meets the set conditions of the similarity between the user question and the user question obtained from the external knowledge base into the third prompt word template to obtain the third prompt word, and input the third prompt word into the LLM to obtain the reply information output by the LLM.
[0227] The execution process of the first intelligent agent can refer to the relevant descriptions in the foregoing other embodiments and will not be elaborated here.
[0228] Figure 6 It is a flowchart of a question and answer processing method provided by an embodiment of the present invention. This method can be executed by the above-mentioned second intelligent agent. As Figure 6 shown, this method includes the following steps:
[0229] 601. Receive the current solution task information corresponding to the user question sent by the first intelligent agent.
[0230] Among them, the first intelligent agent obtains multiple solution task information that needs to be called by the second intelligent agent and summary task information that needs to be executed by the first intelligent agent, which are successively generated by the large language model (LLM) based on the user question, the first prompt word template corresponding to the first intelligent agent, and the LLM. The first prompt word template contains the first tool set information. The current solution task information is the solution task information that is currently sent to the second intelligent agent among the multiple solution task information.
[0231] 602. Concatenate the current solution task information to the second prompt word template corresponding to the second intelligent agent, so as to obtain the task execution information corresponding to the current solution task information generated by the LLM based on the second tool set information contained in the second prompt word template and the LLM.
[0232] 603. Obtain the task execution result of the current solution task information from the application server corresponding to the user question according to the task execution information.
[0233] 604. Send the task execution result of the current solution task information to the first intelligent agent, so that the first intelligent agent summarizes the task execution results of the multiple solution task information based on the summary task information to determine the reply information corresponding to the user question.
[0234] The execution process of the second intelligent agent can refer to the relevant descriptions in the foregoing other embodiments, and will not be elaborated here.
[0235] The question-and-answer processing device of one or more embodiments of the present invention will be described in detail below. Those skilled in the art can understand that these devices can all be configured by using commercially available hardware components through the steps taught by this solution.
[0236] Figure 7 For the structural schematic diagram of a question-and-answer processing device provided by an embodiment of the present invention, as Figure 7 shown, this device is applied to the first intelligent agent, and the first intelligent agent interacts with the second intelligent agent to complete question-and-answer processing. The first intelligent agent and the second intelligent agent share the same large language model. This device includes: a first receiving module 11, a generating module 12, a first sending module 13, an obtaining module 14, and a summarizing module 15.
[0237] The first receiving module 11 is used to receive the user question.
[0238] The generating module 12 is used to concatenate the user question to the first prompt word template corresponding to the first intelligent agent, so as to obtain multiple solution task information that needs to be called by the second intelligent agent and summary task information that needs to be executed by the first intelligent agent, which are successively generated by the large language model based on the first tool set information contained in the first prompt word template and the large language model.
[0239] The first sending module 13 is configured to sequentially send the multiple solution task information to the second intelligent agent.
[0240] The obtaining module 14 is configured to obtain the task execution results of the multiple solution task information sent by the second intelligent agent; wherein, the second intelligent agent, based on the second prompt word template corresponding to the second intelligent agent, the received current solution task information, and the large language model, obtains the task execution information corresponding to the current solution task information generated by the large language model, obtains the task execution result of the current solution task information from the application server corresponding to the user question according to the task execution information, and sends the task execution result to the first intelligent agent, so that the first intelligent agent generates the next solution task information through the large language model after obtaining the task execution result of the current solution task information, and the second prompt word template includes second tool set information.
[0241] The summarizing module 15 is configured to summarize the task execution results of the multiple solution task information based on the summary task information to determine the reply information corresponding to the user question.
[0242] Figure 7 The device shown can execute the steps provided by the first intelligent agent in the foregoing embodiments. For the detailed execution process and technical effects, refer to the descriptions in the foregoing embodiments and will not be elaborated herein.
[0243] Figure 8 The structure diagram of another question-answering processing device provided by an embodiment of the present invention is shown in Figure 8 As shown, the device is applied to a second intelligent agent that interacts with the first intelligent agent to complete question-answering processing. The first intelligent agent and the second intelligent agent share the same large language model. The device includes: a second receiving module 21, a processing module 22, and a second sending module 23.
[0244] The second receiving module 21 is configured to receive the current solution task information corresponding to the user question sent by the first intelligent agent; wherein, the first intelligent agent, based on the user question, the first prompt word template corresponding to the first intelligent agent, and the large language model, obtains multiple solution task information that needs to be called by the second intelligent agent and summary task information that needs to be executed by the first intelligent agent sequentially generated by the large language model. The first prompt word template includes first tool set information, and the current solution task information is the solution task information currently sent to the second intelligent agent among the multiple solution task information.
[0245] The processing module 22 is configured to splice the current solution task information into the second prompt word template corresponding to the second agent, so as to obtain, based on the second tool set information included in the second prompt word template and the large language model, the task execution information generated by the large language model corresponding to the current solution task information, and obtain the task execution result of the current solution task information from the application server corresponding to the user question according to the task execution information.
[0246] The second sending module 23 is configured to send the task execution result of the current solution task information to the first agent, so that the first agent summarizes the task execution results of the multiple solution task information based on the summary task information to determine the reply information corresponding to the user question.
[0247] Figure 8 The illustrated device can execute the steps provided by the second agent in the foregoing embodiments. For the detailed execution process and technical effects, refer to the descriptions in the foregoing embodiments, which will not be elaborated here.
[0248] In a possible design, the above Figures 7-8 The structure of the illustrated question and answer processing device can be implemented as an electronic device. As Figure 9 shown, the electronic device may include: a processor 31, a memory 32, and a communication interface 33. Among them, executable code is stored on the memory 32. When the executable code is executed by the processor 31, the processor 31 can at least implement the question and answer processing method provided in the foregoing embodiments.
[0249] In an alternative embodiment, the electronic device for executing the question and answer processing method provided in the embodiments of the present invention may be any user terminal, such as a mobile phone, a laptop computer, a PC, or may also be an Extended Reality (XR) device. XR is a general term for various forms such as virtual reality and augmented reality.
[0250] In addition, embodiments of the present invention provide a non-transitory machine-readable storage medium, on which executable code is stored. When the executable code is executed by a processor of an electronic device, the processor can at least implement the question and answer processing method provided in the foregoing embodiments.
[0251] The device embodiments described above are merely illustrative. The network elements described as separate components may or may not be physically separated. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. Those of ordinary skill in the art can understand and implement it without creative labor.
[0252] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of adding a necessary general hardware platform, and of course, it can also be implemented by a combination of hardware and software. Based on such an understanding, the above technical solution, in essence, or the part that contributes to the prior art can be embodied in the form of a computer product. The present invention can adopt the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk memories, CD-ROMs, optical memories, etc.) that contain computer-usable program codes.
[0253] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them. Although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments or perform equivalent replacements for some of the technical features. These modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A question and answer processing system, characterized in that, Including: A first intelligent agent, a second intelligent agent, and an application server. The first intelligent agent and the second intelligent agent share the same large language model; The first intelligent agent is configured to receive a user question, splice the user question into a first prompt template, and based on the first toolset information included in the first prompt template and the large language model, obtain multiple solution task information and summary task information sequentially generated by the large language model, and sequentially send the multiple solution task information to the second intelligent agent; The second intelligent agent is configured to splice the currently received solution task information among the multiple solution task information into a second prompt template, and based on the second toolset information included in the second prompt template and the large language model, obtain task execution information corresponding to the current solution task information, and according to the task execution information, obtain the task execution result of the current solution task information from the application server corresponding to the user question, and send the task execution result to the first intelligent agent, so that the first intelligent agent generates the next solution task information through the large language model after obtaining the task execution result of the current solution task information; The first intelligent agent is configured to summarize the task execution results of the multiple solution task information based on the summary task information to determine the reply information corresponding to the user question.
2. The system according to claim 1, wherein The first intelligent agent is further configured to obtain background knowledge information whose similarity with the user question meets the set conditions from an external knowledge base, and / or obtain the historical conversation record corresponding to the user question; splice the background knowledge information and / or the historical conversation record into the first prompt template, and the historical conversation record is the questions and corresponding reply information before the user question in multiple rounds of conversations of the same user.
3. The system according to claim 1, wherein In the process of generating the next solution task information through the large language model after obtaining the task execution result of the current solution task information, the first intelligent agent is configured to: splice the current solution task information and the task execution result of the current solution task information into the first prompt template to obtain a first prompt, and input the first prompt into the large language model to obtain the next solution task information.
4. The system according to any one of claims 1 to 3, characterized in that, The first prompt template is used to prompt the format of the task information that needs to be generated when the large language model gradually solves the user question; Each task information in the multiple solution task information and the summary task information generated by the large language model includes first thinking information, first action information, and a first feedback identifier. The first thinking information is used to describe the task to be completed, the first action information is used to describe the tool to be selected from the first toolset information for executing the task and the input parameters of the tool, and the first feedback identifier is used to indicate the filling position of the task execution result.
5. The system according to claim 1, characterized in that, The second intelligent agent is further configured to receive the background knowledge information and / or historical conversation records corresponding to the user question sent by the first intelligent agent, and splice the background knowledge information and / or the historical conversation records into the second prompt template. The historical conversation records are the questions and corresponding reply information before the user question in multiple rounds of conversations of the same user, and the background knowledge information is the background knowledge information in the external knowledge base whose similarity to the user question meets the set conditions.
6. The system according to claim 1, wherein The second prompt template is used to prompt the large language model for the format of the task execution information required for the current solution task information. The task execution information includes second thinking information, second action information, and a second feedback identifier. The second thinking information is used to describe the solution task to be completed, the second action information is used to describe the tools and the input parameters of the tools to be selected from the second tool set information for executing the solution task, and the second feedback identifier is used to indicate the filling position of the execution result of the solution task.
7. The system according to claim 6, wherein The number of task execution information corresponding to the current solution task information is multiple. The second intelligent agent is further configured to: after obtaining the corresponding first task execution result from the application server based on the first task execution information, splice the first task execution information and the first task execution result into the second prompt template, and input the obtained second prompt into the large language model to obtain the second task execution information generated by the large language model.
8. The system according to claim 1, wherein In the process that the first intelligent agent summarizes the task execution results of the multiple solution task information to determine the reply information corresponding to the user question, the first intelligent agent is configured to: splice the user question, the task execution results of the multiple solution task information, and the background knowledge information obtained from the external knowledge base whose similarity to the user question meets the set conditions into a third prompt template to obtain a third prompt, and input the third prompt into the large language model to obtain the reply information output by the large language model.
9. A question and answer processing method, characterized in that, Applied to a first intelligent agent, the first intelligent agent interacts with a second intelligent agent to complete the question and answer processing method. The first intelligent agent and the second intelligent agent share the same large language model. The method includes: Receiving a user question. Splicing the user question into a first prompt template to obtain multiple solution task information and summary task information sequentially generated by the large language model based on the first tool set information included in the first prompt template and the large language model. Send the multiple solution task information to the second intelligent agent in sequence, so that the second intelligent agent obtains the task execution information corresponding to the current solution task information based on the second prompt word template, the received current solution task information, and the large language model, and obtains the task execution result of the current solution task information from the application server corresponding to the user question according to the task execution information, and sends the task execution result to the first intelligent agent. The second prompt word template contains second tool set information; Obtain the task execution results of the multiple solution task information sent by the second intelligent agent; Based on the summary task information, summarize the task execution results of the multiple solution task information to determine the reply information corresponding to the user question.
10. The method according to claim 9, wherein The method further includes: Obtain background knowledge information whose similarity to the user question meets the set conditions from an external knowledge base, and / or obtain the historical conversation record corresponding to the user question; Splice the background knowledge information and / or the historical conversation record into the first prompt word template. The historical conversation record is the questions and corresponding reply information before the user question in multiple rounds of conversations of the same user.
11. The method according to claim 9, characterized in that The generation process of the multiple solution task information includes: After obtaining the task execution result of the current solution task information, splice the current solution task information and the task execution result of the current solution task information into the first prompt word template to obtain a first prompt word; Input the first prompt word into the large language model to obtain the next solution task information.
12. The method according to any one of claims 9 to 11, characterized in that The summarizing the task execution results of the multiple solution task information to determine the reply information corresponding to the user question includes: Splice the user question, the task execution results of the multiple solution task information, and the background knowledge information whose similarity to the user question meets the set conditions obtained from the external knowledge base into the third prompt word template to obtain a third prompt word; Input the third prompt word into the large language model to obtain the reply information output by the large language model.
13. A question-and-answer processing method, characterized in that, Applied to the second intelligent agent that interacts with the first intelligent agent to complete the question and answer processing method. The first intelligent agent and the second intelligent agent share the same large language model. The method includes: Receive the current solution task information corresponding to the user question sent by the first intelligent agent. Among them, the first intelligent agent obtains multiple solution task information and summary task information sequentially generated by the large language model based on the user question, the first prompt word template, and the large language model. The first prompt word template contains first tool set information. The current solution task information is the solution task information currently sent to the second intelligent agent among the multiple solution task information; Splice the current solution task information into the second prompt word template to obtain the task execution information corresponding to the current solution task information generated by the large language model based on the second tool set information contained in the second prompt word template and the large language model; Obtain the task execution result of the current solution task information from the application server corresponding to the user question according to the task execution information; Send the task execution result of the current solution task information to the first intelligent agent, so that the first intelligent agent summarizes the task execution results of the multiple solution task information based on the summary task information to determine the reply information corresponding to the user question.
14. An electronic device, characterized in that, Including: A memory, a processor, and a communication interface; wherein, executable code is stored on the memory, and when the executable code is executed by the processor, the processor executes the question and answer processing method according to any one of claims 9 to 12 or claim 13.
15. A non-transitory machine-readable storage medium, characterized in that, Executable code is stored on the non-transitory machine-readable storage medium, and when the executable code is executed by the processor of the electronic device, the processor executes the question and answer processing method according to any one of claims 9 to 12 or claim 13.
Citation Information
Cited By
Task execution method, device and system, electronic device and storage medium
CN120723411A
Intelligent agent-based information processing method and device and storage medium
CN121031788A