An agent question and answer method and device, a storage medium and an apparatus
Patent Information
- Application Number
- CN202410646816.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-05-23
- Publication Date
- 2026-09-29
- Estimated Expiration
- 2044-05-23
AI Technical Summary
[0003]目前,实现智能体问答的方式通常有两种:一种是通过推理行动(Reasoning andActing,ReAct)方式,该方式的效果完全取决于大语言模型的理解能力,对于大部分的闭源大语言模型以及开源大语言模型,都无法满足ReAct框架的准确率要求,并且该方式中大语言模型相关经验无法融入到模型内部,在行动过程中也无法与用户进行交互,导致效果较差
[0050]本申请实施例提供的一种智能体问答方法、装置、存储介质及设备,首先接收目标用户向智能体问答系统发出的问题指令,并根据问题指令确定待答复的目标问题,然后调用预设的大语言模型对目标问题进行分解,得到目标问题包含的N个目标子问题;其中,N为大于0的正整数;接着在循环遍历N个目标子问题时,将第i个目标子问题和第i-1个目标子问题的答复结果输入智能体问答系统的函数调用单元进行求解,得到第i个目标子问题的答复结果,依次类推,直至得到N个目标子问题各自对应的答复结果;其中,i为大于2且不大于N的正整数。最后将N个目标子问题各自对应的答复结果和目标问题对应的文本,输入至预设的大语言模型,得到大语言模型输出的针对目标问题的答复内容,并通过智能体问答系统展示给目标用户。
Smart Images

Figure CN118469017B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of computer technology, and in particular to an intelligent agent question-answering method, apparatus, storage medium and device. Background Technology
[0002] With the rapid development of information technologies such as artificial intelligence and the Internet of Things, the application scenarios of human-computer interaction are becoming increasingly widespread. Various large language models (LLMs) are appearing in people's lives and work, such as the Chat Generative Pre-trained Transformer (ChatGPT). Intelligent agents can use large language models as their core to complete various language tasks, such as question answering and text generation, to assist users in fulfilling various behavioral intentions.
[0003] Currently, there are generally two ways to implement agent-based question answering: one is through Reasoning and Acting (ReAct), the effectiveness of which depends entirely on the understanding ability of the large language model. Most closed-source and open-source large language models cannot meet the accuracy requirements of the ReAct framework, and the relevant experience of the large language model cannot be integrated into the model's internal structure, nor can it interact with the user during the action, resulting in poor performance. The second method is through function calls, which typically only calls the function once during execution, leading to poor response results for complex multi-task problems. Therefore, both existing methods of implementing agent-based question answering result in poor agent response performance, thus degrading the user's question-and-answer experience. Summary of the Invention
[0004] The main objective of this application is to provide an intelligent agent question-answering method, apparatus, storage medium, and device that can improve the response effect of the intelligent agent.
[0005] This application provides an intelligent agent question-answering method, including:
[0006] Receive a question instruction from a target user to the intelligent agent question-answering system, and determine the target question to be answered based on the question instruction;
[0007] The target problem is decomposed by calling a preset large language model to obtain N target sub-problems contained in the target problem; where N is a positive integer greater than 0.
[0008] When iterating through the N target sub-problems, the answers to the i-th target sub-problem and the (i-1)-th target sub-problem are input into the function call unit of the intelligent agent question answering system for solving, so as to obtain the answer to the i-th target sub-problem. This process is repeated until the answer to each of the N target sub-problems is obtained. Here, i is a positive integer greater than 1 and not greater than N.
[0009] The answers to the N target sub-questions and the text corresponding to the target question are input into the preset large language model to obtain the answer content for the target question output by the large language model, and then displayed to the target user through the intelligent agent question answering system.
[0010] In one possible implementation, the step of inputting the answers to the i-th target sub-question and the (i-1)-th target sub-question into the function call unit of the agent question-answering system for solving, to obtain the answer to the i-th target sub-question, includes:
[0011] After inputting the answers to the i-th target sub-problem and the (i-1)-th target sub-problem into the function call unit of the intelligent agent question answering system, the i-th target sub-problem is solved by calling the corresponding tool function and the parameters required by the tool function, and the answer to the i-th target sub-problem is obtained.
[0012] In one possible implementation, the parameters required by the tool function are extracted from the answers to the i-th target subproblem and the (i-1)-th target subproblem.
[0013] In one possible implementation, the parameters required by the tool function are obtained through interaction with the target user.
[0014] In one possible implementation, the tool function includes at least one of a general tool function, a target enterprise internal system call tool function, and a target enterprise internal professional knowledge base question and answer tool function.
[0015] In one possible implementation, the utility function is generated by user-defined configuration in a page-like manner.
[0016] In one possible implementation, after inputting the answers to the i-th and (i-1)-th target sub-problems into the function calling unit of the intelligent agent question-answering system, the i-th target sub-problem is solved by calling the corresponding tool function and the parameters required by the tool function to obtain the answer to the i-th target sub-problem, including:
[0017] After inputting the answers to the i-th target sub-question and the (i-1)-th target sub-question into the function call unit of the intelligent agent question answering system, the preset large language model is called to integrate the answers to the i-th target sub-question and the (i-1)-th target sub-question to obtain the updated i-th target sub-question;
[0018] By searching the historical information of each target sub-problem recorded in the pre-built agent history table, it is determined whether the completion flag bit in the agent history table is a completion flag;
[0019] If yes, or if there are no question-and-answer records in the session corresponding to the target question, then the historical information of the i-th target sub-question is set to empty; if no, or if there are question-and-answer records in the session corresponding to the target question, then the historical information of the i-th target sub-question is read.
[0020] Based on the updated i-th target sub-question and the historical information of the i-th target sub-question, obtain the parameters required by the called tool function, and determine whether the called tool function is a question-answering tool function;
[0021] If so, then using the parameters required by the question-answering tool function, the preset question-answering tool is called in conjunction with the large language model to solve the updated i-th target sub-question, and the answer result of the i-th target sub-question is obtained; and the historical information of the i-th target sub-question is updated, and the completion flag in the agent's history table is set to the completion flag;
[0022] If not, then determine whether the called utility function is a utility function that requires user information authentication;
[0023] When it is determined that the called utility function is a utility function that requires user information authentication, the user information contained in the target question is verified.
[0024] Upon successful verification, the historical information of the i-th target sub-problem is updated. Based on the updated historical information of the i-th target sub-problem, the utility function to be called is determined to solve the updated i-th target sub-problem and obtain the response result of the i-th target sub-problem. The historical information of the i-th target sub-problem is also updated, and the completion flag in the agent's history table is set to the completion flag.
[0025] This application also provides an intelligent agent question-answering device, including:
[0026] The determining unit is used to receive a question instruction sent by a target user to the intelligent agent question-answering system, and determine the target question to be answered based on the question instruction;
[0027] The decomposition unit is used to call a preset large language model to decompose the target problem into N target sub-problems; where N is a positive integer greater than 0.
[0028] The traversal unit is used to input the answer results of the i-th target sub-problem and the (i-1)-th target sub-problem into the function call unit of the intelligent agent question answering system for solving when iterating through the N target sub-problems, so as to obtain the answer result of the i-th target sub-problem, and so on, until the answer results corresponding to each of the N target sub-problems are obtained; where i is a positive integer greater than 2 and not greater than N;
[0029] The obtaining unit is used to input the answer results corresponding to each of the N target sub-questions and the text corresponding to the target question into the preset large language model, obtain the answer content for the target question output by the large language model, and display it to the target user through the intelligent agent question answering system.
[0030] In one possible implementation, the traversal unit is specifically used for:
[0031] After inputting the answers to the i-th target sub-problem and the (i-1)-th target sub-problem into the function call unit of the intelligent agent question answering system, the i-th target sub-problem is solved by calling the corresponding tool function and the parameters required by the tool function, and the answer to the i-th target sub-problem is obtained.
[0032] In one possible implementation, the parameters required by the tool function are extracted from the answers to the i-th target subproblem and the (i-1)-th target subproblem.
[0033] In one possible implementation, the parameters required by the tool function are obtained through interaction with the target user.
[0034] In one possible implementation, the tool function includes at least one of a general tool function, a target enterprise internal system call tool function, and a target enterprise internal professional knowledge base question and answer tool function.
[0035] In one possible implementation, the utility function is generated by user-defined configuration in a page-like manner.
[0036] In one possible implementation, the traversal unit includes:
[0037] The integration subunit is used to input the answer results of the i-th target sub-question and the (i-1)-th target sub-question into the function call unit of the intelligent agent question answering system, and then call the preset large language model to integrate the answer results of the i-th target sub-question and the (i-1)-th target sub-question to obtain the updated i-th target sub-question;
[0038] The first judgment subunit is used to determine whether the completion flag bit in the agent history table is a completion flag by searching the historical information of each target sub-problem recorded in the pre-built agent history table;
[0039] The reading sub-unit is used to set the historical information of the i-th target sub-question to empty if, in the case of yes, or no question-and-answer records under the session corresponding to the target question; otherwise, if, question-and-answer records exist under the session corresponding to the target question, the historical information of the i-th target sub-question is read.
[0040] The second judgment subunit is used to obtain the parameters required by the called tool function based on the updated i-th target sub-question and the historical information of the i-th target sub-question, and to determine whether the called tool function is a question-answering tool function;
[0041] The first solving subunit is used to, if yes, utilize the parameters required by the question-answering tool function to call a preset question-answering tool in conjunction with a large language model to solve the updated i-th target sub-question, obtain the answer result of the i-th target sub-question; and update the historical information of the i-th target sub-question, and set the completion flag in the agent's history table to a completion flag;
[0042] The third judgment subunit is used to determine whether the called utility function is a utility function that requires user information authentication if the condition is not met.
[0043] The verification subunit is used to verify the user information contained in the target question when it is determined that the called utility function is a utility function that requires user information authentication.
[0044] The second solution subunit is used to update the historical information of the i-th target sub-problem when the verification passes, determine the tool function to be called based on the updated historical information of the i-th target sub-problem, so as to solve the updated i-th target sub-problem and obtain the answer result of the i-th target sub-problem; and update the historical information of the i-th target sub-problem, and set the completion flag in the agent history table to the completion flag.
[0045] This application also provides an intelligent agent question-answering device, including: a processor, a memory, and a system bus;
[0046] The processor and the memory are connected via the system bus;
[0047] The memory is used to store one or more programs, the one or more programs including instructions, which, when executed by the processor, cause the processor to perform any of the above-described implementations of the intelligent agent question-answering method.
[0048] This application also provides a computer-readable storage medium storing instructions that, when executed on a terminal device, cause the terminal device to perform any of the above-described implementations of the intelligent agent question-answering method.
[0049] This application also provides a computer program product, which, when run on a terminal device, causes the terminal device to execute any of the above-described intelligent agent question-answering methods.
[0050] This application provides an intelligent agent question-answering method, apparatus, storage medium, and device. First, it receives a question instruction from a target user to the intelligent agent question-answering system. Based on the question instruction, it determines the target question to be answered. Then, it calls a preset large language model to decompose the target question, obtaining N target sub-questions; where N is a positive integer greater than 0. Next, while iterating through the N target sub-questions, the answer results of the i-th target sub-question and the (i-1)-th target sub-question are input into the function call unit of the intelligent agent question-answering system for solving, obtaining the answer result of the i-th target sub-question. This process continues until the answer results corresponding to each of the N target sub-questions are obtained; where i is a positive integer greater than 2 and not greater than N. Finally, the answer results corresponding to each of the N target sub-questions and the text corresponding to the target question are input into the preset large language model to obtain the answer content for the target question output by the large language model, which is then displayed to the target user through the intelligent agent question-answering system.
[0051] As can be seen, this application combines reasoning action (ReAct) and function call, using the function call unit of the intelligent agent question answering system as the smallest implementation unit, and combining the three steps of thinking, execution and observation of ReAct, to achieve accurate answers to the target question posed by the target user containing N target sub-questions, thereby improving the question answering experience of the target user. Attached Figure Description
[0052] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0053] Figure 1 A flowchart illustrating an intelligent agent question-answering method provided in an embodiment of this application;
[0054] Figure 2 A diagram illustrating the overall implementation process of the intelligent agent question-answering method provided in this application embodiment;
[0055] Figure 3 This is a process diagram illustrating the solution of a target sub-problem through a function call unit of an intelligent agent question-answering system, as provided in an embodiment of this application.
[0056] Figure 4 Example diagram of the overall structure of the intelligent agent question-answering system provided in the embodiments of this application;
[0057] Figure 5 This is a schematic diagram of the composition of an intelligent agent question-answering device provided in an embodiment of this application. Detailed Implementation
[0058] Large Language Model (LLM) agents refer to functional systems that use powerful Large Language Models (LLMs) as their core to perform language tasks such as question answering and text generation. Currently, users obtain answers to questions through agent-based question answering systems primarily in the following two ways:
[0059] (1) The Reasoning-Based Action (ReAct) approach leverages the logical reasoning capabilities of a large language model to construct a complete series of actions (Acts) to achieve the desired goal. A ReAct process typically includes the following three steps:
[0060] Thinking: Generated by the large language model, it serves as the basis for the large language model's behavior. The rationality of the proposed action can be assessed based on the thinking process of the large language model. This is a key criterion for judging the rationality of a decision. Compared to the human brain, the existence of thinking makes the decisions of the large language model more interpretable and credible.
[0061] Action: Specifically, this refers to the specific behavior that the large language model determines needs to be executed. An action generally consists of two parts: the behavior itself and the object, i.e., what tool to use and what the tool's parameters are. The biggest advantage of the large language model is its ability to select the necessary tool and generate the parameters to be entered into it based on its deliberate judgment. This ensures the feasibility of the ReAct framework at the execution level.
[0062] Observation: Synchronize the action results with the large language model to assist the large language model in further analysis or decision-making.
[0063] However, this ReAct framework has the following drawbacks:
[0064] First, ReAct's effectiveness depends entirely on the understanding ability of the large model. Most closed-source and open-source large language models cannot meet the accuracy requirements of the ReAct framework.
[0065] Second, after completing each task, the large language model cannot accumulate experience, meaning that relevant experience cannot be integrated into the model.
[0066] Third, although the functions required by the tool can be extracted during the action, if the user does not provide the relevant parameters in the task, the ReAct execution will fail and the user will no longer be able to interact (i.e., human-computer interaction cannot be achieved).
[0067] (2) Function Call: While this method addresses the aforementioned shortcomings of the ReAct framework, and most open-source large language models currently support Function Call applications well, it also has drawbacks. It internalizes external calls into the model through pre-training / fine-tuning techniques, making it a native capability that ensures iterative optimization. Furthermore, if suitable parameters are not extracted during the function call, it can throw an exception and interact with the user (i.e., enable human-computer interaction). However, it directly matches the function tool based on the user's task, resulting in a single function call. For complex multi-task scenarios, this Function Call approach performs poorly.
[0068] Therefore, both of the above-mentioned methods for implementing intelligent agent question answering result in poor agent response, thereby reducing the user's question answering experience.
[0069] To address the aforementioned shortcomings, this application provides an intelligent agent question-answering method. First, it receives a question instruction from a target user to the intelligent agent question-answering system and determines the target question to be answered based on the instruction. Then, it calls a preset large language model to decompose the target question into N target sub-questions, where N is a positive integer greater than 0. Next, while iterating through the N target sub-questions, the answers to the i-th and (i-1)-th target sub-questions are input into the function call unit of the intelligent agent question-answering system for solving, obtaining the answer to the i-th target sub-question. This process continues until the answer to each of the N target sub-questions is obtained, where i is a positive integer greater than 2 and not greater than N. Finally, the answer to each of the N target sub-questions and the text corresponding to the target question are input into the preset large language model to obtain the response content for the target question output by the large language model, which is then displayed to the target user through the intelligent agent question-answering system.
[0070] As can be seen, this application combines reasoning action (ReAct) and function call, using the function call unit of the intelligent agent question answering system as the smallest implementation unit, and combining the three steps of thinking, execution and observation of ReAct, to achieve accurate answers to the target question posed by the target user containing N target sub-questions, thereby improving the question answering experience of the target user.
[0071] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0072] First Embodiment
[0073] See Figure 1 This is a flowchart illustrating an intelligent agent question-answering method provided in this embodiment. The method includes the following steps:
[0074] S101: Receives a question instruction from the target user to the intelligent agent question-answering system, and determines the target question to be answered based on the question instruction.
[0075] In this embodiment, any user who issues a question command to the intelligent agent question-answering system is defined as the target user. Furthermore, this embodiment does not limit the specific manner in which the target user issues the question command; for example, the target user can issue a question command to the intelligent agent question-answering system via voice or text input. After receiving the command from the target user, the intelligent agent question-answering system can further determine the question to be answered based on the command and define it as the target question for execution of subsequent step S102. This embodiment does not limit the language type of the target question; for example, the target question can be a Chinese question or an English question. This embodiment also does not limit the length of the target question.
[0076] It should be noted that, in order to improve the response effect of agent question answering, after determining the target question to be answered, this application proposes a method that combines reasoning action (ReAct) and function call. The function call unit of the agent question answering system is used as the smallest implementation unit. Combining the three steps of ReAct—thinking, execution, and observation—a more accurate answer to the target question is provided. The overall implementation process is as follows: Figure 2 As shown, this can improve the question-and-answer experience for the target users.
[0077] S102: Call the preset large language model to decompose the target problem and obtain N target subproblems contained in the target problem; where N is a positive integer greater than 0.
[0078] In this embodiment, after the intelligent agent question-answering system determines the target question to be answered in step S101, it can further call a preset large language model (the specific model structure is not limited, such as choosing the open-source large language model ChatGLM3 as the preset large language model, etc.) to decompose the target question into N target sub-questions, which are then used to execute the subsequent step S103. The specific value of N is not limited, but it must be a positive integer greater than 0.
[0079] Large Language Models (LLMs) are deep learning-based language models that can generate new language expressions, such as text, sentences, paragraphs, and even articles, based on input text content. LLMs utilize large-scale language datasets and are trained on language rules and patterns through autoregressive generation. They can simulate human commands to generate language expressions (such as text data). Specifically, when generating new text data, LLMs predict the probability of the next language unit based on previously generated content until complete text data is generated.
[0080] S103: When iterating through N target sub-problems, the answer results of the i-th target sub-problem and the (i-1)-th target sub-problem are input into the function call unit of the intelligent agent question answering system for solving, so as to obtain the answer result of the i-th target sub-problem, and so on, until the answer results corresponding to each of the N target sub-problems are obtained; where i is a positive integer greater than 1 and not greater than N.
[0081] In this embodiment, after obtaining the N target sub-problems contained in the target problem through step S102, the response results of the i-th (i is a positive integer greater than 1 and not greater than N) target sub-problems and the (i-1)-th target sub-problems can be input into the function call unit of the intelligent agent question answering system for solving while iterating through these N target sub-problems. Figure 2 As shown, the answer to the i-th target sub-question is obtained, and so on, until the answer to each of the N target sub-questions is obtained. The answer to each target sub-question can be added to the result list.
[0082] Specifically, one possible implementation is that, after obtaining the N target sub-problems contained in the target problem, the answer results of the i-th target sub-problem and the (i-1)-th target sub-problem can be input into the function call unit of the intelligent agent question answering system. Then, by calling the corresponding tool function and the parameters required by the tool function, the i-th target sub-problem can be solved to obtain the answer result of the i-th target sub-problem.
[0083] The parameters required for the tool function can be extracted from the answers to the i-th and (i-1)-th target sub-questions, or obtained through human-computer interaction with the target user. Furthermore, this application does not limit the specific content or generation method of the tool function. For example, the target function can include, but is not limited to, at least one of the following: general tool functions (such as weather query, text-to-image tool functions, etc.), internal system call tool functions of the target enterprise (such as employee holiday day query tool functions, etc.), and question-and-answer tool functions of the target enterprise's internal professional knowledge base. The tool function can be an existing, mature, general tool function, or a personalized tool function generated by user-defined configuration via a page.
[0084] In this implementation, to more accurately determine the answer to the i-th target sub-question, one possible implementation is as follows: Figure 3 As shown, the intelligent agent question answering system first receives the N target sub-questions contained in the target question returned by the preset large language model, and at the same time carries information such as the username of the target user, the session identifier (id) of the conversation with the target user, and the business robot identifier (id) involved in the target question.
[0085] Secondly, after inputting the responses to the i-th and (i-1)-th target sub-questions into the function call unit of the intelligent agent question-answering system, a pre-defined large language model (such as ChatGLM3) is invoked to integrate the responses to the i-th and (i-1)-th target sub-questions, resulting in the updated i-th target sub-question. This resolves the result dependency issue between the subtasks corresponding to each target sub-question, ensuring the integrity of the parameters provided by the i-th subtask corresponding to the i-th target sub-question.
[0086] Next, by searching the historical information of each target sub-problem recorded in the pre-built agent history table, it is determined whether the completion flag in the agent history table is set to complete. The agent history table records the intermediate process of each sub-task, and the main information may include, but is not limited to, the user input question for the i-th target sub-problem, the response returned by the large model, function calls, and whether the current sub-task has been completed. Generally, executing a function call once is sufficient to complete the i-th sub-task. However, if the target user has not entered relevant parameters or has entered incorrect parameters in the current sub-task, the system can prompt the target user to enter parameters, at which point the completion flag will be set to incomplete.
[0087] Next, if it is determined that the completion flag in the agent's history table is a completion flag, or if there are no question-and-answer records under the session corresponding to the target question, then the historical information of the i-th subtask corresponding to the i-th target sub-question (represented using History, such as...) is... Figure 3 (As shown) is set to empty; if it is determined that the completion flag in the agent's history table is an incomplete flag, or if there is a question-and-answer record under the session corresponding to the target question, the historical information of the i-th target sub-question can be read. Furthermore, the tool configuration table supported by the currently selected business robot is read, and these tools are loaded. The target question raised by the target user is added to the History. Then, the History and Tools information are used as parameters to call the function_call function of the large language model, returning the name of the called tool function and the parameters required for its extraction.
[0088] For example: Suppose the tool list includes two tools: weather search and merchant search. The target user's question is: "Please help me check the weather in Shanghai." The prompt word input to the large language model could be: "Please choose one tool from weather_search (weather search, parameter is city)" and "mchnt_no_search (merchant search, parameter is mchnt (merchant)) to handle the user's question, and return the tool name and parameters. The user's question is: 'Please help me check the weather in Shanghai.'" After receiving this prompt word, the large language model will return the result: `{tool_name:weather_search,params:{city:Shanghai}}`.
[0089] Furthermore, it can be determined whether the called utility function is a question-and-answer utility function (represented by Chat_tool, e.g.) Figure 3 If so, then using the parameters required by the question-answering tool function, the preset question-answering tool is called in conjunction with the large language model to solve the updated i-th target sub-question, and the answer result of the i-th target sub-question is obtained; and the historical information of the i-th target sub-question is updated, and the completion flag in the agent's history table is set to the completion flag.
[0090] For example, such as Figure 3 As shown, the business robot ID information can be used as a parameter to call a preset question-and-answer tool (such as the question-and-answer service tool function of Yinshang Tianyan). By supporting FAQ (Frequently Asked Questions) and knowledge base document question-and-answer functions, and combining Internet question-and-answer / large language model, the updated i-th target sub-question is solved to obtain the answer result of the i-th target sub-question; and the history information (History) of the i-th target sub-question is updated, and the completion flag in the agent's history table is set to the completion flag (set to 1).
[0091] Conversely, if the called utility function is not a question-and-answer utility function (Chat_tool), then a list of tools requiring user authentication is loaded. These tools indicate that the current user can only query information related to themselves, not others, such as checking their own vacation balance or salary. This is used to determine if the currently called utility function requires authentication. If so, the user information contained in the target question is validated to see if it includes other people's information, such as name, employee ID, and username. If not, the validation passes. If it does, further information verification is required. If the verification shows the user is the same person, the verification passes, and the questioner's username is added to the beginning of the question. If the user is not the same person, the verification fails, and the message "You can only query your own information!" is returned. Figure 3 As shown.
[0092] Finally, upon successful verification, the historical information of the i-th target sub-problem can be updated. Based on the updated historical information of the i-th target sub-problem, the utility function to be called is determined to solve the updated i-th target sub-problem and obtain the answer result of the i-th target sub-problem. The historical information of the i-th target sub-problem is also updated, and the completion flag in the agent's history table is set to the completion flag.
[0093] Specifically, such as Figure 3 As shown, upon successful verification, the question-and-answer tool function (Chat_tool) in the Tools list can be deleted. This is because it has been determined that no business question-and-answer is needed, and the agent completion flag can be set to incomplete. Then, the History and Tools information are used as parameters to call the large language model's function_call again. The large language model will then stream the Token character. It checks if the Token is a special character. If it is, the current output of the large language model ends; otherwise, it continues outputting. If it is a special character, it needs to determine what that special character is. If it is <|user|>, it is recorded in the History, and the user is prompted to enter the relevant information again. The process then returns to the step of searching the pre-built agent history table and restarts execution. If it is <|assistent|>, the content output by the large language model before <|assistent|> can be directly recorded in the History, and the function_call function can then be executed. If it's `<|obvervation|>`, then based on the previous output of `<|obvervation|>`, the function and parameters are parsed, and the function is executed by a code executor (such as a Python interpreter). During execution, if the function executes normally, the agent's completion flag is set; if execution fails, the agent's completion flag is set to incomplete, and the Token information describing the role and the function execution information are saved to the history. Finally, the History and Tools information are used as parameters to call the `function_call` function of the large language model to summarize the final output. It should be noted that, usually, the Token will not be a special flag when summarizing in `function_call`, so the last step only needs to save the Token information describing the role and the result to the history and output the result.
[0094] For example: Suppose the i-th sub-task corresponding to the i-th target sub-problem is to query the weather, and the question is "What's the weather like today?", then the possible steps to execute are:
[0095] (1) After selecting the weather query tool, the large language model will output something like "Please enter the city to query <|user|>", where <|user|> is a special character that indicates that the user needs to input content. The user enters "Province B" and submits the query.
[0096] (2) The large language model will determine again that a weather query tool is needed and will output something like "Province B is not a specific city <|assistent|>". <|assistent|> is a special character. At this time, the result "Province B is not a specific city" will continue to be included in the historical information record and will continue to be automatically input into the large language model.
[0097] (3) The large language model can continue to output something like “Province B is not a specific city, please enter the city you want to query <|user|>” based on the new historical information recorded in step (2). Just like in step (1), the user is required to re-enter the information. At this time, the user enters “City A” and continues to ask questions.
[0098] (4) The large language model again determines that a weather query tool is needed, and outputs something like: {tool_name:weather_search,params:{city:city A}}<|obvervation|>, where <|obvervation|> is a special character. Based on the function name weather_search and the parameter {city:Nanjing}, the weather query tool is called. If the tool call fails, the system can output the message "Tool unavailable, call failed" and add it to the historical record. The large language model is called again to summarize, and it will output the result "Tool unavailable, call failed" to inform the user. If the tool executes successfully, the system can output the message "City A has sunny weather today, temperature 23 degrees, humidity 36%". The large language model is called again to summarize, and it will output the result "City A has sunny weather today, temperature 23 degrees, humidity 36%" to inform the user.
[0099] S104: Input the response results corresponding to each of the N target sub-questions and the text corresponding to the target question into the preset large language model, obtain the response content of the large language model for the target question, and display it to the target user through the intelligent agent question answering system.
[0100] In this embodiment, after obtaining the response results corresponding to each of the N target sub-questions through step S103, the response results corresponding to each of the N target sub-questions and the text corresponding to the target question can be further integrated into the prompt instruction and input into a preset large language model (such as ChatGLM3, etc.) to output the response content for the target question through the large language model and display it to the target user through the intelligent agent question answering system. The specific display method is not limited, and it can be text display, animation display, and / or voice broadcast of the response content for the target question through the system display screen, etc.
[0101] For example, suppose the target user is Zhang San. After clicking into the intelligent agent question-answering system, a chat page is created, and the system assigns a session identifier (session_id) of 001. Since Zhang San has multiple business robots, for example, if he selects a robot with the robot identifier (id) abc, based on that robot ID, the system can query the tools available to that business robot, such as: transaction inquiry, legal representative inquiry, and network access time inquiry, etc. When Zhang San sends the target question to the intelligent agent question-answering system with the text "Help me find out who the merchant with the highest transaction amount in the A province branch is, and when they joined the network," firstly, a preset large language model (such as ChatGLM3) can be called to decompose the target question into three sub-questions: "1. Find the merchant with the highest transaction amount in the A province branch; 2. Find the legal representative of the merchant; 3. Find the merchant's network access time."
[0102] Then, while iterating through these three target sub-problems, the three sub-tasks corresponding to these three target sub-problems are executed as follows: First, query the merchant with the highest transaction amount in the A province branch. For example, this would call the transaction query tool, which returns the name: Merchant B. Then, combine the output of the first sub-task with the original second target sub-problem to form a new problem: Given [Query the merchant with the highest transaction amount in the A province branch, the merchant name is Merchant B]. Next, query the merchant's legal representative. This would call the legal representative query tool, which returns the answer: The legal representative is xxx. Finally, combine the outputs of the two sub-tasks corresponding to the first two target sub-problems with the third target sub-problem to form a new problem: Given [Query the merchant with the highest transaction amount in the A province branch, the merchant name is Merchant B; query the merchant's legal representative, the merchant's legal representative is xxx]. Finally, query the merchant's network access date. This would call the network access date query tool, which returns the answer: The network access date is: March 19, 2021.
[0103] Next, the results of the three subtasks [find the merchant with the highest transaction amount in the A province branch, the merchant name is Merchant B; find the legal representative of the merchant, the legal representative is xxx; find the merchant's network access date, the network access date is March 19, 2021] and the question "Help me find the merchant with the highest transaction amount in the A province branch, who is its legal representative, and when did it join the network." are integrated into the prompt input of a large preset language model (such as ChatGLM3). The large language model returns the answer as: After querying, the merchant with the highest transaction amount in the A province branch is Merchant B, the legal representative is xxx, and the network access date is March 19, 2021.
[0104] To facilitate understanding of the intelligent agent question-answering method provided in this application, the following embodiment will introduce the overall structure of the intelligent agent question-answering system mentioned in this application.
[0105] For intelligent agent question-answering systems based on open-source large language models, users can interact with the system through text, voice, and other question-and-answer methods. The system decomposes the complex tasks corresponding to the user's questions by calling a pre-defined large language model and then invokes appropriate tools to complete each sub-task. The tool functions supported by the intelligent agent question-answering system can include, but are not limited to, at least one of the following: general-purpose tool functions (such as weather query and text-to-image tool functions), internal system call tool functions of the target enterprise (such as employee holiday day query tool functions), and question-and-answer tool functions of the target enterprise's internal professional knowledge base. These tool functions can be existing, mature, general-purpose tool functions, or personalized tool functions that are user-defined and configured via a webpage.
[0106] Specifically, such as Figure 4 As shown, the preset large language model can be the ChatGLM3 large language model, which is used to decompose the task corresponding to the user's question, select tools in the Function Call unit, and understand and integrate the execution results.
[0107] The intelligent agent question-answering system is used to receive user question commands, coordinate and call a pre-set large language model to decompose and solve the question. The core of the system is to combine ReAct and FunctionCall for question-answering processing.
[0108] The tool management system includes general tools and user-defined tool functions displayed in a web-based interface. These functions include their functionality, parameters, and other information, which are used to automatically convert the configuration information into the required code when the tool list is loaded. The system also configures the scope of tools needed by each business function. When defining a user-defined tool function, the system needs to specify the function's name, description, parameters, parameter descriptions, whether parameters are required, and the specific implementation interface of the function.
[0109] For example: If a user needs to customize and add a weather query tool function, the configuration process on the page can be as follows: (1) The user needs to provide an English name for this tool function, such as weather_search; (2) Add a description to the tool function so that the large language model can understand the function's function, such as "This tool function is used to query the weather of a city"; (3) Input the parameters of the tool function, such as the parameter "city" which is the English name of this tool function; (4) Input the meaning of the parameter, such as "the name of the city to be queried", and specify whether the parameter is required, such as the parameter is required; (5) The user also needs to input the specific interface for weather query, such as https: / / wttr.in / {city_name}?format=j1; (6) The user clicks to generate the tool, and the system will generate the tool registration code based on the user's configuration information and insert it into the system code for execution using the code executor.
[0110] A code executor (such as a Python interpreter) is a runtime environment (such as a Python runtime environment) for a specified type of code language, used to execute user-defined tools as functions of a specified type (such as Python) and return results.
[0111] The tool implementation service / system provides the specific capabilities for the tool functions needed to execute the subtasks corresponding to each target subproblem. This part can be horizontally expanded as needed. Common implementation methods include the following:
[0112] Weather query API: A free API interface for the internet;
[0113] The Wenshengtu interface: draws corresponding images based on user descriptions;
[0114] Human Resources System: Query human resources-related information, including remaining vacation days, issuance of employment certificates, attendance information, etc.
[0115] Operations Service Desk: Provides specific business inquiries such as sub-merchant number inquiry and centralized payment inquiry;
[0116] The professional knowledge base Q&A interface supports users uploading business documents. The system's answers to users will refer to these uploaded documents, ensuring the accuracy of the responses. In addition, it also supports general free-form Q&A.
[0117] In addition, the historical question-and-answer data and tool function call data stored in the agent question-answering system can be used to fine-tune the preset large language model (such as ChatGLM3), thus forming a virtuous cycle of model performance and question-and-answer behavior, thereby improving the accuracy of the agent question-answering system's answers to user questions.
[0118] In summary, the agent-based question-answering method provided in this embodiment first receives a question instruction from a target user to the agent-based question-answering system, determines the target question to be answered based on the question instruction, and then calls a preset large language model to decompose the target question into N target sub-questions; where N is a positive integer greater than 0. Next, while iterating through the N target sub-questions, the answer results of the i-th target sub-question and the (i-1)-th target sub-question are input into the function call unit of the agent-based question-answering system for solving, obtaining the answer result of the i-th target sub-question, and so on, until the answer results corresponding to each of the N target sub-questions are obtained; where i is a positive integer greater than 2 and not greater than N. Finally, the answer results corresponding to each of the N target sub-questions and the text corresponding to the target question are input into the preset large language model to obtain the answer content for the target question output by the large language model, and then displayed to the target user through the agent-based question-answering system.
[0119] As can be seen, this application combines reasoning action (ReAct) and function call, using the function call unit of the intelligent agent question answering system as the smallest implementation unit, and combining the three steps of thinking, execution and observation of ReAct, to achieve accurate answers to the target question posed by the target user containing N target sub-questions, thereby improving the question answering experience of the target user.
[0120] Second Embodiment
[0121] This embodiment will introduce an intelligent agent question-answering device; please refer to the above method embodiment for related content.
[0122] See Figure 5 This is a schematic diagram of the composition of an intelligent agent question-answering device provided in this embodiment. The device 500 includes:
[0123] The determining unit 501 is used to receive a question instruction sent by a target user to the intelligent agent question-answering system, and determine the target question to be answered based on the question instruction;
[0124] The decomposition unit 502 is used to call a preset large language model to decompose the target problem into N target sub-problems; where N is a positive integer greater than 0.
[0125] The traversal unit 503 is used to input the answer results of the i-th target sub-problem and the (i-1)-th target sub-problem into the function call unit of the intelligent agent question answering system for solving when iterating through the N target sub-problems, so as to obtain the answer result of the i-th target sub-problem, and so on, until the answer results corresponding to each of the N target sub-problems are obtained; where i is a positive integer greater than 2 and not greater than N;
[0126] The obtaining unit 504 is used to input the answer results corresponding to each of the N target sub-questions and the text corresponding to the target question into the preset large language model, obtain the answer content for the target question output by the large language model, and display it to the target user through the intelligent agent question answering system.
[0127] In one implementation of this embodiment, the traversal unit 503 is specifically used for:
[0128] After inputting the answers to the i-th target sub-problem and the (i-1)-th target sub-problem into the function call unit of the intelligent agent question answering system, the i-th target sub-problem is solved by calling the corresponding tool function and the parameters required by the tool function, and the answer to the i-th target sub-problem is obtained.
[0129] In one implementation of this embodiment, the parameters required by the tool function are extracted from the answers to the i-th target sub-problem and the (i-1)-th target sub-problem.
[0130] In one implementation of this embodiment, the parameters required by the tool function are obtained through interaction with the target user.
[0131] In one implementation of this embodiment, the tool function includes at least one of a general tool function, a target enterprise internal system call tool function, and a target enterprise internal professional knowledge base question and answer tool function.
[0132] In one implementation of this embodiment, the utility function is generated by user-defined configuration in the form of a page.
[0133] In one implementation of this embodiment, the traversal unit 503 includes:
[0134] The integration subunit is used to input the answer results of the i-th target sub-question and the (i-1)-th target sub-question into the function call unit of the intelligent agent question answering system, and then call the preset large language model to integrate the answer results of the i-th target sub-question and the (i-1)-th target sub-question to obtain the updated i-th target sub-question;
[0135] The first judgment subunit is used to determine whether the completion flag bit in the agent history table is a completion flag by searching the historical information of each target sub-problem recorded in the pre-built agent history table;
[0136] The reading sub-unit is used to set the historical information of the i-th target sub-question to empty if, in the case of yes, or no question-and-answer records under the session corresponding to the target question; otherwise, if, question-and-answer records exist under the session corresponding to the target question, the historical information of the i-th target sub-question is read.
[0137] The second judgment subunit is used to obtain the parameters required by the called tool function based on the updated i-th target sub-question and the historical information of the i-th target sub-question, and to determine whether the called tool function is a question-answering tool function;
[0138] The first solving subunit is used to, if yes, utilize the parameters required by the question-answering tool function to call a preset question-answering tool in conjunction with a large language model to solve the updated i-th target sub-question, obtain the answer result of the i-th target sub-question; and update the historical information of the i-th target sub-question, and set the completion flag in the agent's history table to a completion flag;
[0139] The third judgment subunit is used to determine whether the called utility function is a utility function that requires user information authentication if the condition is not met.
[0140] The verification subunit is used to verify the user information contained in the target question when it is determined that the called utility function is a utility function that requires user information authentication.
[0141] The second solution subunit is used to update the historical information of the i-th target sub-problem when the verification passes, determine the tool function to be called based on the updated historical information of the i-th target sub-problem, so as to solve the updated i-th target sub-problem and obtain the answer result of the i-th target sub-problem; and update the historical information of the i-th target sub-problem, and set the completion flag in the agent history table to the completion flag.
[0142] Furthermore, embodiments of this application also provide an intelligent agent question-answering device, including: a processor, a memory, and a system bus;
[0143] The processor and the memory are connected via the system bus;
[0144] The memory is used to store one or more programs, the one or more programs including instructions, which, when executed by the processor, cause the processor to perform any of the above-described implementations of the intelligent agent question-answering method.
[0145] Furthermore, embodiments of this application also provide a computer-readable storage medium storing instructions that, when executed on a terminal device, cause the terminal device to execute any of the above-described implementations of the intelligent agent question-answering method.
[0146] Furthermore, this application embodiment also provides a computer program product, which, when run on a terminal device, causes the terminal device to execute any of the above-described implementation methods of the intelligent agent question-answering method.
[0147] As can be seen from the above description of the embodiments, those skilled in the art can clearly understand that all or part of the steps in the methods of the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to cause a computer device (which may be a personal computer, a server, or a network communication device such as a media gateway, etc.) to execute the methods described in various embodiments or some parts of the embodiments of this application.
[0148] It should be noted that the various embodiments in this specification are described in a progressive manner, with each embodiment focusing on the differences from other embodiments. Similar or identical parts between embodiments can be referred to interchangeably. For the apparatus disclosed in the embodiments, since it corresponds to the method disclosed in the embodiments, the description is relatively simple; relevant parts can be referred to in the method section.
[0149] It should also be noted that, in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.
[0150] The above description of the disclosed embodiments enables those skilled in the art to make or use this application. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of this application. Therefore, this application is not to be limited to the embodiments shown herein, but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.
Claims
1. A question-answering method for intelligent agents, characterized in that, include: Receive a question instruction from a target user to the intelligent agent question-answering system, and determine the target question to be answered based on the question instruction; The target problem is decomposed by calling a preset large language model to obtain N target sub-problems contained in the target problem; where N is a positive integer greater than 0. When iterating through the N target sub-problems, the answers to the i-th and (i-1)-th target sub-problems are input into the function call unit of the intelligent agent question-answering system. Then, by calling the corresponding tool function and the parameters required by the tool function, the i-th target sub-problem is solved to obtain the answer to the i-th target sub-problem. This process is repeated until the answer to each of the N target sub-problems is obtained; where i is a positive integer greater than 1 and not greater than N. After inputting the answers to the i-th and (i-1)-th target sub-questions into the function call unit of the intelligent agent question-answering system, the i-th target sub-question is solved by calling the corresponding tool function and the parameters required by the tool function, and the answer to the i-th target sub-question is obtained, including: After inputting the answers to the i-th target sub-question and the (i-1)-th target sub-question into the function call unit of the intelligent agent question answering system, the preset large language model is called to integrate the answers to the i-th target sub-question and the (i-1)-th target sub-question to obtain the updated i-th target sub-question; By searching the historical information of each target sub-problem recorded in the pre-built agent history table, it is determined whether the completion flag bit in the agent history table is a completion flag; If yes, or if there are no question-and-answer records in the session corresponding to the target question, then the historical information of the i-th target sub-question is set to empty; if no, or if there are question-and-answer records in the session corresponding to the target question, then the historical information of the i-th target sub-question is read. Based on the updated i-th target sub-question and the historical information of the i-th target sub-question, obtain the parameters required by the called tool function, and determine whether the called tool function is a question-answering tool function; If so, then using the parameters required by the question-answering tool function, the preset question-answering tool is called in conjunction with the large language model to solve the updated i-th target sub-question, and the answer result of the i-th target sub-question is obtained; and the historical information of the i-th target sub-question is updated, and the completion flag in the agent's history table is set to the completion flag; If not, then determine whether the called utility function is a utility function that requires user information authentication; When it is determined that the called utility function is a utility function that requires user information authentication, the user information contained in the target question is verified. Upon successful verification, the historical information of the i-th target sub-problem is updated. Based on the updated historical information of the i-th target sub-problem, the utility function to be called is determined to solve the updated i-th target sub-problem and obtain the answer result of the i-th target sub-problem. The historical information of the i-th target sub-problem is also updated, and the completion flag in the agent's history table is set to the completion flag. The answers to the N target sub-questions and the text corresponding to the target question are input into the preset large language model to obtain the answer content for the target question output by the large language model, and then displayed to the target user through the intelligent agent question answering system.
2. The method according to claim 1, characterized in that, The parameters required by the tool function are extracted from the answers to the i-th and (i-1)-th target sub-problems.
3. The method according to claim 1, characterized in that, The parameters required by the tool function are obtained through interaction with the target user.
4. The method according to any one of claims 1-3, characterized in that, The utility functions include at least one of the following: general utility functions, internal system call utility functions of the target enterprise, and question-and-answer utility functions of the target enterprise's internal professional knowledge base.
5. The method according to any one of claims 1-3, characterized in that, The utility functions are generated by user-defined configurations in a page-based manner.
6. A smart agent question-answering device, characterized in that, include: The determining unit is used to receive a question instruction sent by a target user to the intelligent agent question-answering system, and determine the target question to be answered based on the question instruction; The decomposition unit is used to call a preset large language model to decompose the target problem into N target sub-problems; where N is a positive integer greater than 0. The traversal unit is used to input the answer results of the i-th target sub-problem and the (i-1)-th target sub-problem into the function call unit of the intelligent agent question answering system for solving when iterating through the N target sub-problems, so as to obtain the answer result of the i-th target sub-problem, and so on, until the answer results corresponding to each of the N target sub-problems are obtained; where i is a positive integer greater than 2 and not greater than N; The obtaining unit is used to input the answer results corresponding to each of the N target sub-questions and the text corresponding to the target question into the preset large language model, obtain the answer content for the target question output by the large language model, and display it to the target user through the intelligent agent question answering system.
7. An intelligent agent question-answering device, characterized in that, include: Processor, memory, system bus; The processor and the memory are connected via the system bus; The memory is used to store one or more programs, the one or more programs including instructions that, when executed by the processor, cause the processor to perform the method according to any one of claims 1-5.
8. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores instructions that, when executed on a terminal device, cause the terminal device to perform the method described in any one of claims 1-5.
Citation Information
Patent Citations
Question and answer method and device, electronic equipment and medium
CN117521625A