Intelligent agent question and answer method and device fusing different large models, medium and product
By integrating the agent's question-and-answer method of different large models, using the first large model to plan the task and performing external tool calls to the second large model, the agent's lack of understanding when facing fuzzy or complex problems is solved, and faster and more accurate answers are achieved, improving the user experience.
Patent Information
- Application Number
- CN202510849838.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-23
- Publication Date
- 2025-07-22
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
When an agent faces vague or complex problems, it is difficult for him to accurately understand the user's intentions, resulting in slow output responses or inability to meet user needs, affecting the Q&A experience.
By integrating agent question and answer methods of different large models, the first large model is used to plan tasks, generate target processing tasks, and call the second large model to execute external tool calls to generate answers. The first big model has a stronger inference ability than the second big model, and is used to more accurately understand user intentions, and the second big model responds quickly.
It improves the execution efficiency of the agent's Q&A process, reduces user waiting time, improves user's Q&A experience, and ensures that the output answers are more accurate and reasonable.
Smart Images

Figure CN120353909A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the fields of large models, agents, and artificial intelligence. Specifically, it relates to an agent question-answering method, device, medium, and product that integrates different large models. Background Art
[0002] An agent refers to a system or program that can autonomously execute tasks on behalf of a user or other system by designing its workflow and utilizing available tools. With the rapid development of artificial intelligence technology, agents are increasingly widely used in daily life and have gradually become a useful assistant in people's lives. For example, in a question-answering scenario, a user can input a question to the agent in natural language form, and the agent can output a corresponding reply based on the associated large model.
[0003] However, the capabilities of agents are limited in specific scenarios. For example, when faced with ambiguous questions, it is often difficult to understand the user's intent, resulting in slow output of reply content or the output reply content not meeting the user's needs, thereby affecting the user's question-answering experience. Summary of the Invention
[0004] This Summary of the Invention section is provided to introduce concepts in a brief form, which will be described in detail in the following Detailed Description section. This Summary of the Invention section is not intended to identify key features or essential features of the claimed technical solution, nor is it intended to be used to limit the scope of the claimed technical solution.
[0005] In a first aspect, the present disclosure provides an agent question-answering method that integrates different large models. The agent question-answering method that integrates different large models includes: In response to obtaining a question input by a user to the agent, call a first large model associated with the agent to perform task planning based on the question, obtain a target processing task for answering the question, and send the target processing task to a second large model associated with the agent, where the reasoning ability of the first large model is stronger than that of the second large model, and the target processing task includes steps of calling external tools associated with the agent; Call the second large model to execute the target processing task to obtain external knowledge required to answer the question, where the external knowledge represents the knowledge obtained by the second large model by calling the external tools; Generate an answer to the question based on the question and the external knowledge.
[0006] In a second aspect, the present disclosure provides an agent question-answering device that integrates different large models. The agent question-answering device that integrates different large models includes: A first processing module, configured to, in response to obtaining a question input by a user to an intelligent agent, call a first large model associated with the intelligent agent to perform task planning based on the question, obtain a target processing task for answering the question, and send the target processing task to a second large model associated with the intelligent agent, where the reasoning ability of the first large model is stronger than that of the second large model, and the target processing task includes a step of calling an external tool associated with the intelligent agent; A second processing module, configured to call the second large model to execute the target processing task to obtain external knowledge required to answer the question, where the external knowledge represents knowledge obtained by the second large model calling the external tool; A generation module, configured to generate an answer to the question based on the question and the external knowledge.
[0007] In a third aspect, the present disclosure provides a computer-readable medium, on which a computer program is stored, and when the computer program is executed by a processing device, the steps of the method described in the first aspect are implemented.
[0008] In a fourth aspect, the present disclosure provides an electronic device, including: A storage device, on which a computer program is stored; A processing device, configured to execute the computer program in the storage device to implement the steps of the method described in the first aspect.
[0009] In a fifth aspect, the present disclosure provides a computer program product, including a computer program, and when the computer program is executed by a processor, the steps of the method described in the first aspect are implemented.
[0010] Through the above technical solution, the first large model associated with the intelligent agent can perform task planning on the obtained question to obtain a target processing task for answering the question, and the second large model can be called to execute the target processing task to obtain external knowledge required for answering the question. Additionally, based on the question and the external knowledge, an answer to the question can be generated. Among them, the reasoning ability of the first large model is stronger than that of the second large model. Therefore, through the first large model, the question can be analyzed and logically deduced more deeply, thereby understanding the user's intention more accurately, and then planning a more accurate and reasonable target processing task for answering the user's question, enabling the intelligent agent to output an accurate and reasonable answer based on the target processing task. Moreover, since the reasoning ability of the second large model is weaker than that of the first large model, when the second large model executes the target processing task, it generally does not think deeply about the target processing task, thereby being able to improve the response speed to the target processing task, and further improving the execution efficiency of the intelligent agent's question-and-answer process and reducing the user's waiting time. Thus, by combining the first large model and the second large model, not only can the user's intention be understood more accurately, but also the user's waiting time can be reduced, thereby better meeting the user's needs and enhancing the user's question-and-answer experience.
[0011] Other features and advantages of the present disclosure will be described in detail in the subsequent specific implementation section. BRIEF DESCRIPTION OF THE DRAWINGS
[0012] In combination with the accompanying drawings and with reference to the following specific implementation, the above and other features, advantages, and aspects of the various embodiments of the present disclosure will become more apparent. Throughout the drawings, the same or similar reference numerals represent the same or similar elements. It should be understood that the drawings are schematic, and the original components and elements are not necessarily drawn to scale. In the drawings: Figure 1 Shows a flowchart of question-and-answer based on an intelligent agent in the related art; Figure 2 Is a flowchart of an intelligent agent question-and-answer method that integrates different large models according to an exemplary embodiment of the present disclosure; Figure 3 Is a schematic diagram of obtaining a question input by a user to the intelligent agent according to an exemplary embodiment of the present disclosure; Figure 4 Is a flowchart block diagram of an intelligent agent question-and-answer method that integrates different large models according to an exemplary embodiment of the present disclosure; Figure 5 Is a schematic diagram of obtaining feedback information on a processing task from a user according to an exemplary embodiment of the present disclosure; Figure 6 Is a schematic diagram of obtaining feedback information on a processing task from a user according to another exemplary embodiment of the present disclosure; Figure 7 It is a structural block diagram of an intelligent agent question - answering device that integrates different large - language models shown according to an exemplary embodiment of the present disclosure; Figure 8 It is a schematic structural diagram of an electronic device shown according to an exemplary embodiment of the present disclosure. Detailed implementation manners
[0013] Embodiments of the present disclosure will be described in more detail below with reference to the accompanying drawings. Although some embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. On the contrary, these embodiments are provided to more thoroughly and completely understand the present disclosure. It should be understood that the drawings and embodiments of the present disclosure are only for exemplary purposes and are not used to limit the protection scope of the present disclosure.
[0014] It should be understood that the steps recorded in the method embodiments of the present disclosure can be executed in different orders and / or in parallel. In addition, the method embodiments may include additional steps and / or omit the steps shown. The scope of the present disclosure is not limited in this regard.
[0015] The term "including" and its variants used herein are open - ended, that is, "including but not limited to". The term "based on" means "at least partially based on". The term "one embodiment" means "at least one embodiment"; the term "another embodiment" means "at least one additional embodiment"; the term "some embodiments" means "at least some embodiments". The relevant definitions of other terms will be given in the following description.
[0016] It should be noted that the concepts such as "first" and "second" mentioned in the present disclosure are only used to distinguish different devices, modules or units, and are not used to limit the order of functions executed by these devices, modules or units or their interdependent relationships.
[0017] It should be noted that the modifications of "one" and "multiple" mentioned in the present disclosure are illustrative rather than restrictive. Those skilled in the art should understand that, unless otherwise clearly indicated in the context, it should be understood as "one or more".
[0018] The names of the messages or information exchanged between multiple devices in the embodiments of the present disclosure are only for illustrative purposes and are not used to limit the scope of these messages or information.
[0019] It can be understood that before using the technical solutions disclosed in the embodiments of the present disclosure, the types, usage scopes, usage scenarios, etc. of the personal information involved in the present disclosure should be informed to users and user authorization should be obtained in an appropriate manner in accordance with relevant laws and regulations.
[0020] For example, when responding to an active request from a user, a prompt message is sent to the user to clearly prompt the user that the operation requested by the user will require obtaining and using the user's personal information. Thus, the user can autonomously choose whether to provide personal information to software or hardware such as an electronic device, application program, server, or storage medium that performs the operations of the present disclosure's technical solution based on the prompt message.
[0021] As an optional but non-limiting implementation manner, the way of sending a prompt message to the user in response to receiving an active request from the user can be, for example, in the form of a pop-up window. The prompt message can be presented in text in the pop-up window. In addition, the pop-up window can also carry selection controls for the user to choose "agree" or "disagree" to provide personal information to the electronic device.
[0022] It can be understood that the above process of notifying and obtaining user authorization is only illustrative and does not limit the implementation manner of the present disclosure. Other ways that comply with relevant laws and regulations can also be applied to the implementation manner of the present disclosure.
[0023] At the same time, it can be understood that the data involved in the present technical solution (including but not limited to the data itself, the acquisition or use of the data) should comply with the requirements of corresponding laws, regulations and related provisions.
[0024] It should be understood that an intelligent agent usually includes a task planning and execution stage and a thinking and summarizing stage when working, as Figure 1 shown. In the task planning and execution stage, the intelligent agent mainly performs intent recognition and tool invocation through an associated large model to obtain additional context information. Specifically, the large model can select a suitable tool from multiple associated external tools according to the prompt words and the questions input by the user, and obtain knowledge from them. In the thinking and summarizing stage, the intelligent agent mainly generates a final answer through the associated large model based on the prompt words, the additional context obtained in the task planning and execution stage, and the questions input by the user. Specifically, it will summarize by combining the obtained knowledge with the prompt words and use the summary result as the final response to the user's question.
[0025] Among them, the large model associated with the intelligent agent is generally one of a general large model and an inference large model. When the large model associated with the intelligent agent is a general large model, since the general large model is a traditional large model without an explicit multi-step reasoning process and is more focused on directly generating answers according to existing patterns, the response speed is fast. However, the general large model has limited capabilities in dealing with complex tasks or ambiguous tasks, which leads to poor final solution effects. For example, when dealing with complex tasks, complex tasks usually need to be solved in multiple steps. The general large model sometimes cannot correctly plan multiple steps or has poor effects when executing specific steps. Another example is that when dealing with ambiguous tasks, the general large model often has difficulty understanding the user's intent.
[0026] When the large model associated with the agent is an inference large model, since the inference large model often requires in-depth reasoning and multi-dimensional analysis, the accuracy of task planning is better. However, because the inference large model needs to think deeply every time it calls an external tool, the time-consuming for generating and processing tasks is longer, and users need to spend more time waiting for the large model to output the response content, thus affecting the user's Q&A experience.
[0027] In summary, during the task planning and execution stage of the agent, the general large model has limited ability to recognize the user's intention, which may lead to slow output of responses or inability to meet the user's needs, affecting the Q&A experience; although the inference large model has high accuracy, it takes a long time, which also affects the user experience.
[0028] In view of this, the present disclosure provides an intelligent agent Q&A method, device, medium and product that integrates different large models to solve the above technical problems.
[0029] The following further explains the embodiments of the present disclosure with reference to the accompanying drawings.
[0030] Figure 2 is a flowchart of an intelligent agent Q&A method that integrates different large models shown according to an exemplary embodiment of the present disclosure. Referring to Figure 2 , the intelligent agent Q&A method that integrates different large models may include the following steps: S201: In response to obtaining a question input by the user to the agent, call the first large model associated with the agent to perform task planning based on the question, obtain a target processing task for answering the question, and send the target processing task to the second large model associated with the agent, where the inference ability of the first large model is stronger than that of the second large model, and the target processing task includes steps for calling external tools associated with the agent.
[0031] It should be understood that an agent refers to a system, hardware or program that can autonomously execute tasks on behalf of a user or other system by designing its workflow and using available tools. Thus, the agent involved in the present disclosure may be a server deployed with a large model, or a terminal device deployed with a large model, and of course, it may also be others. The embodiments of the present disclosure do not make any restrictions on this.
[0032] Exemplarily, when the intelligent agent is a server deployed with a large model, the problem of obtaining user input to the intelligent agent can be: displaying an intelligent interaction page for the user to input questions in the terminal device that has established a communication connection with the intelligent agent. The intelligent interaction page may include an information input box and an information display box. Among them, a "Send" control for sending the question input by the user to the intelligent agent is displayed in the information input box. Thus, when the user inputs a question in the information input box, for example, after inputting "Based on a certain hot topic today, send a relevant article to content platform a and content platform b", the user can trigger the "Send" control through operations such as clicking, and send "Based on a certain hot topic today, send a relevant article to content platform a and content platform b" to the intelligent agent and display it in the information display box, thereby obtaining the question input by the user to the intelligent agent, such as Figure 3 as shown
[0033] After the intelligent agent receives "Based on a certain hot topic today, send a relevant article to content platform a and content platform b", it can call the first large model associated with the intelligent agent to perform task planning based on this question, obtain the target processing task for answering the question, and send the target processing task to the second large model associated with the intelligent agent.
[0034] It should be understood that due to the diverse nature of the questions themselves, the questions input by users to the intelligent agent may be complex questions, ambiguous questions, or simple questions. For simple questions, the second large model with relatively weak reasoning ability can generally understand the user's intention more accurately and quickly output the corresponding response content. Therefore, in order to further improve the execution efficiency of the intelligent agent's question-and-answer process and reduce the user's waiting time. In a possible way, after obtaining the question input by the user to the intelligent agent, the category of the question can be judged first. If the question is a complex question or an ambiguous question, the first large model associated with the intelligent agent is called to perform task planning based on the question; if the question is not a complex question or an ambiguous question, the second large model associated with the intelligent agent is called to perform task planning based on the question, thereby reducing the inference time of the large model. That is to say, in a possible way, in response to obtaining the question input by the user to the intelligent agent, calling the first large model associated with the intelligent agent to perform task planning based on the question may include: In response to obtaining the question input by the user to the intelligent agent, calling the second large model associated with the intelligent agent to classify the question to obtain the question category corresponding to the question; in the case where the question category indicates that the question belongs to an ambiguous question or a complex question, calling the first large model associated with the intelligent agent to perform task planning based on the question.
[0035] Exemplarily, different categories of problems can be predefined. For example, ambiguous problems, complex problems, and simple problems can be predefined in advance, and then a first prompt word template can be constructed based on the definition of the problems. For example, the constructed first prompt word template can be: You are a professional classifier. Based on the following definitions: "Ambiguous problem: AAAAA; Complex problem: BBBBB; Simple problem: CCCCCC", classify the user input problem: "DDDDDD".
[0036] Thus, after obtaining the problem input by the user to the intelligent agent, the problem can be filled into the corresponding position of the first prompt word template to obtain the first prompt word, and the first prompt word is input into the second large model to obtain the problem category corresponding to the problem. If the problem category indicates that the problem belongs to a simple problem, the second large model associated with the intelligent agent is called to perform task planning based on the problem; if the problem category indicates that the problem belongs to an ambiguous problem or a complex problem, the first large model associated with the intelligent agent is called to perform task planning based on the problem, as Figure 4 shown.
[0037] S202: Call the second large model to execute the target processing task to obtain the external knowledge required to answer the question, where the external knowledge represents the knowledge obtained by the second large model by calling external tools.
[0038] It should be understood that the target processing task can include one sub-processing task or multiple sub-processing tasks with a preset execution order. When the target processing task includes multiple sub-processing tasks with a preset execution order, in view of the fact that the next sub-processing task may depend on the external knowledge obtained by the previous sub-processing task when obtaining external knowledge, in possible ways, in order to improve the accuracy of the external knowledge, when calling the second large model to execute the target processing task to obtain the external knowledge required to answer the question, the second large model can be called to execute multiple sub-processing tasks in a preset execution order to obtain multiple external knowledge required to answer the question. That is to say, in possible ways, the target processing task can include multiple sub-processing tasks, and the multiple sub-processing tasks have a preset execution order. Correspondingly, calling the second large model to execute the target processing task to obtain the external knowledge required to answer the question can include: Call the second large model to execute multiple sub-processing tasks in a preset execution order to obtain multiple external knowledge required to answer the question; Correspondingly, generating an answer to the question based on the question and the external knowledge can include: Integrate the multiple external knowledge obtained through multiple sub-processing tasks into the target external knowledge, and fill the question and the target external knowledge into the preset prompt word template to obtain the target prompt word; call the first large model through the target prompt word to generate an answer to the question.
[0039] In this embodiment, the preset prompt word template can be determined according to the actual situation, and the present disclosure does not impose any restrictions on this. By way of example, the preset prompt word template can be: You are an intelligent summarization assistant responsible for summarizing based on the questions input by the user and the external knowledge obtained.
[0040] The problem to be solved is: XXXXXX; The external knowledge obtained is: TTTTTTTT; The summarization requirement is: NNNNNNN.
[0041] Thus, by filling the problem and the target external knowledge into the corresponding positions of the preset prompt word template, the target prompt word can be obtained. After inputting the target prompt word into the first large model, on the one hand, the first large model can generate an answer to the problem according to the summarization requirement; on the other hand, since the first large model has reasoning ability, when generating an answer based on the problem and external knowledge through the first large model, it can conduct in-depth analysis and thinking based on the problem and external knowledge, thereby improving the richness of the answer and further enhancing the user experience.
[0042] S203: Generate an answer to the question based on the question and external knowledge.
[0043] By way of example, continuing to refer to Figure 3 As shown, after generating an answer to the question based on the question and external knowledge, the answer can be displayed on the intelligent interaction page.
[0044] Through the above technical solution, the first large model associated with the intelligent agent can perform task planning on the obtained question to obtain the target processing task for answering the question, and can call the second large model to execute the target processing task to obtain the external knowledge required to answer the question; it can also generate an answer to the question based on the question and external knowledge. Among them, the reasoning ability of the first large model is stronger than that of the second large model. Thus, through the first large model, the question can be analyzed and logically deduced more deeply, so as to more accurately understand the user's intention, and then plan a more accurate and reasonable target processing task for answering the user's question, enabling the intelligent agent to output an accurate and reasonable answer based on this target processing task. In addition, since the reasoning ability of the second large model is weaker than that of the first large model, when the second large model executes the target processing task, it generally does not conduct in-depth thinking on the target processing task, thereby being able to improve the response speed to the target processing task, and then improving the execution efficiency of the intelligent agent's question-and-answer process and reducing the user's waiting time. Thus, by combining the first large model and the second large model, not only can the user's intention be understood more accurately, but also the user's waiting time can be reduced, so as to better meet the user's needs and enhance the user's question-and-answer experience.
[0045] To facilitate the understanding of the intelligent agent question - answering method that integrates different large - language models provided by this disclosure, the possible implementation manners in this disclosure are described below.
[0046] In a possible manner, the first large - language model associated with the intelligent agent is called to perform task planning based on the question, and an objective processing task for answering the question can be obtained, including: Call the first large - language model associated with the intelligent agent to perform the following task - planning process: According to the intention recognition result of the question, break the question into multiple sub - questions; according to the dependency relationship between the multiple sub - questions, determine the execution order between the subtasks corresponding to the multiple sub - questions, where one sub - question corresponds to one subtask; for each subtask, generate the execution steps required for the subtask, where the execution steps include steps for calling external tools and / or steps for calling the second large - language model; generate an objective processing task for answering the question at least based on the execution order between the multiple subtasks and the execution steps of the multiple subtasks.
[0047] Exemplarily, a first prompt - word template can be preset in advance. Thus, after obtaining the question input by the user to the intelligent agent, the question can be filled into the corresponding position of the first prompt - word template to obtain a first prompt word, and the first prompt word is input into the first large - language model to obtain an objective processing task for answering the question. For example, the first prompt - word template can be: You are an intelligent planning assistant responsible for generating processing tasks to solve the user's question.
[0048] You have the following external tools. Try to use these external tools to generate corresponding processing tasks, and each processing task includes one external - tool call.
[0049] Internet search tool: Search the Internet according to keywords; Content platform a publishing tool: Publish an article to content platform a; Content platform b publishing tool: Publish an article to content platform b.
[0050] The problem to be solved is: XXXXX.
[0051] Thus, after obtaining the question "Based on a certain current hot topic, publish a relevant article to content platform a and content platform b" input by the user to the intelligent agent, "Based on a certain current hot topic, publish a relevant article to content platform a and content platform b" can be filled into the first prompt - word template to obtain the following first prompt word: You are an intelligent planning assistant responsible for generating processing tasks to solve the user's question.
[0052] You have the following external tools. Try to use these external tools to generate corresponding processing tasks, and each processing task includes one call to an external tool.
[0053] Online Search Tool: Search the Internet based on keywords; Content Platform A Publishing Tool: Publish articles to Content Platform A; Content Platform B Publishing Tool: Publish articles to Content Platform B.
[0054] The problem to be solved is: Based on a certain hot topic today, publish a relevant article to Content Platform A and Content Platform B.
[0055] After that, the first prompt can be input into the first large model. The first large model can identify the user's intention based on the first prompt and obtain the intention of "the user hopes to publish articles on two different content platforms by leveraging hot events". After clarifying the user's intention, the statement "Based on a certain hot topic today, publish a relevant article to Content Platform A and Content Platform B" can be disassembled into the following sub-questions: Sub-question a: Selection of hot events; Sub-question b: Obtaining detailed information about the hot event; Sub-question c: Creation and publication of article content.
[0056] Since Sub-question c depends on Sub-question b, Sub-question b depends on Sub-question a, and the sub-tasks corresponding to Sub-question a and Sub-question b need to call the "Online Search Tool" during the execution process, and the sub-task corresponding to Sub-question c needs to call the "Content Platform A Publishing Tool and Content Platform B Publishing Tool" during the execution process. Therefore, according to the execution order and execution steps of multiple sub-tasks, the following target processing tasks can be input: Sub-task a: Call the "Online Search Tool" and use the keyword "Today's Hot Search List" to search for the current hottest topic; Reason: First, determine the most discussed hot event of the day to ensure the timeliness and attention of the selected topic.
[0057] Sub-task b: Call the "Online Search Tool" to conduct in-depth searches using the obtained hot keywords; Reason: Collect materials such as event details, comments from all parties, and relevant data to prepare the article content.
[0058] Sub-task c: Call the "Content Platform A Publishing Tool" to create content that conforms to the platform's tone, and call the "Content Platform B Publishing Tool" to create content that conforms to the platform's tone.
[0059] It should be understood that when the first large model generates a target processing task based on a question, due to possible deviations in the understanding of the question or limitations in the logical reasoning process, the generated target processing task may be incomplete or unreasonable. If the second large model directly obtains the external knowledge required to answer the question based on the incomplete or unreasonable target processing task, there may be a situation where the obtained external knowledge is inaccurate, resulting in inaccurate or unreasonable answers when the intelligent agent generates an answer to the question based on the external knowledge, thereby affecting the user experience. Therefore, in order to further improve the accuracy and reasonableness of the answers output by the intelligent agent, among possible methods, after the first large model outputs the target processing task, it can first feedback the target processing task to the user for confirmation. After the user confirms that it is correct, then call the second large model to execute the target processing task. When the user confirms that there is a problem, the target processing task can be dynamically adjusted based on the user's feedback information, and after the dynamic adjustment, it is fed back to the user for confirmation again until the user confirms that it is correct and then call the second large model to execute the target processing task. That is to say, among possible methods, calling the first large model associated with the intelligent agent to perform task planning based on the question to obtain the target processing task for answering the question may include: Call the first large model associated with the intelligent agent to perform task planning based on the question to obtain an initial processing task for answering the question, and loop through the following process: Display the processing task to the user; in response to obtaining the user's feedback information on the processing task, parse the feedback information to obtain a task evaluation result. If the task evaluation result indicates that the processing task does not include all the complete task steps required to answer the question, adjust the processing task according to the task evaluation result to obtain a new processing task, and return to the step of displaying the processing task to the user until the task evaluation result indicates that the processing task includes all the complete task steps; use the processing task after the loop ends as the target processing task for answering the question.
[0060] Exemplarily, continuing to refer to the above example, if the first large model generates the following target processing task based on the question "Post an article related to a certain current hot topic to content platform a and content platform b today": Sub-task d: Call the "online search tool" and use the keyword "Today's hot search list" to search for the current hottest topic; Reason: First determine the most discussed hot event of the day to ensure the timeliness and attention of the topic selection.
[0061] Sub-task e: Call the "Content platform a publishing tool" to create content that conforms to the platform's tone, and call the "Content platform b publishing tool" to create content that conforms to the platform's tone.
[0062] Then, the target processing task can be fed back to the intelligent interaction page, and a confirmation control for the user to confirm whether the target processing task is complete or reasonable is displayed on the intelligent interaction page. If the user confirms through the confirmation control that the target processing task is complete or reasonable, the second large model is called to execute the target processing task. If the user confirms through the confirmation control that the target processing task is incomplete or unreasonable, the first large model can be called to adjust the above target processing task based on the question "Post an article related to a certain hot topic today on content platform a and content platform b", and the adjustment result is fed back to the intelligent interaction page for the user to confirm again, as Figure 5 shown. Or, in the case where the user confirms through the confirmation control that the target processing task is incomplete or unreasonable, a prompt for the user to input the reason for the incompleteness or unreasonableness can be displayed on the intelligent interaction page. After the user inputs the reason for the incompleteness or unreasonableness on the intelligent interaction page, the first large model adjusts the above target processing task based on the question "Post an article related to a certain hot topic today on content platform a and content platform b" and the reason input by the user, and feeds back the adjustment result to the intelligent interaction page for the user to confirm again, as Figure 6 shown.
[0063] In a possible way, adjusting the processing task according to the task evaluation result may include: In the case where the task evaluation result indicates a lack of task steps, according to the first subtask included in the processing task and the sub-questions obtained by decomposing the question, determine the second subtask missing from the processing task, and identify the target execution order between the second subtask and the first subtask, and add the second subtask to the processing task in accordance with the target execution order, where one sub-question corresponds to one subtask; and / or, in the case where the task evaluation result indicates an incorrect task execution order, according to the execution order between the first subtasks and the dependency relationship between the sub-questions, identify the third sub-processing task with an incorrect execution order in the processing task, and adjust the execution order of the third sub-processing task in the processing task according to the dependency relationship.
[0064] Exemplarily, continuing with the above example, the sub-questions obtained by decomposing the question may include: Sub-question a: Selection of hot events; Sub-question b: Obtaining detailed information about hot events; Sub-question c: Creation and publication of article content.
[0065] If the first subtask included in the processing task is: Subtask f: Call the "online search tool" and use the keyword "Today's hot search list" to search for the current hottest topic; Sub-task g: Invoke the "Content Platform A Publishing Tool" to create content that conforms to the platform's tone, and invoke the "Content Platform B Publishing Tool" to create content that conforms to the platform's tone.
[0066] By comparison, it can be seen that the sub-task corresponding to sub-problem b is missing in this processing task, and since the sub-task corresponding to sub-problem b is located between sub-task f and sub-task g, therefore, a sub-task corresponding to sub-problem h can be added between sub-task f and sub-task g. That is, the following processing task can be obtained: Sub-task f: Invoke the "Online Search Tool" and use the keyword "Today's Hot Search List" to search for the current hottest topics; Sub-task h: Invoke the "Online Search Tool" to conduct in-depth searches using the obtained hot keywords; Sub-task g: Invoke the "Content Platform A Publishing Tool" to create content that conforms to the platform's tone, and invoke the "Content Platform B Publishing Tool" to create content that conforms to the platform's tone.
[0067] Exemplarily, continuing to refer to the above example, the sub-problems obtained by decomposing the problem may include: Sub-problem a: Selection of hot events; Sub-problem b: Obtaining detailed information about hot events; Sub-problem c: Creation and publication of article content.
[0068] If the first sub-task included in the processing task is: Sub-task i: Invoke the "Online Search Tool" and use the keyword "Today's Hot Search List" to search for the current hottest topics; Reason: First determine the hottest events of the day to ensure the timeliness and attention of the selected topics.
[0069] Sub-task j: Invoke the "Content Platform A Publishing Tool" to create content that conforms to the platform's tone, and invoke the "Content Platform B Publishing Tool" to create content that conforms to the platform's tone.
[0070] Sub-task k: Invoke the "Online Search Tool" to conduct in-depth searches using the obtained hot keywords.
[0071] By comparison, it can be seen that the execution order between sub-task j and sub-task k is incorrect. Therefore, the execution order between sub-task j and sub-task k can be adjusted so that the execution order of sub-task j is after sub-task k. That is, the following processing task can be obtained: Sub-task i: Invoke the "Online Search Tool" and use the keyword "Today's Hot Search List" to search for the current hottest topics; Sub-task k: Invoke the "Online Search Tool" to conduct in-depth searches using the obtained hot keywords; Sub-task j: Call the "Content Platform a Publishing Tool" to create content that conforms to the platform's tone, and call the "Content Platform b Publishing Tool" to create content that conforms to the platform's tone.
[0072] After determining the target processing task for answering the question, the second largest model associated with the agent can be called to execute the target processing task to obtain the external knowledge required for answering the question.
[0073] Among possible ways, calling the second largest model to execute the target processing task may include: Determine the target sub-task to be executed according to the execution order among multiple sub-tasks; call the second largest model to generate step description text based on the question, the description information of the agent, and the task information of the target sub-task, where the task information includes the execution steps of the target sub-task and the execution order between the execution steps, and the description information is used to describe the functions of the agent and the external tools associated with the agent, and the step description text is used to describe the target execution steps to be executed in the target sub-task; parse the step description text into function call code, where the function call code includes the call logic for the external tool; execute the function call code through the code execution tool.
[0074] It should be understood that in this embodiment, the target sub-task to be executed refers to the sub-task that needs to be executed currently. By way of example, continuing to refer to the above example, the target processing task includes sub-task a, sub-task b, and sub-task c, where the execution order among sub-task a, sub-task b, and sub-task c is: first execute sub-task a, then execute task b, and finally execute sub-task c. If sub-task a has been executed and sub-task b and sub-task c have not been executed, then the target sub-task to be executed is sub-task b.
[0075] By way of example, a second prompt word template can be set in advance. Thus, after determining the target sub-task, the question, the description information of the agent, and the task information of the target sub-task can be filled into the corresponding positions of the second prompt word template to obtain the second prompt word, and the second prompt word is input into the second largest model to obtain the generated step description text. Then, the description text is parsed into function call code by the code parsing module, and finally, the function call code is executed through the code execution tool to obtain the corresponding external knowledge.
[0076] Given that the description information of the agent generally does not change, thus among possible ways, in order to improve the generation efficiency of the prompt word, the second prompt word template can include the description information of the agent. Thus, after determining the target sub-task, the question and the task information of the target sub-task can be filled into the corresponding positions of the second prompt word template, thereby reducing the amount of information written in the second prompt word template, improving the generation efficiency of the second prompt word, and further improving the execution efficiency of the agent question and answer process and reducing the user waiting time.
[0077] Exemplarily, the second prompt template can be: You are a task execution assistant responsible for obtaining the external knowledge required for the task.
[0078] You have the following external tools. Try to use these external tools to obtain the required external knowledge. Each task involves one call to an external tool.
[0079] Internet search tool: Search the Internet according to keywords.
[0080] Content platform a publishing tool: Publish articles to content platform a; Content platform b publishing tool: Publish articles to content platform b; Knowledge base tool: Used to query knowledge from the knowledge base.
[0081] The problem to be solved is: XXXXX.
[0082] The task to be completed is: YYYYY.
[0083] Thus, if the target subtask obtained is: "Call the 'Internet search tool' to perform a deep search using the obtained hot keyword 'MM box office'." Then, based on the problem, the target subtask, and the second prompt template, the following second prompt can be generated: You are a task execution assistant responsible for obtaining the external knowledge required for the task.
[0084] You have the following external tools. Try to use these external tools to obtain the required external knowledge. Each task involves one call to an external tool.
[0085] Internet search tool: Search the Internet according to keywords.
[0086] Content platform a publishing tool: Publish articles to content platform a; Content platform b publishing tool: Publish articles to content platform b; Knowledge base tool: Used to query knowledge from the knowledge base.
[0087] The problem to be solved is: Publish a relevant article to content platform a and content platform b based on a certain hot topic today.
[0088] The task to be completed is: Call the Internet search tool to perform a deep search on the obtained hot keyword 'MM box office'.
[0089] After that, the second prompt can be input into the second large model to obtain the step description text. For example, the step description text can be: "Search for the real-time box office, number of moviegoers, movie-watching experience, plot, and behind-the-scenes stories of MM through the Internet search tool."
[0090] Finally, the step description text can be parsed into function call code, and the function call code can be executed through a code execution tool to obtain content such as the real-time box office of the MM, the number of moviegoers, the movie-watching experience, the plot, and the behind-the-scenes story.
[0091] After obtaining the external knowledge required to answer the question, the question and the external knowledge can be combined to generate an answer to the question.
[0092] Based on the same concept, an embodiment of the present disclosure also provides an intelligent agent question-answering device that integrates different large models, as Figure 7 shown. The intelligent agent question-answering device 700 that integrates different large models may include: A first processing module 701, configured to, in response to obtaining a question input by a user to the intelligent agent, call a first large model associated with the intelligent agent to perform task planning based on the question, obtain a target processing task for answering the question, and send the target processing task to a second large model associated with the intelligent agent, where the reasoning ability of the first large model is stronger than that of the second large model, and the target processing task includes steps of calling external tools associated with the intelligent agent; A second processing module 702, configured to call the second large model to execute the target processing task to obtain external knowledge required to answer the question, where the external knowledge represents knowledge obtained by the second large model by calling external tools; A generation module 703, configured to generate an answer to the question based on the question and the external knowledge.
[0093] Through the above intelligent agent question-answering device 700 that integrates different large models, the first large model associated with the intelligent agent can perform task planning on the obtained question to obtain a target processing task for answering the question, and the second large model can be called to execute the target processing task to obtain external knowledge required to answer the question; an answer to the question can also be generated based on the question and the external knowledge. Among them, the reasoning ability of the first large model is stronger than that of the second large model. Therefore, through the first large model, the question can be analyzed and logically deduced more deeply, so as to more accurately understand the user's intention, and then plan a more accurate and reasonable target processing task for answering the user's question, so that the intelligent agent can output an accurate and reasonable answer based on the target processing task. In addition, since the reasoning ability of the second large model is weaker than that of the first large model, when the second large model executes the target processing task, it generally does not think deeply about the target processing task, so that the response speed to the target processing task can be improved, and then the execution efficiency of the intelligent agent question-answering process can be improved, and the user waiting time can be reduced. Thus, by combining the first large model and the second large model, not only can the user's intention be understood more accurately, but also the user waiting time can be reduced, so that the user's needs can be better met and the user's question-answering experience can be improved.
[0094] In a possible way, the first processing module 701 may include: A classification sub-module, configured to, in response to obtaining a question input by a user to the intelligent agent, call a second large model associated with the intelligent agent to classify the question and obtain a question category corresponding to the question; A first processing sub-module, configured to, when the question category indicates that the question belongs to a fuzzy question or a complex question, call a first large model associated with the intelligent agent to perform task planning based on the question.
[0095] In a possible way, the first processing module 701 may include: A second processing sub-module, configured to call a first large model associated with the intelligent agent to perform the following task planning process: According to the result of intent recognition of the question, disassemble the question into multiple sub-questions; according to the dependency relationship between the multiple sub-questions, determine the execution order between the subtasks corresponding to the multiple sub-questions, where one sub-question corresponds to one subtask; for each subtask, generate the execution steps required for the subtask, where the execution steps include steps for calling external tools and / or steps for calling the second large model; at least according to the execution order between the multiple subtasks and the execution steps of the multiple subtasks, generate a target processing task for answering the question.
[0096] In a possible way, the second processing module 702 may include: A first determination sub-module, configured to determine a target subtask to be executed according to the execution order between the multiple subtasks; A first generation sub-module, configured to call the second large model to generate a step description text based on the question, the description information of the intelligent agent, and the task information of the target subtask, where the task information includes the execution steps of the target subtask and the execution order between the execution steps, the description information is used to describe the functions of the intelligent agent and the external tools associated with the intelligent agent, and the step description text is used to describe the target execution steps to be executed in the target subtask; A parsing sub-module, configured to parse the step description text into function call code, where the function call code includes the call logic for external tools; An execution sub-module, configured to execute the function call code through a code execution tool.
[0097] In a possible way, the first processing module 701 may include: A third processing sub-module, configured to call a first large model associated with the intelligent agent to perform task planning based on the question, obtain an initial processing task for answering the question, and loop through the following process: Display a processing task to the user; in response to obtaining feedback information of the user on the processing task, parse the feedback information to obtain a task evaluation result. If the task evaluation result indicates that the processing task does not include all the task steps required to answer the question, adjust the processing task according to the task evaluation result to obtain a new processing task, and return to the step of displaying the processing task to the user until the task evaluation result indicates that the processing task includes all the task steps; use the processing task after the loop ends as the target processing task for answering the question.
[0098] In a possible way, the third processing sub-module can also be used to: in the case where the task evaluation result indicates that a task step is missing, determine the second sub-task missing from the processing task according to the first sub-task included in the processing task and the sub-questions obtained by disassembling the question, and identify the target execution order between the second sub-task and the first sub-task, and add the second sub-task to the processing task according to the target execution order, where one sub-question corresponds to one sub-task; and / or, in the case where the task evaluation result indicates that the task execution order is incorrect, identify the third sub-processing task with incorrect execution order in the processing task according to the execution order between the first sub-tasks and the dependency relationship between the sub-questions, and adjust the execution order of the third sub-processing task in the processing task according to the dependency relationship.
[0099] In a possible way, the target processing task includes multiple sub-processing tasks, and the multiple sub-processing tasks have a preset execution order. Correspondingly, the second processing module 702 can be used to call the second large model to execute the multiple sub-processing tasks according to the preset execution order to obtain multiple external knowledge required to answer the question; Correspondingly, the generation module 703 can be used to: integrate the multiple external knowledge obtained through the multiple sub-processing tasks into target external knowledge, and fill the question and the target external knowledge into a preset prompt word template to obtain a target prompt word; call the first large model through the target prompt word to generate an answer to the question.
[0100] Based on the same concept, an embodiment of the present disclosure also provides a computer-readable medium, on which a computer program is stored, and when the program is executed by a processing device, the steps of any of the above intelligent agent question-answering methods for integrating different large models are implemented.
[0101] Based on the same concept, an embodiment of the present disclosure also provides an electronic device, which may include: A storage device, on which a computer program is stored; A processing device, configured to execute the computer program in the storage device to implement the steps of any of the above intelligent agent question-answering methods for integrating different large models.
[0102] Based on the same concept, embodiments of the present disclosure also provide a computer program product, including a computer program which, when executed by a processor, implements the steps of any of the above-mentioned intelligent agent question-answering methods for integrating different large models.
[0103] Reference is made below to Figure 8 , which shows a schematic structural diagram of an electronic device 800 suitable for implementing embodiments of the present disclosure. The terminal device in the embodiments of the present disclosure may include, but is not limited to, mobile terminals such as mobile phones, laptop computers, digital broadcast receivers, PDAs (Personal Digital Assistants), PADs (Tablet Computers), PMPs (Portable Multimedia Players), in-vehicle terminals (such as in-vehicle navigation terminals), etc., and fixed terminals such as digital TVs, desktop computers, etc. Figure 8 The electronic device shown is only an example and should not impose any limitation on the functions and usage scope of the embodiments of the present disclosure.
[0104] As Figure 8 shown, the electronic device 800 may include a processing device (such as a central processing unit, a graphics processing unit, etc.) 801, which may perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 802 or a program loaded from a storage device 808 into a random access memory (RAM) 803. In the RAM 803, various programs and data required for the operation of the electronic device 800 are also stored. The processing device 801, the ROM 802, and the RAM 803 are connected to each other through a bus 804. An input / output (I / O) interface 805 is also connected to the bus 804.
[0105] Generally, the following devices may be connected to the I / O interface 805: an input device 806 including, for example, a touch screen, a touchpad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, etc.; an output device 807 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; a storage device 808 including, for example, a magnetic tape, a hard disk, etc.; and a communication device 809. The communication device 809 may allow the electronic device 800 to communicate with other devices wirelessly or wirelessly to exchange data. Although Figure 8 shows the electronic device 800 having various devices, it should be understood that it is not required to implement or include all the shown devices. Instead, more or fewer devices may be implemented or included.
[0106] In particular, according to an embodiment of the present disclosure, the processes described above with reference to the flowchart can be implemented as computer software programs. For example, an embodiment of the present disclosure includes a computer program product that includes a computer program carried on a non-transitory computer-readable medium, and the computer program includes program code for performing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from the network via the communication device 809, or installed from the storage device 808, or installed from the ROM 802. When the computer program is executed by the processing device 801, the above-mentioned functions defined in the method of the embodiment of the present disclosure are performed.
[0107] It should be noted that the computer-readable medium in the present disclosure can be a computer-readable signal medium, a computer-readable storage medium, or any combination of the two. The computer-readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples of the computer-readable storage medium can include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present disclosure, the computer-readable storage medium can be any tangible medium that contains or stores a program, and the program can be used by or in combination with an instruction execution system, apparatus, or device. In the present disclosure, the computer-readable signal medium can include a data signal propagated in a baseband or as part of a carrier wave, which carries computer-readable program code. Such a propagated data signal can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. The computer-readable signal medium can also be any computer-readable medium other than the computer-readable storage medium, and the computer-readable signal medium can send, propagate, or transmit a program for use by or in combination with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted by any appropriate medium, including but not limited to: wires, optical cables, RF (radio frequency), etc., or any suitable combination of the above.
[0108] In some embodiments, any currently known or future-developed network protocol such as HTTP (HyperText Transfer Protocol) can be used for communication, and it can be interconnected with digital data communication in any form or medium (e.g., a communication network). Examples of communication networks include local area networks (“LANs”), wide area networks (“WANs”), the Internet (e.g., the Internet), and end-to-end networks (e.g., ad hoc end-to-end networks), as well as any currently known or future-developed networks.
[0109] The above computer-readable medium can be included in the above electronic device; it can also exist separately without being assembled into the electronic device.
[0110] The above computer-readable medium carries one or more programs. When the above one or more programs are executed by the electronic device, the electronic device is caused to: in response to obtaining a question of a user input agent, call a first large model associated with the agent to perform task planning based on the question, obtain a target processing task for answering the question, and send the target processing task to a second large model associated with the agent, where the reasoning ability of the first large model is stronger than that of the second large model, and the target processing task includes steps of calling external tools associated with the agent; call the second large model to execute the target processing task to obtain external knowledge required for answering the question, where the external knowledge represents knowledge obtained by the second large model calling external tools; generate an answer to the question based on the question and the external knowledge.
[0111] Computer program code for performing the operations of the present disclosure can be written in one or more programming languages or combinations thereof. The above programming languages include, but are not limited to, object-oriented programming languages such as Java, Smalltalk, C++, and also include conventional procedural programming languages such as the “C” language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, executed as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving a remote computer, the remote computer can be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or can be connected to an external computer (e.g., by using an Internet service provider to connect through the Internet).
[0112] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each block in the flowchart or block diagram may represent a module, a segment of a program, or a portion of code that contains one or more executable instructions for implementing a specified logical function. It should also be noted that, in some alternative implementations, the functions noted in the blocks may occur in a different order than noted in the drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, or they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented by a dedicated hardware-based system that performs the specified functions or operations, or by a combination of dedicated hardware and computer instructions.
[0113] The modules described in the embodiments of the present disclosure can be implemented in software or in hardware. In some cases, the name of a module does not constitute a limitation on the module itself.
[0114] The functions described above herein can be performed, at least in part, by one or more hardware logic components. By way of example, and without limitation, exemplary types of hardware logic components that may be used include: field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on a chip (SOCs), complex programmable logic devices (CPLDs), and the like.
[0115] In the context of the present disclosure, a machine-readable medium may be a tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. A machine-readable medium may be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium may include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of a machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0116] The above description is only a preferred embodiment of the present disclosure and an explanation of the applied technical principles. Those skilled in the art should understand that the scope of disclosure involved in the present disclosure is not limited to the technical solutions formed by the specific combination of the above technical features, and should also cover other technical solutions formed by any combination of the above technical features or their equivalent features without departing from the above disclosure concept. For example, the technical solutions formed by mutually replacing the above features with the technical features (but not limited to) having similar functions disclosed in the present disclosure.
[0117] In addition, although the operations are depicted in a particular order, this should not be understood as requiring that the operations be performed in the particular order shown or in sequential order. In certain environments, multitasking and parallel processing may be advantageous. Similarly, although several specific implementation details are included in the above discussion, these should not be construed as limiting the scope of the present disclosure. Certain features described in the context of separate embodiments may also be implemented in combination in a single embodiment. Conversely, the various features described in the context of a single embodiment may also be implemented separately or in any suitable sub-combination in multiple embodiments.
[0118] Although the subject matter has been described in language specific to structural features and / or methodological logical acts, it should be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or acts described above. On the contrary, the specific features and acts described above are merely example forms for implementing the claims. Regarding the apparatus in the above embodiments, the specific manner in which each module performs operations has been described in detail in the embodiments related to the method, and will not be elaborated herein.
Claims
1. An intelligent agent question-answering method for integrating different large models, characterized in that, The intelligent agent question-answering method for integrating different large models includes: In response to obtaining a question input by the user to the intelligent agent, call the first large model associated with the intelligent agent to perform task planning based on the question, obtain a target processing task for answering the question, and send the target processing task to the second large model associated with the intelligent agent, where the reasoning ability of the first large model is stronger than that of the second large model, and the target processing task includes a step of calling an external tool associated with the intelligent agent; Call the second large model to execute the target processing task to obtain external knowledge required to answer the question, where the external knowledge represents the knowledge obtained by the second large model by calling the external tool; Generate an answer to the question based on the question and the external knowledge.
2. The intelligent agent question-answering method for integrating different large models according to claim 1, wherein The step of, in response to obtaining a question input by the user to the intelligent agent, calling the first large model associated with the intelligent agent to perform task planning based on the question includes: In response to obtaining a question input by the user to the intelligent agent, call the second large model associated with the intelligent agent to classify the question to obtain the question category corresponding to the question; In the case where the question category indicates that the question belongs to a fuzzy question or a complex question, call the first large model associated with the intelligent agent to perform task planning based on the question.
3. The intelligent agent question-answering method for integrating different large models according to claim 1, wherein The step of calling the first large model associated with the intelligent agent to perform task planning based on the question to obtain a target processing task for answering the question includes: Call the first large model associated with the intelligent agent to perform the following task planning process: Decompose the question into multiple sub-questions according to the result of intent recognition of the question; Determine the execution order between the subtasks corresponding to the multiple sub-questions according to the dependency relationship between the multiple sub-questions, where one sub-question corresponds to one subtask; For each subtask, generate the execution steps required for the subtask, where the execution steps include a step of calling the external tool and / or a step of calling the second large model; Generate a target processing task for answering the question at least according to the execution order between the multiple sub-tasks and the execution steps of the multiple sub-tasks.
4. The intelligent agent question and answer method for integrating different large models according to claim 3, wherein, The step of calling the second large model to execute the target processing task includes: Determine the target subtask to be executed according to the execution order between the multiple sub-tasks; Call the second large model to generate a step description text based on the question, the description information of the intelligent agent, and the task information of the target subtask, where the task information includes the execution steps of the target subtask and the execution order between the execution steps, the description information is used to describe the function of the intelligent agent and the external tool associated with the intelligent agent, and the step description text is used to describe the target execution steps to be executed in the target subtask; Parse the step description text into function call code, where the function call code includes the call logic of the external tool; Execute the function call code through a code execution tool.
5. The intelligent agent question-answering method for integrating different large language models according to any one of claims 1-4, characterized in that, Invoking the first large model associated with the agent to perform task planning based on the question to obtain a target processing task for answering the question, including: Invoking the first large model associated with the agent to perform task planning based on the question to obtain an initial processing task for answering the question, and looping through the following process: Displaying the processing task to the user; In response to obtaining the user's feedback information on the processing task, parsing the feedback information to obtain a task evaluation result. If the task evaluation result indicates that the processing task does not include all the complete task steps required to answer the question, adjusting the processing task according to the task evaluation result to obtain a new processing task, and returning to the step of displaying the processing task to the user until the task evaluation result indicates that the processing task includes all the complete task steps; Taking the processing task after the loop ends as the target processing task for answering the question.
6. The intelligent agent question answering method for integrating different large language models according to claim 5, characterized in that, The adjusting the processing task according to the task evaluation result includes: In the case where the task evaluation result indicates a lack of task steps, determining a second subtask missing from the processing task according to the first subtask included in the processing task and the sub-questions obtained by decomposing the question, and identifying the target execution order between the second subtask and the first subtask, and adding the second subtask to the processing task in accordance with the target execution order, where one sub-question corresponds to one subtask; and / or, In the case where the task evaluation result indicates an incorrect task execution order, identifying a third sub-processing task with an incorrect execution order in the processing task according to the execution order between the first subtasks and the dependency relationship between the sub-questions, and adjusting the execution order of the third sub-processing task in the processing task according to the dependency relationship.
7. The intelligent agent question-answering method for integrating different large language models according to any one of claims 1-4, characterized in that, The target processing task includes multiple sub-processing tasks, and the multiple sub-processing tasks have a preset execution order. Invoking the second large model to execute the target processing task to obtain external knowledge required to answer the question, including: Invoking the second large model to execute the multiple sub-processing tasks in accordance with the preset execution order to obtain multiple external knowledge required to answer the question; Generating an answer to the question based on the question and the external knowledge, including: Integrating the multiple external knowledge obtained through the multiple sub-processing tasks into target external knowledge, and filling the question and the target external knowledge into a preset prompt template to obtain a target prompt; Invoking the first large model through the target prompt to generate an answer to the question.
8. An intelligent agent question-answering device that integrates different large models, characterized in that, The intelligent agent question-answering device that fuses different large models includes: The first processing module is configured to, in response to obtaining a question input by a user to an intelligent agent, call a first large model associated with the intelligent agent to perform task planning based on the question, obtain a target processing task for answering the question, and send the target processing task to a second large model associated with the intelligent agent, where the reasoning ability of the first large model is stronger than that of the second large model, and the target processing task includes a step of calling an external tool associated with the intelligent agent; The second processing module is configured to call the second large model to execute the target processing task to obtain external knowledge required to answer the question, where the external knowledge represents the knowledge obtained by the second large model calling the external tool; The generation module is configured to generate an answer to the question based on the question and the external knowledge.
9. A computer-readable medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processing device, it implements the steps of the method according to any one of claims 1-7.
10. An electronic device, characterized in that, It includes: A storage device on which a computer program is stored; A processing device configured to execute the computer program in the storage device to implement the steps of the method according to any one of claims 1-7.
11. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1-7.
Citation Information
Patent Citations
Question and answer task processing method, device and equipment and readable storage medium
CN117520514A
Generative target planning model construction method and device
CN117681902A
Intelligent agent question and answer method and device, storage medium and equipment
CN118469017A
Knowledge question and answer method and device, readable medium, electronic equipment and program product
CN118964694A
Interactive question and answer task processing method based on multi-agent cooperation and related device
CN119537542A
Cited By
Multi-agent information processing method and device based on large model, electronic equipment, storage medium and program product
CN120763321A
Multi-agent information processing methods, devices, electronic devices, storage media, and program products based on large models.
CN120763321B
Task processing method and device, equipment, storage medium and product
CN121300986A
Intelligent question and answer method based on multi-agent collaboration and related device
CN121457482A
Medical business system processing method based on multi-agent cooperation, medium and equipment
CN121528472A