Interaction method and apparatus, and computer-readable storage medium
By combining multiple intelligent agents and scheduling different types of intelligent agents to perform tasks, the problem of single-model AI assistants being unable to meet diverse needs and performance bottlenecks has been solved, achieving more efficient and lower-cost user interaction and improving the performance and adaptability of AI assistants.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- BEIJING ZITIAO NETWORK TECH CO LTD
- Filing Date
- 2025-08-05
- Publication Date
- 2026-04-23
AI Technical Summary
Existing single-model AI assistants are unable to meet the diverse needs of users in different scenarios, have obvious performance bottlenecks, and are costly, making further development and application difficult.
By combining multiple agents and scheduling different types of agents to perform different tasks, the task order and agent combination are determined according to the user's interaction intent, thereby achieving efficient task allocation and collaboration.
It improves the performance and adaptability of the AI assistant, enabling it to handle more complex and varied user needs, reduces overall costs, and improves the accuracy and naturalness of responses.
Smart Images

Figure CN2025112625_23042026_PF_FP_ABST
Abstract
Description
Interaction methods and apparatus, computer-readable storage media
[0001] Cross-references to related applications
[0002] This application is based on and claims priority to Chinese application No. 202411440940.2, filed on October 15, 2024, the disclosure of which is incorporated herein by reference in its entirety. Technical Field
[0003] This disclosure relates to the field of artificial intelligence, and in particular to interactive methods and apparatuses, and computer-readable storage media. Background Technology
[0004] Artificial intelligence (AI) agents are intelligent entities capable of perceiving their environment, making decisions, performing actions, and interacting with their environment. AI agents are a crucial concept in the field of artificial intelligence and play a key role in achieving automated decision-making and task execution.
[0005] With the continuous development and popularization of artificial intelligence technology, AI entities will play an increasingly important role in people's lives and work. Summary of the Invention
[0006] According to a first aspect of this disclosure, an interaction method is provided, comprising: in response to receiving input information from a user, determining the user's interaction intent; determining a plurality of tasks and their execution order based on the interaction intent; for each of the plurality of tasks, determining an agent corresponding to each task to obtain a combination of the plurality of agents, wherein different types of agents perform different types of tasks; and scheduling the combination of the plurality of agents to execute the plurality of tasks according to the execution order.
[0007] According to a second aspect of this disclosure, an interaction device is provided, comprising: an intent determination module configured to determine a user's interaction intent in response to receiving input information from a user; a task determination module configured to determine a plurality of tasks and their execution order based on the interaction intent; an agent determination module configured to determine an agent for each of the plurality of tasks, thereby obtaining a combination of the plurality of agents, wherein different types of agents perform different types of tasks; and a scheduling module configured to schedule the combination of the plurality of agents to execute the plurality of tasks according to the execution order.
[0008] According to a third aspect of this disclosure, an interactive device is provided, comprising: a memory; and a processor coupled to the memory, the processor being configured to execute an interactive method according to any of the embodiments of this disclosure based on instructions stored in the memory.
[0009] According to a fourth aspect of this disclosure, an interactive system is provided, comprising: an interactive device according to any of the embodiments of this disclosure; and a plurality of intelligent agents.
[0010] According to a fifth aspect of this disclosure, a computer-readable storage medium is provided that stores computer program instructions thereon, which, when executed by a processor, implement the interaction method described in any of the embodiments of this disclosure.
[0011] According to a sixth aspect of this disclosure, a computer program product is provided, including computer program instructions that, when executed by a processor, implement the interaction method described in any of the embodiments of this disclosure. Attached Figure Description
[0012] Preferred embodiments of the present disclosure are described below with reference to the accompanying drawings. The accompanying drawings, which are included to provide a further understanding of the present disclosure, and which, together with the following detailed description, are incorporated in and form a part of this specification and are used to explain the present disclosure. It should be understood that the drawings described below only relate to some embodiments of the present disclosure and are not intended to limit the present disclosure. In the drawings:
[0013] Figure 1 shows a flowchart of an interaction method according to some embodiments of the present disclosure;
[0014] Figure 2 illustrates a flowchart of determining an intelligent agent according to some embodiments of the present disclosure;
[0015] Figure 3 illustrates a schematic diagram of an interaction method according to some embodiments of the present disclosure;
[0016] Figure 4 shows a block diagram of an interactive device according to some embodiments of the present disclosure;
[0017] Figure 5 shows a block diagram of an interactive device according to some other embodiments of the present disclosure;
[0018] Figure 6 shows a block diagram of an electronic device according to some embodiments of the present disclosure.
[0019] It should be understood that, for ease of description, the dimensions of the various parts shown in the accompanying drawings are not necessarily drawn to actual scale. The same or similar reference numerals are used in the various drawings to denote the same or similar parts. Therefore, once an item is defined in one drawing, it may not be discussed further in subsequent drawings. Detailed Implementation
[0020] The technical solutions of the embodiments of this disclosure will be clearly and completely described below with reference to the accompanying drawings. However, it is obvious that the described embodiments are only some embodiments of this disclosure, and not all embodiments. The following description of the embodiments is merely illustrative and is in no way intended to limit this disclosure or its application or use. It should be understood that this disclosure can be implemented in various forms and should not be construed as limited to the embodiments set forth herein.
[0021] It should be understood that the various steps described in the method embodiments of this disclosure may be performed in different orders and / or in parallel. Furthermore, method embodiments may include additional steps and / or omit the steps shown. The scope of this disclosure is not limited in this respect. Unless otherwise specifically stated, the relative arrangement, numerical expressions, and values of components and steps set forth in these embodiments should be interpreted as merely exemplary and do not limit the scope of this disclosure.
[0022] As used in this disclosure, the term "comprising" and its variations are open-ended terms that include at least the following elements / features but do not exclude other elements / features, i.e., "including but not limited to". Furthermore, as used in this disclosure, the term "including" and its variations are open-ended terms that include at least the following elements / features but do not exclude other elements / features, i.e., "including but not limited to". Therefore, "comprising" and "including" are synonymous. The term "based on" means "at least partially based on".
[0023] Throughout this specification, the terms "one embodiment," "some embodiments," or "embodiment" mean that a specific feature, structure, or characteristic described in connection with an embodiment is included in at least one embodiment of the invention. For example, the term "one embodiment" means "at least one embodiment"; the term "another embodiment" means "at least one additional embodiment"; and the term "some embodiments" means "at least some embodiments." Furthermore, the appearance of the phrases "in one embodiment," "in some embodiments," or "in an embodiment" in various places throughout the specification does not necessarily refer to the same embodiment, but may refer to the same embodiment.
[0024] It should be noted that the concepts of "first," "second," etc., used in this disclosure are used only to distinguish different devices, modules, or units, and are not intended to define the order of functions performed by these devices, modules, or units or their interdependencies. Unless otherwise specified, the concepts of "first," "second," etc., are not intended to imply that the objects described herein must be in a given temporal, spatial, rank, or any other given order.
[0025] It should be noted that the terms "a" and "a plurality of" used in this disclosure are illustrative rather than restrictive, and those skilled in the art should understand that, unless otherwise expressly indicated in the context, they should be understood as "one or more".
[0026] The names of messages or information exchanged between multiple devices in the embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of such messages or information.
[0027] The embodiments of this disclosure are described in detail below with reference to the accompanying drawings; however, this disclosure is not limited to these specific embodiments. These specific embodiments can be combined with each other, and the same or similar concepts or processes may not be described again in some embodiments. Furthermore, in one or more embodiments, specific features, structures, or characteristics can be combined in any suitable manner that will be apparent to those skilled in the art from this disclosure.
[0028] In the field of AI assistants, as user needs become increasingly complex and diverse, the limitations of single models are becoming more and more apparent. These limitations are mainly manifested in the following ways: limited functionality, making it difficult to simultaneously meet the diverse needs of users in different scenarios; performance bottlenecks, as a single intelligent agent may perform poorly when handling certain specific tasks; and high costs, as larger and more expensive models are often required to improve overall performance. These shortcomings limit the further development and application of AI assistants.
[0029] This disclosure provides an interaction method, including: in response to receiving input information from a user, determining the user's interaction intent; determining multiple tasks and their execution order based on the interaction intent; for each of the multiple tasks, determining an agent corresponding to each task to obtain a combination of multiple agents, wherein different types of agents execute different types of tasks; and scheduling the combination of multiple agents to execute the multiple tasks according to the execution order.
[0030] Compared to a single model, combining multiple agents and scheduling agents of different types to perform different tasks allows the combination of agents to handle more complex and varied user needs, improving the performance and adaptability of AI assistants. The collaboration of multiple specialized agents also results in more accurate and natural responses.
[0031] Figure 1 shows a flowchart of an interaction method according to some embodiments of the present disclosure.
[0032] As shown in Figure 1, the interaction method includes: Step S1, in response to receiving user input information, determining the user's interaction intent; Step S2, determining multiple tasks and their execution order according to the interaction intent; Step S3, for each of the multiple tasks, determining the agent corresponding to each task, obtaining a combination of multiple agents, wherein different types of agents execute different types of tasks; Step S4, scheduling the combination of multiple agents to execute multiple tasks according to the execution order.
[0033] User input can include text, URLs, images, videos, etc.
[0034] User interaction intent refers to the underlying purpose or desired goal of a user's interaction behavior.
[0035] Multiple tasks may include, for example, multiple steps that need to be performed to fulfill a user's interactive intent. These tasks can be executed in parallel, sequentially, or some tasks can be executed in parallel while others are executed sequentially.
[0036] Intelligent agents, for example, are agents that perform specific tasks, also known as specialized agents. Each type of agent possesses specific expertise, thus refining the division of labor among agents. During the scheduling of a combination of multiple agents to perform multiple tasks, each agent performs the task it excels at, and the combination of multiple agents cooperates with each other.
[0037] The interaction method in this embodiment can be executed by an interaction device. The interaction device is, for example, a scheduler, and the above steps S1-S4 can be executed by the scheduler. By introducing the role of a "scheduler," intelligent task allocation can be achieved.
[0038] The interaction method in this embodiment can be executed on the client side or partially on the server side.
[0039] The process of determining an intelligent agent according to some embodiments of this disclosure will be described below with reference to Figure 2.
[0040] As shown in Figure 2, step S3, for each of the multiple tasks, determines the agent corresponding to each task, including: step S31, determining the type and requirements of each task; step S32, determining one or more candidate agents that match the type of each task; and step S33, selecting the agent corresponding to each task from one or more candidate agents according to the requirements of each task.
[0041] If the task type is, for example, chatting, then the corresponding agent type is a chatter-type agent (e.g., a role-playing agent). One or more candidate agents for a task may be agents of the same type but with different tones and styles. For example, different chatter-type agents have different tones and styles, such as friendly and approachable, humorous and witty, calm and rational, or artistic and emotional.
[0042] The requirements of a task can be determined based on the user's intent. For example, if the user desires comfort, the task requirement is to be warm and friendly during the chat. In this case, a friendly and approachable agent is selected from the chat-type agents to perform the chat task.
[0043] In some embodiments, in response to receiving user input information, determining the user's interaction intent includes: in response to receiving user input information, determining the type of input information; and determining the interaction intent based on the type of input information.
[0044] The type of input information refers to the medium in which the input information is presented, such as text, URLs, images, videos, etc. Different types of input information may have different interactive intentions.
[0045] If the user input includes an image, the user's intent is likely image processing or image-to-text conversion. If the user input is in text format, natural language processing techniques can be used to analyze the user's intent.
[0046] In some embodiments, in response to receiving user input information, determining the user's interaction intent includes: in response to determining that the input information does not meet a second specified condition, outputting a second guidance question to the user; and determining the interaction intent based on the user's feedback to the second guidance question.
[0047] For example, if a user's question is not clear enough and it is difficult to determine the user's interaction intent, the input information does not meet the second specified condition. Therefore, it is necessary to output a second guiding question to the user to guide them to clarify their intent.
[0048] In some embodiments, determining multiple tasks and their execution order based on the interaction intent includes: determining the complex task corresponding to the interaction intent; splitting the complex task into multiple tasks; and determining the execution order based on the relationship between the multiple tasks.
[0049] For example, the scheduler determines how many intentions are contained in a user's input sentence, which tasks need to be performed to fulfill these intentions, and determines the order in which these tasks are executed. If a user wants to create an article that meets their needs, they first need to perform an online search, and then write based on the search results. Therefore, the task execution order is: search first, then create.
[0050] One approach is to first determine the execution order of the tasks, and then determine the corresponding agent for each task based on the tasks and their execution order. Alternatively, one can first determine the agent for each task, and then determine the execution order of the tasks based on the characteristics of the agent, user intent, task requirements, etc. The relationship between multiple tasks can be, for example, an input-output relationship. For instance, if executing task B requires the processing result of task A, then task A should be executed first, followed by task B.
[0051] By breaking down complex tasks and refining the division of labor and collaboration among intelligent agents, efficient collaboration can be achieved. Simultaneously, through granular task allocation of complex tasks, a basic model can solve a basic task derived from the complex task, reducing the overuse of models with large parameters and high costs.
[0052] In some embodiments, the interaction method further includes: providing task information to at least one of the plurality of agents based on input information.
[0053] For example, the scheduler assigns tasks to selected agents and processes the user's input information, breaking it down into the context information needed to execute the task and providing it to the corresponding agent.
[0054] The scheduler can provide task information to all agents or only to a subset of them. For example, if tasks are executed sequentially, and each subsequent agent relies solely on the output of the preceding agent to complete its task, then only the task information needs to be provided to the first agent in the sequence. If the first few tasks are executed in parallel, then task information is provided to all of the first few agents in the sequence.
[0055] In some embodiments, the interaction method further includes: determining a base model for at least one of the multiple agents based on input information.
[0056] The base model and the agent do not have to be bound together. The base model of the agent can be flexibly adjusted according to the task requirements of different scenarios.
[0057] The foundational models include, for example, models that excel at tool invocation, models that excel at verbal communication, models that excel at text processing, and models that excel at mathematical reasoning. For tasks that require interaction with external systems, models that excel at tool invocation are selected; for tasks that require natural and fluent dialogue, models that excel at verbal communication are selected; for tasks that process and generate large amounts of text, models that excel at text processing are selected; and for scientific computing and logical analysis tasks, models that excel at mathematical reasoning are selected.
[0058] If the user's input is in text format, the base can also be selected based on the text's token. For example, if the user's input is long, a base model with a long context window can be selected for the agent.
[0059] The multiple intelligent agents include at least one of the following: a role-playing intelligent agent, an information retrieval intelligent agent, a creative intelligent agent, and a question-answering intelligent agent. Among these, the different creative intelligent agents have different language styles.
[0060] The characteristics of an information retrieval agent include rigor and expertise, proficiency in finding and organizing information, and the ability to perform web searches, database queries, and information aggregation. The foundational model of an information retrieval agent is, for example, a model with broad knowledge.
[0061] Characteristics of creative agents include being imaginative, skilled in content creation, and capable of writing, storytelling, and poetry. The foundational model for creative agents is, for example, a model with strong creativity.
[0062] A question-answering intelligent agent, also known as a problem-solving intelligent agent, is characterized by strong logic, a knack for analyzing and solving complex problems, and the ability to perform mathematical calculations, logical reasoning, and solution design. The base model of a question-answering intelligent agent may employ a model with strong reasoning capabilities.
[0063] Multiple intelligent agents can also include chatter agents. Chatter agents are characterized by their proficiency in everyday conversations, handling greetings, small talk, and simple questions and answers. The foundational model for chatter agents is, for example, a model adept at natural dialogue.
[0064] In addition, multiple intelligent agents can include translation agents, text augmentation agents, image augmentation agents, text-to-image agents, etc.
[0065] In some embodiments, the interaction method further includes: if the output of multiple agents does not meet a first specified condition, redetermining the agent for each task to obtain a combination of the redetermined agents; and scheduling the combination of the redetermined agents to re-execute the multiple tasks according to the execution order.
[0066] The first specified condition is, for example, that the outputs of multiple agents are consistent, or that the final result generated based on the outputs of multiple agents satisfies the user's interaction intent.
[0067] If the first specified condition is not met, the scheduler can reselect an agent, change the type of the selected agent, or select a different agent within the same type. For example, if the outputs of agent A executing task 1 and agent B executing task 2 are inconsistent, an agent C of the same type as agent B but with a different style can be used to execute task 2, or an agent D of a different type than agent B can be used to execute task 2.
[0068] In other words, during runtime, the scheduler can dynamically adjust the combination and working methods of intelligent agents, flexibly configuring multiple agents to work collaboratively according to different business scenarios. Working methods include, for example, perceiving the environment, planning and decision-making, and executing actions. Simultaneously, it can iteratively optimize and continuously adjust the collaborative methods of the intelligent agents based on user feedback.
[0069] An additional detector can be added to check whether the tasks, task order, and selected agents determined by the scheduler are reasonable. If they are not reasonable, the scheduler will rearrange the tasks, task order, and selected agents.
[0070] In some embodiments, the interaction method further includes: when the outputs of multiple agents do not meet a first specified condition, re-determining multiple tasks and their execution order, determining the agent corresponding to each task, and scheduling a combination of multiple agents to execute multiple tasks.
[0071] For example, if the outputs of multiple agents do not meet the first specified condition, it may not only be due to an unreasonable combination of agents, but also to an unreasonable task arrangement. Therefore, multiple tasks and their execution order can be redefined, and agents can be reassigned.
[0072] In some embodiments, the interaction method further includes: outputting a guidance question to the user when the outputs of multiple agents do not meet a first specified condition; redetermining the user's interaction intent based on the user's feedback on the guidance question; redetermining multiple tasks and their execution order, determining the agent corresponding to each task, and scheduling a combination of multiple agents to execute multiple tasks.
[0073] For example, if the outputs of multiple agents do not meet the first specified condition, it may be due to an inaccurate determination of the user's intent. In this case, the user can be asked to clarify their intent by asking questions, and the tasks can be rearranged and reassigned.
[0074] In some embodiments, the interaction method further includes: in the event of conflicting outputs from multiple agents, re-determining multiple tasks and their execution order, determining the agent corresponding to each task, and scheduling a combination of multiple agents to execute multiple tasks.
[0075] For example, if the outputs of multiple agents conflict, it could be due to an inappropriate combination of agents or an unreasonable task assignment. Therefore, tasks can be rearranged and reassigned to resolve the conflict. The agent responsible for final integration can determine whether there is a conflict in the outputs of all agents, or other detection modules can be used to determine this.
[0076] If the outputs of multiple agents conflict, the combination of agents can be changed first. If the conflict is resolved, the result can be generated and output to the user. If the conflict is not resolved, and the number of times the agents are recombined reaches a threshold, multiple tasks can be redefined, and a combination of agents can be selected for the redefined tasks.
[0077] In some embodiments, the interaction method further includes: outputting a first guiding question to the user when there is a conflict in the outputs of multiple agents; redetermining the user's interaction intent based on the user's feedback on the first guiding question; redetermining multiple tasks and their execution order, determining the agent corresponding to each task, and scheduling a combination of multiple agents to execute multiple tasks.
[0078] For example, if the outputs of multiple agents conflict, the user can be consulted to decide which part they need.
[0079] After receiving user feedback, the system can output the portion of the conflicting content that the user needs. Based on this feedback, the system can also redefine the user's interaction intent and reassign tasks. Furthermore, when assigning tasks to agents, agents that generate content that the user does not need can be excluded.
[0080] If the outputs of multiple agents conflict, the combination of agents can be changed first. If the conflict is resolved, the result can be generated and output to the user. If the conflict is not resolved, and the number of times the agents are recombined reaches a threshold, multiple tasks can be redefined. For each redefined task, a combination of agents can be selected. If the number of times the task is redefined reaches the threshold and the conflict is still not resolved, a first guiding question will be output to the user to seek their help.
[0081] After multiple agents output their results, the outputs of these agents can be integrated to obtain an integrated result, which is then output. The integration process is described below.
[0082] For example, a scheduler determines a first target agent, which then collects and integrates the outputs of multiple agents. The first target agent can be one of the multiple agents performing the aforementioned task, or it can be a dedicated agent for integration.
[0083] The scheduler can also identify a second target agent and schedule that agent to output the integrated result. The second target agent can be one of the aforementioned agents, or it can be a dedicated output agent. The second target agent can be the same as the first target agent, or it can be a different agent.
[0084] The scheduler can determine the first and second target agents based on the tasks and their execution order. For example, the agent that executes the last task can be designated as the first and second target agents, responsible for integration and output.
[0085] It can also perform a consistency check on the integration results, and output the results that have passed the consistency check if the consistency check passes.
[0086] Consistency checks, for example, ensure that the integrated result is consistent in logic and tone. For instance, if the integrated result is an article, one could check for contradictions between paragraphs or differences in writing style.
[0087] The scheduler schedules the third target agent to perform a consistency check on the integrated results. If the consistency check passes, the scheduler schedules the second target agent to output the results that have passed the consistency check.
[0088] If the consistency check fails, the outputs of multiple agents are re-integrated, and the consistency check is performed on the re-integrated result.
[0089] For example, if the consistency check fails, it may be due to a problem in the integration process. In this case, the outputs of multiple agents can be re-integrated, and only the consistent content can be selected for integration.
[0090] If the consistency check fails, the system can be reassembled first. If reassembly resolves the consistency issue, the reassembled result is output to the user. If the number of reassemblies reaches a threshold without resolving the consistency issue, the agent combination can be changed. If the recombination of agents resolves the consistency issue, the result is generated and output to the user. If the consistency issue remains unresolved and the number of agent recombinations reaches the threshold, multiple tasks can be redefined. For each redefined task, an agent combination is selected. If the number of redefined tasks reaches the threshold and the consistency issue is still not resolved, a guidance question is output to the user for assistance.
[0091] Multiple intelligent agents can exchange information by sharing a knowledge base. For example, multiple intelligent agents can share some common knowledge by sharing a knowledge base. By centrally storing common knowledge and making it available to multiple intelligent agents, it is beneficial to update and maintain the common knowledge in real time.
[0092] According to some embodiments of this disclosure, by determining the user's interaction intent, determining multiple tasks and their execution order, and determining a corresponding intelligent agent for each task, a combination of multiple intelligent agents is obtained, and the combination of multiple intelligent agents is scheduled to execute multiple tasks, thereby satisfying the user's interaction intent.
[0093] Compared to a single model, combining multiple agents, with different types of agents performing different tasks, allows the combination of agents to handle more complex and varied user needs, improving the performance and adaptability of AI assistants. The collaboration of multiple specialized agents also results in more accurate and natural responses.
[0094] Multi-agent collaboration can fully leverage the strengths of different models and allow individual agents to optimize, improving the performance of a particular agent in a specific task, thereby enhancing overall performance and reducing costs.
[0095] Figure 3 illustrates a schematic diagram of an interaction method according to some embodiments of the present disclosure.
[0096] As shown in Figure 3, the system can provide prompts to users using preset personas to assist them with input. Simultaneously, it can also verify the user's identity to enhance security.
[0097] Determine whether the user input is long text, an image, a video, or a link. If it is long text, select a model appropriate for its token length. If it is an image, video, or link, you can also determine the user's intent based on the type of input and select the corresponding model.
[0098] If the carrier is not of the above type, but is text of normal length, then determine whether the user's meaning is clear. If not, use a question to help the user clarify their intent. If the user's intent can be clearly determined, then determine the task to be performed and its corresponding agent.
[0099] Figure 4 shows a block diagram of an interactive device according to some embodiments of the present disclosure.
[0100] As shown in Figure 4, the interaction device 4 includes: an intent determination module 41, configured to determine the user's interaction intent in response to receiving user input information; a task determination module 42, configured to determine multiple tasks and their execution order based on the interaction intent; an agent determination module 43, configured to determine the agent for each of the multiple tasks, thereby obtaining a combination of multiple agents, wherein different types of agents perform different types of tasks; and a scheduling module 44, configured to schedule the combination of multiple agents to execute multiple tasks according to the execution order.
[0101] The intent determination module 41 of the interaction device 4 can be used to execute step S1 of FIG1. The task determination module 42 of the interaction device 4 can be used to execute step S2 of FIG1. The agent determination module 43 can be used to execute step S3 of FIG1. The scheduling module 44 can be used to execute step S4 of FIG1.
[0102] In some embodiments, the interaction device 4 further includes a providing module configured to provide task information to at least one of a plurality of agents based on input information.
[0103] In some embodiments, the interaction device 4 further includes a base determination module configured to determine a base model for at least one of the plurality of agents based on input information.
[0104] In some embodiments, the interaction device 4 further includes: a loop module configured to, when the output of multiple agents does not meet a first specified condition, redetermine the agent for each task to obtain a combination of the redetermined agents; and schedule the combination of the redetermined agents to re-execute the multiple tasks according to the execution order.
[0105] In some embodiments, the interaction device 4 further includes a loop module configured to, in the event of a conflict between the outputs of multiple agents, re-determine multiple tasks and their execution order, determine the agent corresponding to each task, and schedule a combination of multiple agents to execute multiple tasks.
[0106] In some embodiments, the interaction device 4 further includes: a loop module configured to output a first guiding question to the user when there is a conflict in the outputs of multiple agents; the user's feedback on the first guiding question to redetermine the user's interaction intent; redetermine multiple tasks and their execution order, determine the agent corresponding to each task, and schedule the combination of multiple agents to execute multiple tasks.
[0107] In some embodiments, the interaction device 4 further includes: an integration module configured to integrate the outputs of multiple agents to obtain an integrated result; and output the integrated result.
[0108] In some embodiments, the integration module is further configured to perform a consistency check on the integration result; if the consistency check passes, the result of passing the consistency check is output.
[0109] In some embodiments, the integration module is further configured to re-integrate the outputs of multiple agents if the consistency check fails, and to perform a consistency check on the re-integrated result.
[0110] Figure 5 shows a block diagram of an interactive device according to other embodiments of the present disclosure.
[0111] As shown in FIG5, the interactive device 5 includes: a memory 51; and a processor 52 coupled to the memory 51, the processor 52 being configured to execute the interactive method of any of the foregoing embodiments based on instructions stored in the memory 51.
[0112] Memory 51 is used to store one or more computer-readable instructions. Memory 51 may include any combination of various forms of computer-readable storage media, such as volatile memory and / or non-volatile memory, including but not limited to random access memory (RAM), dynamic random access memory (DRAM), static random access memory (SRAM), read-only memory (ROM), and flash memory. Memory 51 may, for example, store operating systems, application programs, boot loaders, databases, and other programs, as well as various application programs and various data.
[0113] The processor 52 is configured to execute computer-readable instructions to implement the interaction method of any of the foregoing embodiments. Specific implementation details of each step of the interaction method can be found in the above embodiments; repeated details will not be elaborated upon here.
[0114] Processor 52 can be configured to execute steps S1-S4 of Figure 1. Processor 52 can be various processing devices, such as a central processing unit (CPU), a network processor (NP), etc.; it can also be a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components. The central processing unit (CPU) can be an x86 or ARM architecture, etc.
[0115] The processor 52 and the memory 51 can communicate with each other directly or indirectly. For example, the processor 52 and the memory 51 can communicate via a network. The network can include a wireless network, a wired network, and / or any combination of wireless and wired networks. The processor 52 and the memory 51 can also communicate with each other via a system bus, which is not limited in this disclosure.
[0116] It should be noted that the components of the interactive device 5 shown in Figure 5 are merely exemplary and not limiting. Depending on the specific application requirements, the interactive device 5 may also have other components. The processor 52 can control other components in the interactive device 5 to perform the desired functions.
[0117] Interactive devices can be implemented by software, firmware, and / or hardware, and can be integrated into electronic devices with relevant applications installed.
[0118] This disclosure provides an interactive system, including: an interactive device according to any of the embodiments of this disclosure; and multiple intelligent agents. The interactive system according to this disclosure enables multi-agent cooperation and provides a flexible and efficient multi-agent cooperation framework.
[0119] In some embodiments, the interactive system further includes a user interface configured to: receive user input; and display the output of the interactive device.
[0120] In some embodiments, the interactive system also includes a knowledge base configured to store knowledge and information shared by the agents.
[0121] In some embodiments, the interactive system also includes an external API interface configured to connect to external services and data sources.
[0122] According to some embodiments of this disclosure, a system for constructing a multi-agent AI assistant can be built, the system including a scheduler and multiple selectable agents to provide services to users.
[0123] The system framework of a multi-agent AI assistant can adopt an extensible architecture, such as a plug-in system, which can support the flexible addition of new agents and base models.
[0124] Each agent can be an independent module, making it easy to add, remove, or update. Furthermore, standardized interfaces can be used to define a unified data exchange format, facilitating the integration of new agents and models.
[0125] Developers can create their own intelligent agents. Furthermore, developers can integrate their agents into the AI assistant's system through dialogue. For example, if a developer creates an agent to search for news and adds it to the system, the user-created agent can be invoked to handle user intents related to news during conversations with the AI assistant. By adopting a scalable architecture, the system's scalability is enhanced, facilitating the integration of new models and functions and adapting to future technological advancements.
[0126] Figure 6 shows a block diagram of an electronic device according to some embodiments of the present disclosure.
[0127] The electronic device 6 shown in Figure 6 can be a computer system with a dedicated hardware structure, capable of performing corresponding functions when relevant applications are installed.
[0128] Electronic devices include, but are not limited to, mobile terminals such as smartphones, laptops, personal digital assistants (PDAs), tablet PCs (Tablet Personal Computers), PMPs (Portable Multimedia Players), in-vehicle terminals (such as in-vehicle navigation terminals), wearable devices, and fixed terminals such as digital televisions and desktop computers.
[0129] As shown in Figure 6, the Central Processing Unit (CPU) 61 executes various processes based on the program stored in the Read-Only Memory (ROM) 62 or the program loaded from the storage section 68 into the Random Access Memory (RAM) 63. The RAM 63 stores data required as needed when the CPU 61 executes various processes. The CPU is merely exemplary and can also be other types of processors, such as the various processors described above. The ROM 62, RAM 63, and storage section 68 can be various forms of computer-readable storage media. It should be noted that although the ROM 62, RAM 63, and storage section 68 are shown separately in Figure 6, one or more of them can be combined or located in the same or different memories or storage modules.
[0130] CPU 61, ROM 62 and RAM 63 are interconnected via bus 64. Input / output interface 65 is also connected to bus 64.
[0131] The following components are connected to the input / output interface 65: input section 66, such as a touchscreen, touchpad, keyboard, mouse, image sensor, microphone, accelerometer, gyroscope, etc.; output section 67, including displays such as cathode ray tube (CRT), liquid crystal display (LCD), speakers, vibrators, etc.; storage section 68, including hard disks, magnetic tapes, etc.; and communication section 69, including network interface cards such as LAN cards, modems, etc. The communication section 69 allows communication processing to be performed via a network such as the Internet. It is readily understood that although the various devices or modules in the electronic device 6 shown in Figure 6 communicate via bus 64, they can also communicate via a network or other means, wherein the network can include wireless networks, wired networks, and / or any combination of wireless and wired networks.
[0132] As needed, drive 610 is also connected to input / output interface 65. Removable media 611, such as disks, optical disks, magneto-optical disks, semiconductor memories, etc., are installed on drive 610 as needed, so that computer programs read from them can be installed into storage section 68 as needed.
[0133] When the above series of processes are implemented through software, the program constituting the software can be installed from a network such as the Internet or a storage medium such as removable medium 611.
[0134] According to embodiments of this disclosure, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, some embodiments of this disclosure include a computer program product that, when run on a computer, causes the computer to implement the interactive methods of any of the foregoing embodiments. The computer program product includes a computer program carried on a computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via communication section 69, or installed from storage section 68, or installed from ROM 62. When the computer program is executed by CPU 61, the interactive methods of embodiments of this disclosure are performed.
[0135] It should be noted that, in the context of this disclosure, a computer-readable medium can be a tangible medium that may contain or store programs for use by or in conjunction with an instruction execution system, apparatus, or device.
[0136] A computer-readable medium may be a computer-readable storage medium, a computer-readable signal medium, or any combination thereof.
[0137] Computer-readable storage media include, but are not limited to, systems, apparatuses, or devices that are electrical, magnetic, optical, electromagnetic, infrared, or semiconductor, or any combination thereof. More specific examples of computer-readable storage media may include, but are not limited to, electrical connections having one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof. In this disclosure, a computer-readable storage medium can be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution system, apparatus, or device. A computer program is stored on the computer-readable storage medium that, when executed by a processor, implements the interactive methods of any of the foregoing embodiments.
[0138] Computer-readable signal media may include data signals propagated in baseband or as part of a carrier wave, carrying computer-readable program code. Such propagated data signals may take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. Computer-readable signal media may also be any computer-readable medium other than computer-readable storage media, capable of sending, propagating, or transmitting programs for use by or in connection with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium may be transmitted using any suitable medium, including but not limited to: wires, optical fibers, RF (radio frequency), etc., or any suitable combination thereof.
[0139] The aforementioned computer-readable medium may be included in the aforementioned electronic device; or it may exist independently and not assembled into the electronic device.
[0140] In some embodiments, a computer program is also provided, comprising: instructions that, when executed by a processor, cause the processor to perform the interactive method of any of the above embodiments. For example, the instructions may be embodied in computer program code.
[0141] In embodiments of this disclosure, computer program code for performing the operations of this disclosure can be written in one or more programming languages or a combination thereof. These programming languages include, but are not limited to, object-oriented programming languages such as Java, Smalltalk, and C++, as well as conventional procedural programming languages such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network (including a local area network (LAN) or a wide area network (WAN)), or it can be connected to an external computer (e.g., via the Internet using an Internet service provider).
[0142] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.
[0143] The functions described above can be performed, at least in part, by one or more hardware logic components. For example, without limitation, exemplary hardware logic components that can be used include: Field Programmable Gate Arrays (FPGAs), Application-Specific Integrated Circuits (ASICs), Application Standard Products (ASSPs), System-on-Chip (SoCs), Complex Programmable Logic Devices (CPLDs), and so on.
[0144] While specific embodiments of this disclosure have been described in detail by way of example, those skilled in the art should understand that the examples are for illustrative purposes only and not intended to limit the scope of this disclosure. Those skilled in the art should understand that modifications can be made to the above embodiments without departing from the scope and spirit of this disclosure. The scope of this disclosure is defined by the appended claims.
Claims
1. An interaction method, comprising: In response to receiving user input, determine the user's interaction intent; Based on the interaction intent, multiple tasks and their execution order are determined; For each of the multiple tasks, determine the corresponding intelligent agent for each task to obtain a combination of multiple intelligent agents, wherein different types of intelligent agents perform different types of tasks; According to the execution order, the multiple agents are scheduled to perform the multiple tasks in combination.
2. The interaction method of claim 1, wherein, For each of the plurality of tasks, determining the corresponding intelligent agent for each task includes: Determine the type and requirements of each task; Identify one or more candidate agents that match the type of each task; Based on the requirements of each task, select the agent corresponding to each task from the one or more candidate agents.
3. The interaction method according to claim 1 or 2, wherein, The step of determining multiple tasks and their execution order based on the interaction intent includes: Determine the complex task corresponding to the interaction intent; The complex task is broken down into the multiple tasks; The execution order is determined based on the relationship between the multiple tasks.
4. The interaction method according to any one of claims 1 to 3, further comprising: Based on the input information, task information is provided to at least one of the plurality of intelligent agents.
5. The interaction method according to any one of claims 1 to 4, further comprising: Based on the input information, a base model is determined for at least one of the plurality of agents.
6. The interaction method according to any one of claims 1 to 5, further comprising: If the output of the plurality of agents does not meet the first specified condition, the agent for each task is re-determined to obtain a combination of the re-determined agents; According to the execution order, the newly determined combination of multiple agents is scheduled to re-execute the multiple tasks.
7. The interaction method according to any one of claims 1 to 6, further comprising: In the event of a conflict in the outputs of the multiple agents, the multiple tasks and their execution order are redefined, the agent corresponding to each task is determined, and the combination of the multiple agents is scheduled to execute the multiple tasks.
8. The interaction method according to any one of claims 1 to 7, further comprising: If the outputs of the multiple agents conflict, a first guidance question is output to the user. Based on the user's feedback on the first guidance question, the user's interaction intent is redefined; The process involves re-determining multiple tasks and their execution order, identifying the agent corresponding to each task, and scheduling the combined execution of the multiple tasks by the agents.
9. The interaction method according to any one of claims 1 to 8, further comprising: The outputs of the multiple agents are integrated to obtain the integrated result; Output the integrated result.
10. The interaction method of claim 9, wherein, The output of the integrated result includes: The integration results are then subjected to a consistency check. If the consistency check passes, output the result of passing the consistency check.
11. The interaction method according to claim 10, further comprising: If the consistency check fails, the outputs of the multiple agents are re-integrated, and the consistency check is performed on the re-integrated result.
12. The interaction method according to any one of claims 1 to 11, wherein, The step of responding to receiving user input information and determining the user's interaction intent includes: In response to determining that the input information does not meet the second specified condition, a second guidance question is output to the user; The interaction intent is determined based on the user's feedback on the second guidance question.
13. The interaction method according to any one of claims 1 to 12, wherein, The step of responding to receiving user input information and determining the user's interaction intent includes: In response to receiving the input information, determine the type of the input information; The interaction intent is determined based on the type of input information.
14. The interaction method according to any one of claims 1 to 13, wherein, The plurality of intelligent agents includes at least one of the following: Role-playing intelligent agents; Information retrieval intelligent agent; Creative agents, among which different creative agents have different language styles; The intelligent agent that answers questions.
15. The interaction method according to any one of claims 1 to 14, wherein, The multiple intelligent agents exchange information through a shared knowledge base.
16. An interactive device, comprising: The intent determination module is configured to determine the user's interaction intent in response to receiving user input information; The task determination module is configured to determine multiple tasks and their execution order based on the interaction intent; The agent determination module is configured to determine an agent for each of the plurality of tasks, thereby obtaining a combination of the plurality of agents, wherein different types of agents perform different types of tasks; The scheduling module is configured to schedule a combination of the multiple agents to execute the multiple tasks according to the execution order.
17. An interactive device, comprising: Memory; as well as A processor coupled to the memory, the processor being configured to execute the interaction method according to any one of claims 1 to 15 based on instructions stored in the memory.
18. An interactive system, comprising: The interactive device according to claim 16 or 17; as well as Multiple intelligent agents.
19. A computer-readable storage medium having stored thereon computer program instructions that, when executed by a processor, implement the interactive method according to any one of claims 1 to 15.
20. A computer program comprising instructions that, when executed by a processor, cause the processor to perform the interactive method according to any one of claims 1 to 15.
Citation Information
Patent Citations
Intelligent agent scheduling method, system and equipment based on large language model and medium
CN118132227A
Financial task execution method and device, equipment, medium and program product
CN118312599A
Task processing method and device, equipment and computer readable storage medium
CN118626690A
Task execution method and related device, equipment, storage medium and agent platform
CN118689563A
Interaction method and device and computer readable storage medium
CN118964688A