A task execution method and computer device based on cooperation of multiple LLM intelligent agents
Patent Information
- Application Number
- CN202610720331.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-05-22
- Publication Date
- 2026-09-29
AI Technical Summary
[0003]然而,在实际应用中,许多复杂任务(例如,撰写一份包含市场调研、数据分析、图表生成和报告撰写的行业分析报告)难以由一个单一的LLM智能体高效、准确地完成
[0024]在以上的技术方案中,通过允许子LLM智能体在接收到子任务后,基于与其它共同执行同一任务的智能体所共享的“任务上下文”进行推理判断,从而自主地从多种预定义的协作模式中选择最合适的一种(例如,选择直接将结果返回给用户,或者选择将结果返回给主LLM智能体进行整合)。
Smart Images

Figure CN122838009A_ABST
Abstract
Description
Technical Field
[0001] The embodiments in this specification relate to the field of artificial intelligence technology, and in particular to a task execution method and computer device based on the collaboration of multiple LLM agents. Background Technology
[0002] With the rapid development of artificial intelligence technology, Large Language Models (LLMs) have demonstrated powerful natural language understanding and generation capabilities. To apply the capabilities of LLMs to solve more complex real-world problems, a mainstream technical approach is to build LLM-based agents. An LLM agent can be understood as a software entity that uses a large language model as its core decision-making and reasoning engine. It can perceive the environment, plan action steps, and invoke external tools (such as search engines, databases, and application programming interfaces) to execute tasks, ultimately achieving the user's preset goals.
[0003] However, in practical applications, many complex tasks (such as writing an industry analysis report that includes market research, data analysis, chart generation, and report writing) are difficult to complete efficiently and accurately by a single LLM agent. Summary of the Invention
[0004] According to a first aspect of one or more embodiments of this specification, a task execution method based on the collaboration of multiple LLM agents is proposed, applied to a main LLM agent in a service system constructed based on multi-level LLM agents; wherein the multi-level LLM agents include a main LLM agent and multiple sub-LLM agents that have established communication connections with the main LLM agent; different sub-LLM agents are used to execute different sub-tasks; LLM agents in the multi-level LLM agents that execute the same task share the task context associated with that task; the method includes: Obtain task execution information input by the user; Based on the task execution information, the target task to be executed as indicated by the user is determined, and the sub-LLM agents corresponding to the sub-tasks contained in the target task are determined from the plurality of sub-LLM agents. The multiple subtasks are respectively scheduled to corresponding sub-LLM agents. The sub-LLM agents, based on the context related to the target task shared with other LLM agents jointly executing the target task, determine the target task collaboration mode for their corresponding subtasks from multiple task collaboration modes supported by the service system, and execute the subtasks according to the determined target task collaboration mode. The task collaboration mode is used to represent the collaboration method between the sub-LLM agents and the main LLM agent in jointly executing the target task.
[0005] Optionally, the multiple task collaboration modes include a first collaboration mode and a second collaboration mode; The first collaboration mode refers to a task collaboration mode in which the sub-LLM agent independently executes its corresponding sub-task and returns the execution result of the sub-task to the user. The second collaboration mode refers to a task collaboration mode in which the sub-LLM agent returns the execution result of its corresponding sub-task to the main LLM agent, and the main LLM agent integrates the execution result with the execution results of other sub-tasks and returns the integrated execution result to the user.
[0006] Optionally, the multi-level LLM agents may further include target LLM agents that have established a communication connection with the master LLM agent; wherein the target LLM agent is used to perform intent recognition on the task execution information; Based on the task execution information, the target task to be executed as indicated by the user is determined, and the sub-LLM agents corresponding to the sub-tasks included in the target task are determined from the plurality of sub-LLM agents, including: The task execution information is sent to the target LLM agent, which then performs inference calculations on the task execution information to determine the user's intent. Based on the user intent, the target task to be executed as indicated by the user is determined, and structured task scheduling information corresponding to the target task is generated. The structured task scheduling information includes agent identifiers of the sub-LLM agents determined from the plurality of sub-LLM agents, which correspond to the multiple sub-tasks included in the target task. Obtain the structured task scheduling information returned by the target LLM agent, and parse the structured task scheduling information to obtain the agent identifier contained in the structured task scheduling information; Based on the agent identifier, the sub-LLM agents corresponding to the multiple sub-tasks are determined from the multiple sub-LLM agents.
[0007] Optionally, the structured task scheduling information further includes prompt words generated based on the task execution information; wherein, the prompt words are used to schedule the sub-LLM agents corresponding to the plurality of sub-tasks to execute their respective sub-tasks; The multiple subtasks are respectively scheduled to their corresponding sub-LLM agents. Each sub-LLM agent, based on the context related to the target task shared with other LLM agents jointly executing the target task, further determines the target task collaboration mode for its corresponding subtask from among the various task collaboration modes supported by the service system, including: Obtain the prompt words contained in the structured task scheduling information; The prompt words are sent to the sub-LLM agents corresponding to the multiple sub-tasks respectively. The sub-LLM agents respond to the prompt words, further obtain the context related to the target task shared with the other LLM agents that jointly execute the target task, and determine the target task collaboration mode for the corresponding sub-task from the multiple task collaboration modes supported by the service system based on the context.
[0008] Optionally, the structured task scheduling information may also include scheduling reasons for sub-LLM agents corresponding to the multiple sub-tasks, generated by reasoning and calculating the task execution information.
[0009] Optionally, the context related to the target task includes the context related to the user session to which the target task belongs; the prompt word also includes a session identifier of the user session to which the target task belongs; The prompt words are sent to the sub-LLM agents corresponding to the plurality of sub-tasks, respectively. The sub-LLM agents respond to the prompt words by further acquiring context related to the target task, and based on the context, determine the target task collaboration mode for the corresponding sub-task from among the various task collaboration modes supported by the service system, including: The prompt words are sent to the sub-LLM agents corresponding to the multiple sub-tasks respectively. The sub-LLM agents respond to the prompt words, further obtain the context related to the user session corresponding to the session identifier contained in the prompt words, and determine the target task collaboration mode for the corresponding sub-task from the multiple task collaboration modes supported by the service system based on the context.
[0010] Optionally, the multi-level LLM agent includes a master LLM agent, first-level sub-LLM agents that have established communication connections with the master LLM agent, and second-level sub-LLM agents that have established communication connections with at least some of the first-level sub-LLM agents.
[0011] Optionally, the subtasks corresponding to each sub-LLM agent in the first-level sub-LLM agent include subtasks that do not involve modifying the context related to the target task and can be executed in parallel; the subtasks corresponding to each sub-LLM agent in the second-level sub-LLM agent include subtasks that involve modifying the context related to the target task and need to be executed sequentially.
[0012] Optionally, any sub-LLM agent in the multi-level LLM can register with the previous-level sub-LLM agent or the main LLM agent in the form of an MCP service that can be invoked by the previous-level LLM agent.
[0013] According to a second aspect of one or more embodiments of this specification, a task execution method based on the collaboration of multiple LLM agents is also proposed, applied to any target sub-LLM agent in a service system constructed based on multi-level LLM agents; wherein the multi-level LLM agents include a master LLM agent and multiple sub-LLM agents that have established communication connections with the master LLM agent; different sub-LLM agents are used to execute different sub-tasks; LLM agents in the multi-level LLM agents that execute the same task share the task context associated with that task; the method includes: Obtain the target subtask scheduled to the local machine by the main LLM agent; wherein, the target subtask is the subtask corresponding to the target LLM agent determined by the main LLM agent from among the multiple subtasks included in the target task to be executed as indicated by the user; Based on the task context related to the target task shared with other LLM agents jointly executing the target task, a target task collaboration mode is determined for the target sub-task from multiple task collaboration modes supported by the service system; wherein, the task collaboration mode is used to represent the collaboration method in which the sub-LLM agent and the main LLM agent jointly execute the target task; The target sub-tasks are executed according to the determined target task collaboration mode.
[0014] Optionally, the multiple task collaboration modes include a first collaboration mode and a second collaboration mode; The first collaboration mode refers to a task collaboration mode in which the sub-LMM agent independently executes its corresponding sub-task and returns the execution result of the sub-task to the user. The second collaboration mode refers to a task collaboration mode in which the sub-LLM agent returns the execution result of its corresponding sub-task to the main LLM agent, and the main LLM agent integrates the execution result with the execution results of other sub-tasks and returns the integrated execution result to the user.
[0015] Optionally, obtaining the target subtask scheduled locally by the main LLM agent includes: Obtain the prompt word sent by the main LLM agent; wherein, the prompt word is used to schedule the target sub-LLM agent to execute the target sub-task corresponding to it; Based on the task context related to the target task shared with other LLM agents jointly executing the target task, a target task collaboration mode is determined for the target subtask from multiple task collaboration modes supported by the service system, including: In response to the prompt word, further obtain the context related to the target task shared with other LLM agents jointly executing the target task, and determine the target task collaboration mode for the corresponding subtask from multiple task collaboration modes supported by the service system based on the context.
[0016] Optionally, the context related to the target task includes the context related to the user session to which the target task belongs; the prompt word also includes a session identifier of the user session to which the target task belongs; In response to the prompt word, further obtain the context related to the target task, and based on the context, determine the target task collaboration mode for the corresponding subtask from multiple task collaboration modes supported by the service system, including: In response to the prompt word, the context related to the user session corresponding to the session identifier contained in the prompt word is further obtained, and based on the context, the target task collaboration mode is determined for the corresponding subtask from the multiple task collaboration modes supported by the service system.
[0017] Optionally, the multi-level LLM agent includes a master LLM agent, first-level sub-LLM agents that have established communication connections with the master LLM agent, and second-level sub-LLM agents that have established communication connections with at least some of the first-level sub-LLM agents.
[0018] Optionally, the subtasks corresponding to each sub-LLM agent in the first-level sub-LLM agent include subtasks that do not involve modifying the context related to the target task and can be executed in parallel; the subtasks corresponding to each sub-LLM agent in the second-level sub-LLM agent include subtasks that involve modifying the context related to the target task and need to be executed sequentially.
[0019] Optionally, any sub-LLM agent in the multi-level LLM can register with the previous-level sub-LLM agent or the main LLM agent in the form of an MCP service that can be invoked by the previous-level LLM agent.
[0020] Optionally, the target subtask may further include multiple subtasks; Execute the target sub-tasks according to the determined target task collaboration mode, including: If there are multiple sub-LLM agents at the next level that have established communication connections with the target sub-LLM agent, then the next-level sub-LLM agents corresponding to the multiple sub-tasks included in the target sub-task are determined from among the multiple sub-LLM agents at the next level. The target subtask is divided into multiple subtasks, which are respectively scheduled to the corresponding next-level sub-LLM agents. The next-level sub-LLM agents further determine the task collaboration mode for their corresponding subtasks from the multiple task collaboration modes supported by the service system based on the context related to the target subtask shared with other LLM agents that jointly execute the target subtask. The subtasks are then executed according to the determined task collaboration mode.
[0021] According to a third aspect of one or more embodiments of this specification, a service system is also proposed, said service system being a service system built on a multi-level LLM agent, comprising: A master LLM agent for performing the method as described in any one of the first aspects above; Multiple sub-LLM agents that have established communication connections with the main LLM agent are used to execute the method as described in any one of the second aspects above.
[0022] According to a fourth aspect of one or more embodiments of this specification, a computer device is also provided, comprising: a processor; a memory for storing processor-executable instructions; wherein the processor performs the executable instructions to implement the steps of the method as described in any one of the first or second aspects above.
[0023] According to a fifth aspect of one or more embodiments of this specification, a computer program product is also provided, comprising a computer program / instructions that, when executed by a processor, implement the steps of the method as described in any one of the first or second aspects above.
[0024] In the above technical solution, after receiving a sub-task, the sub-LLM agent can make inferences and judgments based on the "task context" shared with other agents that jointly perform the same task, thereby autonomously selecting the most suitable one from a variety of predefined collaboration modes (for example, choosing to return the result directly to the user, or choosing to return the result to the main LLM agent for integration).
[0025] This "on-demand decision-making" mechanism ensures that the execution path of each subtask matches its own characteristics: for independent subtasks, the main LLM agent can be bypassed, enabling a "fast track" direct response, thereby effectively reducing communication latency within the system, alleviating the computational load of the main LLM agent, and accelerating the user's response time to the final answer; for subtasks that need to be integrated, a standard path can be selected to ensure the integrity and logical coherence of the final output information. Therefore, this scheme optimizes the system's resource scheduling as a whole, improving the parallelism and execution efficiency of task processing. Attached Figure Description
[0026] To more clearly illustrate the technical solutions of the embodiments in this specification, the drawings used in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments recorded in this specification. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0027] Figure 1 This is a schematic diagram of the architecture of a service system based on a multi-level LLM agent, as shown in one embodiment of this specification. Figure 2 This is a flowchart illustrating a task execution method based on the collaboration of multiple LLM agents in one embodiment of this specification; Figure 3 This is a flowchart illustrating another task execution method based on the collaboration of multiple LLM agents, as shown in one embodiment of this specification. Figure 4 This is a schematic structural diagram of an electronic device shown in one embodiment of this specification; Figure 5 This is a block diagram of a task execution device based on the collaboration of multiple LLM agents, as shown in one embodiment of this specification. Figure 6This is a block diagram of a task execution device based on the collaboration of multiple LLM agents, as shown in one embodiment of this specification. Detailed Implementation
[0028] Exemplary embodiments will now be described in detail, examples of which are illustrated in the accompanying drawings. When the following description relates to the drawings, unless otherwise indicated, the same numerals in different drawings denote the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with one or more embodiments of this specification. Rather, they are merely examples of apparatuses and methods consistent with some aspects of one or more embodiments of this specification as detailed in the appended claims.
[0029] It should be noted that the steps of the corresponding methods are not necessarily performed in the order shown and described in this specification in other embodiments. In some other embodiments, the methods may include more or fewer steps than described in this specification. Furthermore, a single step described in this specification may be broken down into multiple steps in other embodiments; and multiple steps described in this specification may be combined into a single step in other embodiments.
[0030] The user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, data stored, data displayed, etc.) involved in this manual are all information and data authorized by the user or fully authorized by all parties. The collection, use and processing of related data shall comply with the relevant laws, regulations and standards of the relevant countries and regions, and corresponding operation portals shall be provided for users to choose to authorize or refuse.
[0031] In related technologies, in order to enable LLM agents to handle complex tasks consisting of multiple sub-tasks, a technical architecture in which multiple LLM agents collaborate to complete the task has been proposed.
[0032] In these architectures, a master LLM agent is typically designated, which is responsible for breaking down the complex task input by the user into multiple relatively independent subtasks. These subtasks are then distributed to different sub-LLM agents with specific specializations for execution. Finally, the master agent collects the execution results from each sub-LLM agent, summarizes and integrates them, and returns the final answer to the user.
[0033] However, the existing multi-LLM agent collaboration methods described above have significant limitations. The most critical issue lies in the relative complexity of their collaboration models.
[0034] Specifically, in these existing technical architectures, the interaction logic between all sub-LLM agents and the master agent is pre-defined and fixed. That is, the sub-LLM agents always return the execution results to the master agent, which then processes them and feeds them back to the user.
[0035] This centralized, mandatory reporting model is effective when dealing with tasks that require the results of multiple subtasks to form the final answer, but it cannot adapt to all types of subtasks.
[0036] For example, when a subtask is to query current weather information, the result of this subtask is an independent answer that can be directly presented to the user. In this case, the path of returning the result to the master agent first and then forwarding it to the user becomes redundant, increasing system processing latency and the computational burden on the master agent. Conversely, when a subtask is to calculate the average of a set of data, and this average needs to be compared with the results of other subtasks (such as the total number of data points) to form a meaningful conclusion, if the sub-LLM agent directly returns the average to the user, the user will receive incomplete information lacking context.
[0037] It is evident that this single, fixed task collaboration mode in related technologies cannot dynamically and flexibly select appropriate task collaboration methods based on the characteristics of the subtasks themselves and their roles in the global task. This results in low efficiency and insufficient resource utilization when the system handles complex tasks containing multiple types of subtasks, and the information ultimately presented to the user may lack necessary integration and coherence.
[0038] Therefore, how to provide a task execution mechanism that enables more flexible and efficient collaboration among multiple LLM agents has become a technical problem that urgently needs to be solved in this field.
[0039] In view of this, this application provides a task execution method based on the collaboration of multiple LLM agents. This method constructs a multi-level LLM agent architecture and introduces an autonomous decision-making mechanism for collaboration modes based on shared task context, enabling sub-LLM agents to select the most appropriate way to interact with the main LLM agent and the user from a variety of predefined task collaboration modes according to the specific circumstances of the sub-tasks they are executing.
[0040] Please see Figure 1 , Figure 1 This is an exemplary embodiment of a service system architecture based on multi-level LLM agents.
[0041] like Figure 1As shown, the service system provided in this specification can be based on a multi-level, hierarchical LLM agent network. This network is designed with full consideration of task decomposition, agent functional specialization, and overall system scalability.
[0042] Specifically, the system includes at least one master LLM agent (also known as the root agent, coordinating agent, or overall control agent), and one or more sub-LLM agents that have established direct communication connections with the master LLM agent. These sub-LLM agents can further have their own subordinate sub-LLM agents, thus forming a deeper hierarchical tree or forest structure. This multi-level structure allows the system to recursively decompose extremely complex tasks into smaller, more manageable subtasks, which are then executed by highly specialized lower-level agents.
[0043] In some embodiments, the system can be deployed entirely on a cloud server cluster.
[0044] The master LLM agent operates as the central control node of the entire system, responsible for receiving all user requests, maintaining the global task state, and coordinating cross-agent communication. Individual sub-LLM agents can run on the same physical server to reduce communication latency, or they can run distributed across different containers, virtual machines, or even different physical data centers, depending on their functional roles and load requirements, to improve system fault tolerance and resource utilization. Data interaction between different agents can be achieved through standard network communication protocols, such as HTTP / 2 for synchronous request-response interaction, or gRPC for high-performance streaming data transmission.
[0045] To efficiently and dynamically manage these distributed, functionally diverse sub-LLM agents, this specification introduces a service-oriented architecture based on MCP (Model Capability Proxy) service. The main function of this service is to encapsulate and abstract the underlying, specific artificial intelligence capabilities, forming a standardized service interface that can be remotely invoked by upper-layer agents.
[0046] Specifically, an MCP service can encapsulate various types of underlying AI capabilities, such as: Intent recognition capability: It can identify the user's true intent (whether it is to query information, perform an operation, or create content) from the text entered by the user.
[0047] Information extraction capability: It can extract structured entities, relationships and attributes from unstructured text (e.g., extract "Zhang San", "buy", and "three books" from "Zhang San bought three books yesterday").
[0048] Text classification capabilities: It can perform sentiment classification, topic classification, etc. on the input text.
[0049] Numerical computation capability: capable of performing precise mathematical operations or statistical analysis.
[0050] Image understanding ability: Ability to identify objects, scenes, or text in images.
[0051] exist Figure 1 In the service system shown, any sub-LLM agent can encapsulate its core functionality into one or more MCP services and register them with its parent LLM agent. The "parent agent" here could be the main LLM agent or another higher-level sub-LLM agent. The registration process can be completed through a service registry or via a direct point-to-point registration method.
[0052] The registration information typically includes: service name, service capability description, calling endpoint (such as a URL or RPC address), JSON schema definition of input parameters, format description of output results, and service version number and health check address, etc.
[0053] For example, consider a specific "weather query sub-LLM agent." Suppose its core capability is to call a third-party weather API and return formatted weather information. To allow other agents in the system (such as the main LLM agent) to call it, it can encapsulate its capabilities into an MCP service called "WeatherQuery." This service accepts a parameter "city_name" and returns a JSON object containing fields such as temperature, humidity, and weather conditions. Subsequently, the sub-LLM agent sends a registration request to the main LLM agent, reporting all metadata about its MCP service. Upon receiving this, the main LLM agent records the service in its local service directory.
[0054] When the main LLM agent processes a user request, if it needs to query the weather, it can search the service directory and find an MCP service named "WeatherQuery". It knows that calling this service requires a "city_name" parameter and the data structure returned upon successful execution. Therefore, the main LLM agent can send a request containing "city_name": "Beijing" to the endpoint of this MCP service over the network, just like calling a regular local function, to obtain weather data asynchronously or synchronously. From the main LLM agent's perspective, it doesn't need to know how the "weather query sub-LLM agent" is implemented internally (which API it calls, how it parses the data); it only needs to interact with it through the standardized MCP service interface.
[0055] This MCP-based encapsulation and registration mechanism enables the entire system to have "plug-and-play" scalability. For example, new sub-LLM agents can be developed and registered into the system at any time, while the replaced or upgraded sub-LLM agents only need to ensure that their MCP service interface contracts remain unchanged, without affecting the upper-layer callers.
[0056] exist Figure 1 In the service system shown, all LLM agents scheduled to jointly execute the same target task, whether they are the main LLM agent or a sub-LLM agent at a certain level, can access a shared context space bound to that task. This shared context space can be implemented using various techniques. For example, a memory-based distributed cache (such as a Redis cluster) can be maintained internally, storing all state information related to the task using "task ID" or "session ID" as the key. This information can include: Conversation history: A complete record of the dialogue between the user and the system, including each instruction entered by the user and each response returned by the system.
[0057] Intermediate calculation results: Intermediate data generated by each sub-LLM agent during the execution of sub-tasks that may be used by other sub-LLM agents. For example, a raw data table retrieved from a database by a "data extraction sub-LLM agent" can be directly read by a subsequent "data analysis sub-LLM agent".
[0058] User preferences: Preference information extracted from user history or explicit settings, such as preferred language style, output format (e.g., tables, paragraphs, bullet points), level of detail, etc.
[0059] The status of executed subtasks: records which subtasks have been completed, which are currently being executed, which have failed and need to be retried, and a summary of the execution results of each subtask. This helps the agent determine dependencies later.
[0060] Global variables: Variables set by the master LLM agent or planning agent to control the task flow, such as "whether the current task has been canceled" or "whether to request fast response mode".
[0061] The implementation of this sharing mechanism typically relies on a "context manager" component. This component is responsible for creating the context space at the start of the task, assigning access permissions (usually read-write or read-only permissions) to each agent participating in the task, and destroying the space to release resources at the end of the task.
[0062] When a child LLM agent is scheduled to execute a subtask, the prompt it receives contains a unique "task identifier" or "session identifier." The child LLM agent can then use this identifier to request the shared context manager to obtain the necessary shared context information. Because all agents use the same identifier to access the same shared data, the context they see is always consistent and real-time.
[0063] This sharing mechanism is the basis for the sub-LLM agent to make autonomous decisions, because it enables the sub-LLM agent to "perceive" the work progress and output of other collaborating agents, thereby determining which collaboration mode it should adopt (e.g., whether it needs to wait for the results of other agents, or whether its own results need to be integrated).
[0064] The above is a detailed description of the system architecture of the service system used in this specification. Those skilled in the art will understand that the above descriptions of the deployment method, the encapsulation details of the MCP service, the implementation method of the shared context, and the interaction process are all exemplary. Without departing from the core inventive concept of this application, other equivalent technical means can be used to achieve similar functions according to the actual application scenario.
[0065] The technical solution of this specification will be described in detail below with reference to the accompanying drawings.
[0066] Please see Figure 2 , Figure 2 This specification illustrates a flowchart of a task execution method based on the collaboration of multiple LLM agents, which can be applied to, for example... Figure 1The service system shown is based on a multi-level LLM agent and includes a main LLM agent and multiple sub-LLM agents that have established communication connections with the main LLM agent. Different sub-LLM agents are used to execute different sub-tasks. LLM agents performing the same task share the task context associated with that task. The execution process includes the following: Step 202: Obtain the task execution information input by the user; In some embodiments, the master LLM agent can obtain task execution information input by the user.
[0067] The task execution information can take many forms; For example, the task execution information could be a text command based on natural language entered by the user through a text input box in a graphical user interface (GUI); or, the task execution information could be a voice command entered by the user, which the system can convert into text and then pass to the main LLM agent.
[0068] In addition, the task execution information can also be a structured data object generated by the system based on user actions performed by the user in the interface. For example, after detecting a user action (such as clicking a button) in the interface, the system can respond to the user action by triggering a call to the backend API to automatically generate a structured data object (such as an API call request) as the aforementioned task execution information. In this case, the structured data object can contain the task type and related parameters.
[0069] Step 204: Based on the task execution information, determine the target task to be executed as indicated by the user, and determine the sub-LLM agents corresponding to the multiple sub-tasks contained in the target task from the multiple sub-LLM agents; After obtaining the user's task execution information, the main LLM agent needs to parse the information to determine what the user really wants to execute, what the target task can be broken down into more granular subtasks, and identify the sub-LLM agents in the system that can execute these subtasks.
[0070] In some embodiments, a dedicated target LLM agent (or planning agent) may be introduced to identify the intent of task execution information.
[0071] In this way, in addition to the main LLM agent, the service system can deploy one or more target LLM agents that have established communication connections with the main LLM agent. These target LLM agents can be specially trained agents skilled in logical reasoning and task planning.
[0072] Specifically, after receiving the user's task execution information (such as "Please help me book a flight from Beijing to Shanghai tomorrow morning and notify my assistant"), the master LLM agent does not break it down itself, but sends it to the target LLM agent as is or after simple preprocessing through an API call.
[0073] In practical applications, to help the target LLM agent better understand the task, the master LLM agent can also attach some system instructions when sending task execution information; For example, in one instance, task execution information accompanied by system instructions could specifically be in the form of text instructions as follows: "Please analyze the following user request, break it down into a series of executable subtasks, assign the most suitable sub-LLM agent to each subtask, and finally return the results in JSON format." After receiving task execution information and system instructions, the target LLM agent will use its internal large language model to perform reasoning calculations to determine the user's intent and, based on the user's intent, determine the target task to be executed as instructed by the user.
[0074] For example, in one instance, the aforementioned target LLM agent might generate the following chain of thought: "The user's intent is to complete a composite task that includes the actions of 'booking a flight' and 'sending a notification'. Therefore, the target task can be named 'Complete flight booking and notification'."
[0075] This objective task comprises two sub-tasks: Subtask 1: Book a flight from Beijing to Shanghai for tomorrow morning; Subtask 2: Notify the assistant.
[0076] The system contains a 'flight booking sub-LLM agent' and a 'message notification sub-LLM agent', which can execute these two subtasks respectively. Therefore, I will assign 'flight_booking_agent' to subtask 1 and 'notification_agent' to subtask 2. Based on the above reasoning, the target LLM agent can generate structured task scheduling information. This structured task scheduling information includes agent identifiers of the sub-LLM agents identified from multiple sub-LLM agents, each corresponding to one of the sub-tasks included in the target task. In practical applications, this structured information can be in a format easily parsed by programs; for example, it can be in JSON format.
[0077] The structured task scheduling information mentioned above may include agent identifiers. Specifically, the agent identifier refers to the agent identifier of the sub-LLM agent that is determined from multiple sub-LLM agents and corresponds to the multiple sub-tasks included in the target task. In practical applications, the agent identifier is specifically used to clearly indicate which sub-LLM agent is responsible for which sub-task.
[0078] Furthermore, in some embodiments, this structured information may also include a "reason" field, recording the logic behind the scheduling decision. While this reason field is not required at runtime, it is valuable for system debugging, decision auditing, and subsequent model optimization.
[0079] For example, in practical applications, including "scheduling reasons" in scheduling information, although it does not directly affect task execution, provides valuable data support for system debugging, monitoring, and post-event auditing of scheduling decisions, thereby enhancing the interpretability and maintainability of the system.
[0080] The target LLM agent can return the structured task scheduling information to the master LLM agent. Upon receiving the structured task scheduling information from the target LLM agent, the master LLM agent can parse the structured task scheduling information to obtain the agent identifier contained within it.
[0081] For example, taking the above-mentioned structured task scheduling information as a JSON data structure, the main LLM agent can call a JSON parser to extract the agent identifier contained in the JSON data structure.
[0082] Then, the main LLM agent can determine the sub-LLM agents corresponding to the aforementioned sub-tasks from among the multiple sub-LLM agents based on the acquired agent identifier.
[0083] For example, in some cases, the master LLM agent can maintain a locally or globally accessible "agent directory" that records information such as identifiers (IDs), calling endpoints, and capability descriptions of all registered child LLM agents. The master LLM agent can then search this directory using the agent identifiers extracted from the structured information to determine the specific child LLM agent instance corresponding to each subtask.
[0084] Of course, in addition to the above-mentioned method of using a dedicated target LLM agent, in some embodiments, the main LLM agent may also not rely on the external target LLM agent, but instead use its own large language model capabilities to perform task decomposition and agent matching.
[0085] For example, in some cases, the primary LLM agent can load a system prompt word specifically for task planning, putting it into "planning mode." In this mode, the primary LLM agent directly infers from user input and generates the structured task scheduling information described above.
[0086] The advantage of this approach is reduced network call latency, while the disadvantages are increased computational burden on the main LLM agent and potentially lower planning capabilities compared to a specially optimized target LLM agent. Those skilled in the art can flexibly choose or combine these two approaches based on the actual system's load and performance requirements.
[0087] Step 206: The multiple subtasks are scheduled to their corresponding sub-LLM agents. Each sub-LLM agent, based on the context related to the target task shared with other LLM agents jointly executing the target task, determines a target task collaboration mode from among the various task collaboration modes supported by the service system, and executes the subtask according to the determined target task collaboration mode. The task collaboration mode represents the collaborative method between the sub-LLM agent and the main LLM agent in jointly executing the target task. After determining the correspondence between subtasks and sub-LLM agents, the main LLM agent can schedule each subtask to its corresponding sub-LLM agent.
[0088] In some embodiments, the structured task scheduling information may further include prompt words generated based on the task execution information; wherein, the prompt words may be used to schedule the sub-LLM agents corresponding to the plurality of sub-tasks to execute their respective sub-tasks.
[0089] The master LLM agent can retrieve the prompts generated for each subtask from the structured task scheduling information. These prompts are then encapsulated into network messages (e.g., an HTTP POST request) and sent to the corresponding child LLM agents.
[0090] Each child LLM agent that receives the prompt does not immediately execute the subtask assigned to it and, by default, return the result to the main LLM agent. Instead, they first access the shared task context bound to the current target task. This shared context can be indexed by a task identifier or a session identifier. For example, the main LLM agent can append a "session_id" or "task_id" field to the prompt it sends.
[0091] Based on information read from the shared context (e.g., whether there are results of other subtasks, whether there are global cooperation mode instructions, characteristics of its own task, etc.), the sub-LLM agent autonomously determines the most suitable target cooperation mode for the subtask it is responsible for from among the various task cooperation modes supported by the service system.
[0092] In some embodiments, the service system described above may support multiple task collaboration modes, which represent the collaborative methods by which the sub-LLM agent and the main LLM agent jointly execute the target task.
[0093] The service system supports multiple task collaboration modes, including at least a first collaboration mode and a second collaboration mode.
[0094] The first collaborative mode mentioned above refers to a task collaboration mode in which the sub-LLM agent independently executes its corresponding sub-task and returns the execution result of the sub-task to the user.
[0095] In this first collaborative mode, the sub-LLM agent can independently execute its corresponding subtask and return the execution result directly to the user after completion, without going through the main LLM agent.
[0096] This mode is suitable for subtasks whose execution results are independent and complete, and which can be presented directly to the user without needing to be integrated with the results of other subtasks.
[0097] For example, these subtasks could be tasks such as "checking the weather" or simple calculations, the results of which can be directly displayed to the user. By bypassing the main LLM agent, one network relay can be reduced, the load on the main LLM agent can be reduced, and the response speed can be increased.
[0098] The aforementioned second collaboration mode refers to a task collaboration mode in which the sub-LLM agent returns the execution result of its corresponding sub-task to the main LLM agent, and the main LLM agent integrates the execution result with the execution results of other sub-tasks, and returns the integrated execution result to the user.
[0099] In this second collaborative mode, after the sub-LLM agent completes its subtask, it does not return the result directly to the user, but instead returns the execution result to the main LLM agent. The main LLM agent is responsible for collecting the execution results of all (or some) subtasks, integrating, sorting, deduplicating, or formatting these results, and finally returning the integrated complete result to the user.
[0100] This model is suitable for scenarios where the results of subtasks are only part of the final answer, or need to be compared and integrated with the results of other subtasks to produce a complete meaning.
[0101] For example, in the aforementioned task of "booking a flight and notifying the assistant," the results of the two subtasks (an order number confirming a successful booking and a notification status indicating successful delivery) are not particularly meaningful to the user individually. The user expects a final, integrated response: "The flight has been successfully booked, and the notification has been sent." In this case, both sub-LLM agents should select the second collaboration mode, returning their respective results to the main LLM agent, which will then combine them into a single sentence and inform the user.
[0102] In some embodiments, the context associated with the target task may include the context associated with the user session to which the target task belongs. This means that the system supports not only context sharing for a single task, but also session-level context sharing across multiple turns of dialogue. The prompt may also include a session identifier associated with the user session to which the target task belongs; for example, the session identifier may be a SessionID.
[0103] For example, in the first round of conversation, a user might say, "Help me find an Italian restaurant near me." The main LLM agent scheduled a "POI query sub-LLM agent" to perform the task, which returned a list of restaurants. The system stored the context of this round of dialogue (including the user's location information, query results, etc.) in a shared context and associated it with a session identifier (session_123).
[0104] Then, in the second round of conversation, the user said, "Book a table at the first restaurant tonight at 7 pm." After receiving new task execution information, the master LLM agent will include the same session identifier "session_123" in the prompt when scheduling the "restaurant reservation sub-LLM agent".
[0105] Once the restaurant reservation agent receives the prompt, it accesses the shared context through the session identifier, allowing it to retrieve the specific restaurant information (such as restaurant name, address, and phone number) determined in the first round of conversation, thus successfully completing the reservation. Without this shared context, the user would need to explicitly state the restaurant name again in the second round of conversation, significantly diminishing the experience.
[0106] In this scenario, after the master LLM agent sends the aforementioned prompt to the sub-LLM agents, each sub-LLM agent can respond to the prompt by further obtaining the context related to the user session corresponding to the session identifier contained in the prompt, and based on the obtained context, determine the target task collaboration mode for its corresponding sub-task from among the various task collaboration modes supported by the service system.
[0107] For example, in practical applications, each sub-LLM agent can query the shared context from the context subsystem based on the session identifier and its own agent identifier. The context subsystem can then query the agent's capability description based on the agent identifier and search for the context related to the agent's capability description from the context of the relevant session.
[0108] In some embodiments, after each sub-LLM agent determines the target task collaboration mode for its corresponding sub-task from the multiple task collaboration modes supported by the service system, it can execute the corresponding sub-task according to the determined target task collaboration mode.
[0109] For example, if any LLM agent determines the target task collaboration mode for its corresponding subtask as the first collaboration mode mentioned above, then the sub-LLM agent can independently execute its corresponding subtask and directly return the execution result of the subtask to the user.
[0110] If any LLM agent determines the target task collaboration mode for its corresponding subtask to be the second collaboration mode mentioned above, then the sub-LLM agent can directly return the execution result of the subtask to the main LLM agent. The main LLM agent can then receive the execution results returned by each sub-LLM agent and, according to a preset integration logic, organize these scattered results into a coherent, fluent, and complete response, which is then returned to the user.
[0111] The aforementioned integration logic can be flexibly configured based on actual needs; for example, another dedicated "result integration LLM agent" can be invoked to complete the result integration; or, the main LLM agent itself can complete the result integration by being guided by prompts.
[0112] Of course, if all sub-LLM agents determine the first collaboration mode for the target task for their corresponding sub-tasks, the main LLM agent does not need to participate in the subsequent execution of the target task. In this case, the user will directly receive multiple parallel responses.
[0113] For example, the subtasks included in the aforementioned target task might be independently executable tasks that do not require interaction with other subtasks (such as the weather query task mentioned earlier). After each subtask executes its corresponding subtask based on the first collaborative mode, the execution results of these subtasks can be returned to the user in parallel without interfering with each other. During this process, the sub-LLM agents do not need to interact with the main LLM agent in any way.
[0114] In some embodiments, the above-described service system can be designed as a multi-level nested architecture.
[0115] Please continue reading Figure 1 In addition to the main LLM agent and the first-level sub-LLM agents that have established communication connections with the main LLM agent, the above service system may also include, in practical applications, a second-level sub-LLM agent that has established communication connections with at least some of the sub-LLM agents in the first-level sub-LLM agent.
[0116] It is important to emphasize that, theoretically, this hierarchy can be extended indefinitely to form a deep tree structure.
[0117] For example, in practical applications, the above service system may also include a third-level sub-LLM agent that has established a communication connection with at least some of the sub-LLM agents in the second-level sub-LLM agent, and so on. In practical applications, the number and depth of the sub-LLM agents included in the above service system can be customized based on specific needs.
[0118] This requires an instruction manual. This multi-layered nested structure can also be associated with the characteristics of the tasks undertaken by each sub-LLM agent, thereby optimizing the execution efficiency of the entire system.
[0119] For example, in some embodiments, the subtasks assigned to the first-level sub-LLM agents can typically be tasks that do not involve modifying the shared context and can be executed in parallel.
[0120] These types of tasks are typically stateless, meaning that the execution result does not depend on the task's historical state, nor does it change the state required by subsequent tasks.
[0121] For example, data can be retrieved in parallel from multiple different data sources (checking the weather, stock prices, and news). These operations do not interfere with each other and can be performed simultaneously. Placing them at the first level and scheduling them all out at once by the main LLM agent can fully utilize parallel computing capabilities and significantly shorten the overall task response time.
[0122] Subtasks assigned to second-level (or deeper-level) sub-LLM agents are typically those that involve modifying the shared context and need to be executed in a strict order.
[0123] These types of tasks are stateful, meaning that the execution of a subsequent task depends on the modifications made to the shared context by the previous task. For example, in an "e-commerce order placement" task, the "create order" subtask updates the "order status" in the shared context to "pending payment" and writes the order number; the "initiate payment" subtask must read this order number to make the payment; and the "confirm inventory" subtask needs to deduct the inventory after successful payment.
[0124] These subtasks typically need to be executed sequentially, and ideally by dedicated agents at the same level or within a deep processing pipeline, to ensure data consistency and logical correctness. Delegating these dependent tasks to a second or deeper level allows the main LLM agent to focus on macro-level scheduling, while delegating complex process control to lower-level agents, thus reducing the complexity of the main LLM agent.
[0125] In some embodiments, to support such a multi-level, dynamic LLM agent network, a registration mechanism based on MCP (Model Capability Agent) service can also be used to register and manage the LLM agents in the above service system.
[0126] In this scenario, any sub-LLM agent in the aforementioned multi-level LLM agent is registered to the previous-level sub-LLM agent or the main LLM agent in the form of an MCP service that can be invoked by the previous-level LLM agent.
[0127] Specifically, MCP services are a standardized encapsulation of underlying AI capabilities (such as intent recognition, information extraction, and calling specific APIs). When a sub-LLM agent starts up, it registers its function description, API calls, input / output formats, and other information as an MCP service in the service directory of its parent agent.
[0128] For example, an "Address Resolution Sub-LLM Agent" can be encapsulated as an MCP service called "AddressParser," which accepts an address string and returns longitude and latitude. It might be registered with a "Map Service Sub-LLM Agent" (level 1), which in turn registers with the main LLM Agent. When the main LLM Agent needs to resolve an address, it simply calls the MCP service provided by the "Map Service Sub-LLM Agent," which in turn calls the "Address Resolution Sub-LLM Agent."
[0129] This chain-like registration and invocation mechanism makes the system's hierarchical structure transparent to the upper layers, and each layer only needs to care about the service contract of its immediate next layer, which greatly simplifies the complexity and maintenance cost of the system.
[0130] Please see Figure 3 , Figure 3 This specification illustrates a flowchart of a task execution method based on the collaboration of multiple LLM agents, which can be applied to, for example... Figure 1 The illustrated example is any target sub-LLM agent in a service system built upon a multi-level LLM agent architecture. The multi-level LLM agent architecture includes a master LLM agent and multiple sub-LLM agents that have established communication connections with the master LLM agent. Different sub-LLM agents are used to execute different sub-tasks. LLM agents within the multi-level LLM agent architecture that execute the same task share the task context associated with that task. The execution process includes the following: Step 302: Obtain the target subtask scheduled to the local machine by the main LLM agent; wherein, the target subtask is the subtask corresponding to the target sub-LLM agent determined by the main LLM agent from among the multiple subtasks included in the target task to be executed as indicated by the user. Step 304: Based on the task context related to the target task shared with other LLM agents jointly executing the target task, determine a target task collaboration mode for the target sub-task from multiple task collaboration modes supported by the service system; wherein, the task collaboration mode is used to represent the collaboration method between the sub-LLM agent and the main LLM agent in jointly executing the target task; Step 306: Execute the target sub-task according to the determined target task collaboration mode.
[0131] It should be noted that steps 202-206 are... Figure 1 The illustrated service system employs a multi-layered LLM agent architecture, with the main LLM agent serving as the execution entity. Steps 302-306 involve... Figure 1The illustrated service system employs a specific sub-LLM agent within a multi-layered LLM agent architecture as the execution entity. Therefore, the implementation details related to steps 302-306 will not be repeated in this embodiment; please refer to the embodiments corresponding to steps 202-206 for details.
[0132] In the above embodiments, by constructing a multi-level LLM agent architecture and introducing an autonomous decision-making mechanism based on a shared context cooperative mode, various technical effects are brought about for multi-agent cooperative execution of complex tasks: First, it significantly improves the flexibility and efficiency of multi-LLM agent collaborative task execution.
[0133] Specifically, in traditional schemes, all sub-LLM agents adopt a single "report to the main LLM agent" collaboration mode, which results in unnecessary communication relays and processing overhead for independent subtasks that can directly return to the user (such as querying the weather).
[0134] The above technical solutions break this rigid pattern. By allowing sub-LLM agents to reason and make judgments based on the "task context" shared with other agents performing the same task after receiving a sub-task, they can autonomously choose the most suitable one from a variety of predefined collaboration modes (e.g., choosing to return the result directly to the user, or choosing to return the result to the main LLM agent for integration).
[0135] This "on-demand decision-making" mechanism ensures that the execution path of each subtask matches its own characteristics: for independent subtasks, the main LLM agent can be bypassed, enabling a "fast track" direct response, thereby effectively reducing communication latency within the system, alleviating the computational load of the main LLM agent, and accelerating the user's response time to the final answer; while for subtasks that need to be integrated, a standard path can be selected to ensure the integrity and logical coherence of the final output information. Therefore, this scheme optimizes the system's resource scheduling as a whole, improving the parallelism and execution efficiency of task processing.
[0136] Second, with extremely low implementation costs, two core collaboration paradigms are clearly defined, providing a clear and operable choice space for the autonomous decision-making of sub-LLM agents.
[0137] By defining a first collaboration mode of "directly returning the results to the user" and a second collaboration mode of "returning the results to the main LLM agent for integration", this application ingeniously standardizes the two most essential interaction logics in multi-agent collaboration.
[0138] On the one hand, it provides a clear binary choice framework for the decision engine of the sub-LLM agent, reducing the complexity of making correct decisions. On the other hand, these two modes cover the collaborative needs of subtasks in most complex tasks, eliminating the need for the system to pre-define complex and personalized collaborative processes for each specific type of subtask. This achieves both system flexibility and the simplicity and scalability of the overall architecture.
[0139] Third, by introducing a target LLM agent specifically designed for intent recognition, task decoupling and task execution are decoupled, improving the accuracy and robustness of the system when handling complex tasks.
[0140] Specifically, the complex computational process of "determining the target task and its corresponding sub-LLM agent" is extracted from the main LLM agent and delegated to an independent target LLM agent (which can be understood as an agent specifically responsible for planning and scheduling). This target LLM agent provides the main LLM agent with clear and directly parsable scheduling instructions by generating structured task scheduling information (including agent identifiers).
[0141] This decoupled design simplifies the primary LLM agent's role to "execution scheduling," allowing it to focus more on interaction with sub-LLM agents and state management, thus improving operational stability. Furthermore, since intent recognition and task planning are handled by a separate, potentially specially optimized agent, the accuracy of its understanding of user intent and the rationality of task decomposition are enhanced, reducing subsequent execution failures or inefficiencies caused by planning errors. Ultimately, the system's ability to handle complex and ambiguous user commands is strengthened.
[0142] Fourth, by associating prompt words with session identifiers, it ensures that the sub-LLM agents can accurately obtain the correct shared task context, thus guaranteeing information consistency during multi-agent collaboration.
[0143] In a complex service scenario involving multi-turn dialogues or concurrent execution of multiple tasks, it is crucial to ensure that different agents access the correct context of the "same task".
[0144] In the above technical solution, by carrying a "session identifier" in the scheduling prompt, a precise index key is provided for the child LLM agent. When the child LLM agent responds to the prompt, it can accurately locate and load the context information related to that specific task session from the system's shared storage or memory based on the session identifier. This effectively avoids execution errors caused by context confusion (e.g., incorrectly using information from the first round of dialogue in the second round).
[0145] Fifth, by constructing a multi-level (especially two-level) LLM agent structure and clarifying the task characteristics of agents at different levels, the system achieves fine-grained decomposition and pipelined processing of complex tasks, greatly improving the system's scalability and processing depth.
[0146] First, by constructing an infinitely scalable tree-like or hierarchical agent collaboration network, the task characteristics of agents at different levels can be distinguished: The first-level (high-level) sub-LLM agents are responsible for executing stateless, parallelizable subtasks, which helps to make full use of distributed computing resources and maximize parallel processing efficiency. The second-level (lower-level) sub-LLM agents are responsible for executing stateful subtasks that need to be executed sequentially. This ensures that tasks that depend on the results of preceding steps or require modification of the shared context can be completed correctly and in an orderly manner.
[0147] Secondly, by further endowing any sub-LLM agent with the ability to recursively call its subordinate agents, it can act as the "main LLM agent" and perform secondary scheduling when faced with a complex task that can itself be decomposed into multiple subtasks. This recursive decomposition capability enables the technical solution of this application to handle tasks of any complexity without fundamentally changing the system architecture.
[0148] In addition, the use of MCP (Model Context Protocol) services for registration provides a standardized technical foundation for building a dynamic, pluggable intelligent agent ecosystem. New sub-LLM agents with specific capabilities can register with the system at any time via MCP services and be discovered and invoked by higher-level agents, which greatly enhances the system's openness and scalability.
[0149] Corresponding to the embodiments of the foregoing methods, this specification also provides embodiments of apparatus, electronic devices, and storage media.
[0150] Figure 4 This is a schematic structural diagram of an electronic device provided in an exemplary embodiment. Please refer to... Figure 4At the hardware level, the device includes a processor 402, an internal bus 404, a network interface 406, memory 408, and non-volatile memory 410, and may also include other necessary hardware. One or more embodiments of this specification can be implemented in software, for example, the processor 402 reads the corresponding computer program from the non-volatile memory 410 into memory 408 and then runs it. Of course, in addition to software implementation, one or more embodiments of this specification do not exclude other implementation methods, such as logic devices or a combination of hardware and software, etc. That is to say, the execution subject of the following processing flow is not limited to each logic unit, but can also be hardware or logic devices.
[0151] like Figure 5 As shown, Figure 5 This specification is a block diagram illustrating a task execution device based on the collaboration of multiple LLM agents according to an exemplary embodiment. This device can be applied to, for example... Figure 4 The illustrated electronic device implements the technical solution of this specification. The multi-level LLM agent includes a master LLM agent and multiple sub-LLM agents that have established communication connections with the master LLM agent; different sub-LLM agents are used to execute different sub-tasks; LLM agents in the multi-level LLM agent system that perform the same task share the task context associated with that task; the device includes: The first acquisition module 501 acquires task execution information input by the user. The first determining module 502 determines the target task to be executed as indicated by the user based on the task execution information, and determines the sub-LLM agents corresponding to the multiple sub-tasks contained in the target task from the multiple sub-LLM agents. The scheduling module 503 schedules the multiple subtasks to their respective sub-LLM agents. Each sub-LLM agent, based on the context related to the target task shared with other LLM agents jointly executing the target task, determines a target task collaboration mode from among the various task collaboration modes supported by the service system for its corresponding subtask, and executes the subtask according to the determined target task collaboration mode. The task collaboration mode represents the collaborative method by which the sub-LLM agent and the main LLM agent jointly execute the target task.
[0152] like Figure 6 As shown, Figure 6 This is a block diagram illustrating another task execution device based on the collaboration of multiple LLM agents according to an exemplary embodiment of this specification. This device can also be applied to, for example... Figure 4The illustrated electronic device implements the technical solution of this specification. The multi-level LLM agent includes a master LLM agent and multiple sub-LLM agents that have established communication connections with the master LLM agent; different sub-LLM agents are used to execute different sub-tasks; LLM agents in the multi-level LLM agent system that perform the same task share the task context associated with that task; the device includes: The second acquisition module 601 acquires the target subtask scheduled to the local machine by the main LLM agent; wherein the target subtask is the subtask corresponding to the target sub-LLM agent determined by the main LLM agent from among the multiple subtasks included in the target task to be executed as indicated by the user. The second determining module 602 determines a target task collaboration mode for the target sub-task from multiple task collaboration modes supported by the service system, based on the task context related to the target task shared with other LLM agents jointly executing the target task; wherein, the task collaboration mode is used to represent the collaboration method in which the sub-LLM agent and the main LLM agent jointly execute the target task. The execution module 603 executes the target sub-task according to the determined target task collaboration mode.
[0153] Accordingly, this specification also provides an electronic device including a processor; a memory for storing processor-executable instructions; wherein the processor is configured to implement all the steps in the previously described method flow.
[0154] Accordingly, this specification also provides a computer-readable storage medium having stored thereon executable computer program instructions; wherein, when executed by a processor, the instructions implement all the steps in the previously described method flow.
[0155] Accordingly, this specification also provides a computer program product having executable computer program instructions stored thereon; wherein, when the computer program instructions are executed by a processor, they implement all the steps in the previously described method flow.
[0156] The systems, devices, modules, or units described in the above embodiments can be implemented by computer chips or physical entities, or by products with certain functions. A typical implementation device is a server system. Of course, it is not excluded that with the future development of computer technology, the computer implementing the functions of the above embodiments may be, for example, a personal computer, a laptop computer, an in-vehicle human-machine interaction device, a cellular phone, a camera phone, a smartphone, a personal digital assistant, a media player, a navigation device, an email device, a game console, a tablet computer, a wearable device, or any combination of these devices.
[0157] While one or more embodiments of this specification provide the operational steps of the methods described in the embodiments or flowcharts, more or fewer operational steps may be included based on conventional or non-inventive means. The order of steps listed in the embodiments is merely one possible order of execution among many steps and does not represent the only possible order. In actual device or end product execution, the methods shown in the embodiments or drawings may be executed sequentially or in parallel (e.g., in a parallel processor or multi-threaded processing environment, or even a distributed data processing environment). The terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, product, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, product, or apparatus. Without further limitations, it is not excluded that the process, method, product, or apparatus that includes the elements may also have other identical or equivalent elements. For example, the use of terms such as "first," "second," etc., is used to indicate names and does not indicate any particular order.
[0158] For ease of description, the above devices are described in terms of function, divided into various modules. Of course, when implementing one or more of these specifications, the functions of each module can be implemented in one or more software and / or hardware components, or a module that performs the same function can be implemented by a combination of multiple sub-modules or sub-units. The device embodiments described above are merely illustrative. For example, the division of units is only a logical functional division; in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces, indirect coupling or communication connection between devices or units, and may be electrical, mechanical, or other forms.
[0159] This invention is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart illustrations and / or block diagrams. Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.
[0160] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.
[0161] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.
[0162] In a typical configuration, a computing device includes one or more processors (CPU), input / output interfaces, network interfaces, and memory.
[0163] Memory may include non-persistent storage in computer-readable media, such as random access memory (RAM) and / or non-volatile memory, such as read-only memory (ROM) or flash RAM. Memory is an example of computer-readable media.
[0164] Computer-readable media, including both permanent and non-permanent, removable and non-removable media, can store information using any method or technology. Information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase-change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, CD-ROM, digital versatile optical disc (DVD) or other optical storage, magnetic tape, magnetic magnetic disk storage, graphene storage or other magnetic storage devices, or any other non-transferable medium that can be used to store information accessible by a computing device. As defined herein, computer-readable media does not include transient computer-readable media, such as modulated data signals and carrier waves.
[0165] Those skilled in the art will understand that one or more embodiments of this specification can be provided as a method, system, or computer program product. Therefore, one or more embodiments of this specification may take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, one or more embodiments of this specification may take the form of a computer program product implemented on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0166] One or more embodiments of this specification can be described in the general context of computer-executable instructions, such as program modules, that are executed by a computer. Generally, program modules include routines, programs, objects, components, data structures, etc., that perform a particular task or implement a particular abstract data type. One or more embodiments of this specification can also be practiced in distributed computing environments where tasks are performed by remote processing devices connected via a communication network. In a distributed computing environment, program modules can reside in local and remote computer storage media, including storage devices.
[0167] The various embodiments in this specification are described in a progressive manner. Similar or identical parts between embodiments can be referred to mutually. Each embodiment focuses on describing the differences from other embodiments. In particular, system embodiments are basically similar to method embodiments, so the description is relatively simple; relevant parts can be referred to the descriptions in the method embodiments. In the description of this specification, the terms "one embodiment," "some embodiments," "example," "specific example," or "some examples," etc., refer to specific features, structures, materials, or characteristics described in connection with that embodiment or example, which are included in at least one embodiment or example of this specification. In this specification, the illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described can be combined in any suitable manner in one or more embodiments or examples. Moreover, without contradiction, those skilled in the art can combine and integrate the different embodiments or examples described in this specification and the features of different embodiments or examples.
[0168] The above description is merely an embodiment of one or more embodiments of this specification and is not intended to limit the scope of this specification. Various modifications and variations can be made to the one or more embodiments of this specification by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this specification should be included within the scope of the claims.
Claims
1. A task execution method based on the collaboration of multiple LLM agents, applied to the main LLM agent in a service system constructed based on multi-level LLM agents; wherein, The multi-level LLM agent includes a master LLM agent and multiple sub-LLM agents that have established communication connections with the master LLM agent; different sub-LLM agents are used to perform different sub-tasks. In the multi-level LLM agents, LLM agents performing the same task share the task context associated with that task; the method includes: Obtain task execution information input by the user; Based on the task execution information, the target task to be executed as indicated by the user is determined, and the sub-LLM agents corresponding to the sub-tasks contained in the target task are determined from the plurality of sub-LLM agents. The multiple subtasks are respectively scheduled to corresponding sub-LLM agents. The sub-LLM agents, based on the context related to the target task shared with other LLM agents jointly executing the target task, determine the target task collaboration mode for their corresponding subtasks from multiple task collaboration modes supported by the service system, and execute the subtasks according to the determined target task collaboration mode. The task collaboration mode is used to represent the collaboration method between the sub-LLM agents and the main LLM agent in jointly executing the target task.
2. The method according to claim 1, wherein the multiple task collaboration modes include a first collaboration mode and a second collaboration mode; in, The first collaborative mode refers to a task collaboration mode in which the sub-LLM agent independently executes its corresponding sub-task and returns the execution result of the sub-task to the user; The second collaboration mode refers to a task collaboration mode in which the sub-LLM agent returns the execution result of its corresponding sub-task to the main LLM agent, and the main LLM agent integrates the execution result with the execution results of other sub-tasks and returns the integrated execution result to the user.
3. The method according to claim 1, wherein the multi-level LLM agents further include target LLM agents that have established a communication connection with the master LLM agent; wherein, The target LLM agent is used to identify the intent of the task execution information; Based on the task execution information, the target task to be executed as indicated by the user is determined, and the sub-LLM agents corresponding to the sub-tasks included in the target task are determined from the plurality of sub-LLM agents, including: The task execution information is sent to the target LLM agent, which then performs inference calculations on the task execution information to determine the user's intent. Based on the user intent, the target task to be executed as indicated by the user is determined, and structured task scheduling information corresponding to the target task is generated. The structured task scheduling information includes agent identifiers of the sub-LLM agents determined from the plurality of sub-LLM agents, which correspond to the multiple sub-tasks included in the target task. Obtain the structured task scheduling information returned by the target LLM agent, and parse the structured task scheduling information to obtain the agent identifier contained in the structured task scheduling information; Based on the agent identifier, the sub-LLM agents corresponding to the multiple sub-tasks are determined from the multiple sub-LLM agents.
4. The method according to claim 3, wherein the structured task scheduling information further includes prompt words generated based on the task execution information; wherein, The prompt words are used to schedule the sub-LLM agents corresponding to the multiple sub-tasks to execute their respective sub-tasks; The multiple subtasks are respectively scheduled to their corresponding sub-LLM agents. Each sub-LLM agent, based on the context related to the target task shared with other LLM agents jointly executing the target task, further determines the target task collaboration mode for its corresponding subtask from among the various task collaboration modes supported by the service system, including: Obtain the prompt words contained in the structured task scheduling information; The prompt words are sent to the sub-LLM agents corresponding to the multiple sub-tasks respectively. The sub-LLM agents respond to the prompt words, further obtain the context related to the target task shared with the other LLM agents that jointly execute the target task, and determine the target task collaboration mode for the corresponding sub-task from the multiple task collaboration modes supported by the service system based on the context.
5. The method according to claim 4, wherein the context related to the target task includes the context related to the user session to which the target task belongs; the prompt word further includes a session identifier of the user session to which the target task belongs; The prompt words are sent to the sub-LLM agents corresponding to the plurality of sub-tasks, respectively. The sub-LLM agents respond to the prompt words by further acquiring context related to the target task, and based on the context, determine the target task collaboration mode for the corresponding sub-task from among the various task collaboration modes supported by the service system, including: The prompt words are sent to the sub-LLM agents corresponding to the multiple sub-tasks respectively. The sub-LLM agents respond to the prompt words, further obtain the context related to the user session corresponding to the session identifier contained in the prompt words, and determine the target task collaboration mode for the corresponding sub-task from the multiple task collaboration modes supported by the service system based on the context.
6. The method according to claim 1, wherein the multi-level LLM agent includes a master LLM agent and a first-level sub-LLM agent that has established a communication connection with the master LLM agent; And, a second-level sub-LLM agent that has established a communication connection with at least some of the sub-LLM agents in the first-level sub-LLM agent; in, Subtasks corresponding to each sub-LLM agent in the first-level sub-LLM agent include subtasks that do not involve modifying the context related to the target task and can be executed in parallel; The subtasks corresponding to each sub-LLM agent in the second-level sub-LLM agent include subtasks involving modifications to the context related to the target task and which need to be executed sequentially.
7. A task execution method based on the collaboration of multiple LLM agents, applied to any target sub-LLM agent in a service system constructed based on multi-level LLM agents; wherein, The multi-level LLM agent includes a master LLM agent and multiple sub-LLM agents that have established communication connections with the master LLM agent; different sub-LLM agents are used to perform different sub-tasks. In the multi-level LLM agents, LLM agents performing the same task share the task context associated with that task; the method includes: Obtain the target subtask scheduled to the local machine by the main LLM agent; wherein, the target subtask is the subtask corresponding to the target LLM agent determined by the main LLM agent from among the multiple subtasks included in the target task to be executed as indicated by the user; Based on the task context related to the target task shared with other LLM agents jointly executing the target task, a target task collaboration mode is determined for the target sub-task from multiple task collaboration modes supported by the service system; wherein, the task collaboration mode is used to represent the collaboration method in which the sub-LLM agent and the main LLM agent jointly execute the target task; The target sub-tasks are executed according to the determined target task collaboration mode.
8. The method according to claim 7, wherein the multi-level LLM agent includes a master LLM agent and a first-level sub-LLM agent that has established a communication connection with the master LLM agent; And, a second-level sub-LLM agent that has established a communication connection with at least some of the sub-LLM agents in the first-level sub-LLM agent; in, Subtasks corresponding to each sub-LLM agent in the first-level sub-LLM agent include subtasks that do not involve modifying the context related to the target task and can be executed in parallel; The subtasks corresponding to each sub-LLM agent in the second-level sub-LLM agent include subtasks involving modifications to the context related to the target task and which need to be executed sequentially.
9. The method according to claim 8, wherein the target subtask further comprises a plurality of subtasks; Execute the target sub-tasks according to the determined target task collaboration mode, including: If there are multiple sub-LLM agents at the next level that have established communication connections with the target sub-LLM agent, then the next-level sub-LLM agents corresponding to the multiple sub-tasks included in the target sub-task are determined from among the multiple sub-LLM agents at the next level. The target subtask is divided into multiple subtasks, which are respectively scheduled to the corresponding next-level sub-LLM agents. The next-level sub-LLM agents further determine the task collaboration mode for their corresponding subtasks from the multiple task collaboration modes supported by the service system based on the context related to the target subtask shared with other LLM agents that jointly execute the target subtask. The subtasks are then executed according to the determined task collaboration mode.
10. A service system, said service system being a service system built on a multi-level LLM agent, comprising: A master LLM agent for performing the method as described in any one of claims 1 to 6; Multiple sub-LLM agents that have established communication connections with the main LLM agent are used to execute the method as described in any one of claims 7 to 9.
11. A computer device, comprising a memory; a processor and a computer program stored in the memory and executable on the processor, wherein, When the processor executes the computer program, it implements the steps of the method according to any one of claims 1 to 9.
12. A computer program product comprising a computer program that, when executed by a processor, implements the steps of the method of any one of claims 1 to 9.