A fully automated method and related apparatus for handling complex tasks based on multi-agent collaboration
By using a multi-agent collaborative system, complete task requirements are obtained through interaction with the user, and then broken down into complex and simple sub-tasks. These sub-tasks are then processed by the target execution intelligent unit and the sub-agents, solving the problem of insufficient efficiency in handling complex tasks by single-agent systems and achieving more efficient and intelligent task processing.
Patent Information
- Application Number
- CN202510347342.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-24
- Publication Date
- 2025-12-02
- Estimated Expiration
- 2045-03-24
AI Technical Summary
Existing single-agent systems struggle to effectively handle complex tasks, resulting in insufficient task completion efficiency and intelligence.
Through a multi-agent collaborative system, the system first interacts with the user to obtain the complete task requirements, breaks them down into complex and simple sub-tasks, and then has them processed by the target execution intelligent unit and the sub-agents respectively, and finally summarizes the results.
It improves the efficiency and intelligence of task processing, reduces the difficulty of task decomposition, and obtains more accurate processing results.
Smart Images

Figure CN120256113B_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to the field of data processing technology, specifically to artificial intelligence technologies such as large language models, generative models, intelligent agents, and intelligent task scheduling, and particularly to a fully automated method, apparatus, electronic device, computer-readable storage medium, and computer program product for processing complex tasks based on multi-agent collaboration. Background Technology
[0002] With the development of artificial intelligence technology, the application of intelligent agents based on large language models is becoming increasingly widespread. Individual operation of intelligent agents is no longer sufficient to meet the needs of complex tasks; therefore, multi-agent collaborative systems have become an important research direction.
[0003] In a multi-agent system, agents can collaborate, communicate with each other, and work together to complete more complex tasks, thereby improving the efficiency and intelligence of task completion. Summary of the Invention
[0004] This disclosure presents a fully automated method, apparatus, electronic device, computer-readable storage medium, and computer program product for processing complex tasks based on multi-agent collaboration.
[0005] In a first aspect, embodiments of this disclosure propose a fully automated method for processing complex tasks based on multi-agent collaboration, comprising: obtaining a complete task requirement by interacting with a target user who proposes the original task requirement; breaking down the complete task requirement into multiple subtasks, each containing at least one complex subtask; wherein a complex subtask refers to a subtask that requires at least two sub-agents to process according to a collaborative process, and a single sub-agent is used to process a simple subtask; distributing the complex subtasks in each subtask to the corresponding target execution intelligent unit, and distributing the simple subtasks in each subtask to the corresponding target sub-agents; wherein the target execution intelligent unit contains at least two sub-agents for processing the complex subtasks; summarizing the subtask execution results returned by each target execution intelligent unit and each target sub-agent, and presenting the summarized task processing results to the target user.
[0006] Secondly, embodiments of this disclosure propose a fully automated complex task processing device based on multi-agent collaboration, comprising: a complete task requirement acquisition unit, configured to obtain the complete task requirement by interacting with a target user who proposes the original task requirement; a complete task splitting unit, configured to split the complete task requirement into multiple subtasks, each containing at least one complex subtask; wherein a complex subtask refers to a subtask that requires at least two sub-agents to process according to a collaborative process, and a single sub-agent is used to process a simple subtask; a subtask allocation unit, configured to distribute the complex subtasks in each subtask to the corresponding target execution intelligent unit and distribute the simple subtasks in each subtask to the corresponding target sub-agent; wherein the target execution intelligent unit contains at least two sub-agents for processing the complex subtasks; and a subtask result summarization and presentation unit, configured to summarize the subtask execution results returned by each target execution intelligent unit and each target sub-agent, and present the summarized task processing results to the target user.
[0007] Thirdly, embodiments of this disclosure provide an electronic device comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, the instructions being executed by the at least one processor to enable the at least one processor to implement the fully automated processing method for complex tasks based on multi-agent cooperation as described in the first aspect.
[0008] Fourthly, embodiments of this disclosure provide a non-transitory computer-readable storage medium storing computer instructions that enable a computer to perform a fully automated processing method for complex tasks based on multi-agent cooperation as described in the first aspect.
[0009] Fifthly, embodiments of this disclosure provide a computer program product including a computer program that, when executed by a processor, can implement the steps of the fully automated processing method for complex tasks based on multi-agent cooperation as described in the first aspect.
[0010] The fully automated complex task processing scheme based on multi-agent collaboration provided in this disclosure first interacts with the target user who initiated the original task request to determine the complete task requirement through the information obtained during the interaction process. Then, the complete task requirement is broken down into multiple subtasks, each containing at least one complex subtask. The complex subtask is used to distinguish it from the simple subtask that can be completed by a single agent. It refers to the subtask that requires at least two agents to process it according to a collaborative process. Next, the complex subtask is assigned to the target execution intelligent unit, and the simple subtask is assigned to the target agent. The target execution intelligent unit contains multiple agents arranged according to a collaborative process to process the complex subtask. Finally, the subtask execution results returned by each target execution intelligent unit and each target agent are summarized. This solution first attempts to supplement any missing information in the original task information by interacting with the target user, thereby obtaining a more accurate and complete task requirement. Then, by introducing an execution intelligent unit specifically for handling complex sub-tasks, the multiple sub-agents within this execution intelligent unit, arranged according to a collaborative process, can handle the complex sub-tasks more centrally. This eliminates the need to break down the complete task requirement into the most granular simple sub-tasks, reducing the difficulty of task decomposition, while achieving better sub-task processing results through the execution intelligent unit integrating multiple sub-agents.
[0011] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of this disclosure, nor is it intended to limit the scope of this disclosure. Other features of this disclosure will become readily apparent from the following description. Attached Figure Description
[0012] Other features, objects, and advantages of this disclosure will become more apparent from the following detailed description of non-limiting embodiments with reference to the accompanying drawings:
[0013] Figure 1 This is an exemplary system architecture to which this disclosure can be applied;
[0014] Figure 2 A flowchart illustrating a fully automated method for handling complex tasks based on multi-agent collaboration, provided in an embodiment of this disclosure;
[0015] Figure 3 This is a schematic diagram of a branching process that splits a complete task requirement through different fixed protocol processes, as provided in an embodiment of this disclosure.
[0016] Figure 4 A flowchart illustrating a method for ensuring the execution of subtasks by completing interrogations, as provided in this embodiment of the disclosure;
[0017] Figure 5-1 and Figure 5-2A flowchart illustrating a fully automated method for handling complex tasks based on multi-agent collaboration in an application scenario, provided by an embodiment of this disclosure.
[0018] Figure 6 A structural block diagram of a fully automated complex task processing device based on multi-agent collaboration provided in this disclosure embodiment;
[0019] Figure 7 This is a schematic diagram of the structure of an electronic device suitable for performing a fully automated processing method for complex tasks based on multi-agent cooperation, provided as an embodiment of the present disclosure. Detailed Implementation
[0020] The exemplary embodiments of this disclosure are described below with reference to the accompanying drawings, including various details of the embodiments to aid understanding; these should be considered merely exemplary. Therefore, those skilled in the art will recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of this disclosure. Similarly, for clarity and brevity, descriptions of well-known functions and structures are omitted in the following description. It should be noted that, unless otherwise specified, the embodiments and features described in this disclosure can be combined with each other.
[0021] The collection, storage, use, processing, transmission, provision, and disclosure of user personal information involved in the technical solution disclosed herein comply with the provisions of relevant laws and regulations and do not violate public order and good morals.
[0022] Figure 1 An exemplary system architecture 100 is shown, in which embodiments of the fully automated processing method, apparatus, electronic device, and computer-readable storage medium for complex tasks based on multi-agent collaboration of the present disclosure can be applied.
[0023] like Figure 1 As shown, system architecture 100 may include terminal devices 101, 102, and 103, a network 104, and a server 105. Network 104 serves as the medium for providing communication links between terminal devices 101, 102, and 103 and server 105. Network 104 may include various connection types, such as wired or wireless communication links, or fiber optic cables, etc.
[0024] Users can use terminal devices 101, 102, and 103 to interact with server 105 via network 104 to receive or send messages, etc. Various applications for communication between the terminal devices 101, 102, and 103 and server 105 can be installed, such as task processing applications, multi-agent collaboration applications, and instant messaging applications. Server 105 can be installed with or host various intelligent agents or intelligent task units to handle tasks of different granularities.
[0025] Terminal devices 101, 102, and 103 and server 105 can be either hardware or software. When terminal devices 101, 102, and 103 are hardware, they can be various electronic devices with displays, including but not limited to smartphones, tablets, laptops, and desktop computers. When terminal devices 101, 102, and 103 are software, they can be installed in the aforementioned electronic devices, and can be implemented as multiple software programs or software modules, or as a single software program or software module; no specific limitation is made here. When server 105 is hardware, it can be implemented as a distributed server cluster composed of multiple servers, or as a single server. When server 105 is software, it can be implemented as multiple software programs or software modules, or as a single software program or software module; no specific limitation is made here.
[0026] Server 105 can provide various services through its built-in applications. Taking a task processing application that can provide complex task processing services as an example, when running this application, server 105 can achieve the following effects: First, it receives the original task requirements submitted by the user using terminal devices 101, 102, and 103 via network 104, and interacts with the user based on these requirements to obtain the complete task requirements. Next, it breaks down the complete task requirements into multiple subtasks, each containing at least one complex subtask. The complex subtask refers to a subtask that requires at least two sub-agents to process according to a collaborative process, while a single sub-agent is used to handle simple subtasks. The next step is to distribute the complex subtasks to the corresponding target execution intelligent units and the simple subtasks to the corresponding target sub-agents. Each target execution intelligent unit contains at least two sub-agents for processing the complex subtasks. Finally, it summarizes the subtask execution results returned by each target execution intelligent unit and each target sub-agent, and returns the summarized task processing results to terminal devices 101, 102, and 103 via network 104 for presentation to the user.
[0027] It should be noted that the original task requirements can be obtained from terminal devices 101, 102, and 103 via network 104, or they can be pre-stored locally on server 105 through various means. Therefore, when server 105 detects that this data is already stored locally (e.g., when it starts processing previously stored pending tasks), it can choose to retrieve this data directly from the local storage. In this case, the exemplary system architecture 100 may not include terminal devices 101, 102, and 103 and network 104.
[0028] Since processing complex tasks requires significant computing resources and power, the fully automated complex task processing method based on multi-agent collaboration provided in the subsequent embodiments of this disclosure is generally executed by a server 105 with strong computing power and abundant computing resources. Correspondingly, the fully automated complex task processing device based on multi-agent collaboration is also generally located within the server 105. However, it should also be noted that when terminal devices 101, 102, and 103 also possess sufficient computing power and resources, they can also perform the aforementioned calculations performed by the server 105 through the task processing applications installed on them, thereby outputting the same results as the server 105. Especially when multiple terminal devices with different computing capabilities exist simultaneously, but the task processing application determines that its terminal device has strong computing power and abundant remaining computing resources, it can allow the terminal device to perform the aforementioned calculations, thereby appropriately reducing the computing pressure on the server 105. Accordingly, the fully automated complex task processing device based on multi-agent collaboration can also be located within the terminal devices 101, 102, and 103. In this case, the exemplary system architecture 100 may also exclude the server 105 and the network 104.
[0029] It should be understood that Figure 1 The number of terminal devices, networks, and servers shown is merely illustrative. Depending on implementation needs, any number of terminal devices, networks, and servers can be included.
[0030] Please refer to Figure 2 , Figure 2 A flowchart of a fully automated method for processing complex tasks based on multi-agent cooperation provided in this disclosure is included, wherein process 200 includes the following steps:
[0031] Step 201: Obtain the complete task requirements by interacting with the target user who proposed the original task requirements;
[0032] This step aims to enable the execution of a fully automated method for handling complex tasks based on multi-agent collaboration (e.g., Figure 1 The server 105 (or the interactive intelligent agent or interactive intelligent unit hosted on the server 105) interacts with the target user to supplement any ambiguous, incomplete, or missing parts of the original task requirements, ensuring that the task requirements accurately reflect the user's true intentions and providing a reliable foundation for subsequent task decomposition and execution. Specifically, this step aims to achieve the following objectives through interaction: 1) Information completion: Identifying and supplementing key information missing from the task requirements; 2) Requirement confirmation: Ensuring that the task requirements are consistent with the user's actual needs; 3) Task optimization: Optimizing the task requirements during the interaction process to better suit actual execution conditions.
[0033] The interactive behaviors mentioned in this step can be implemented in the following ways:
[0034] 1) Proactive questioning: Based on the analysis of the original task requirements, proactively ask users targeted questions; 2) Multi-turn dialogue: Gradually delve deeper through multi-turn dialogue to guide users to provide more detailed information; 3) Example guidance: Provide examples or templates to help users express their needs more clearly; 4) Feedback confirmation: Provide users with real-time feedback on their current understanding of the task requirements and confirm whether it is accurate.
[0035] One possible implementation method, including but not limited to, is as follows: First, based on the original task requirements proposed by the target user, the missing task information items are identified; then, the missing task information is obtained by sending at least one round of missing information request completion questions to the target user; finally, when the missing task information corresponding to all the missing task information items is obtained, the complete task requirements are determined based on all the missing task information and the original task requirements.
[0036] In intelligent customer service scenarios, users often present vague or incomplete requests, requiring interaction to complete the information and provide accurate service. For example, when a user enters, "I want to change my flight," the aforementioned agent can proactively ask, "What is your order number? What date do you need to change to?" After the user provides the information, it can further confirm by asking, "You wish to change your flight from Beijing to Shanghai on March 25th, is that correct?" This results in the final completed task requirement being: change the flight order number 123456 from Beijing to Shanghai on March 25th.
[0037] In intelligent assistant scenarios, users may present complex task requests that require interactive breakdown and completion of information. For example, when a user inputs, "Schedule a meeting for me," the aforementioned executor can proactively ask, "What is the topic of the meeting? Who are the attendees? What time would you like it to be held?" After the user provides the information, it can further ask confirmation questions: "The meeting topic is a project discussion, the attendees include Zhang San, Li Si, and Wang Wu, and the time is this Friday at 3 PM, is that correct?" This results in the final completed task request being: Schedule a project discussion meeting, with attendees Zhang San, Li Si, and Wang Wu, and the time is this Friday at 3 PM.
[0038] In decision support systems, users may raise high-level decision-making needs, requiring interactive completion of specific parameters and constraints. For example, when a user inputs, "Please help me analyze market trends," the aforementioned execution entity can proactively ask, "Which market do you want to analyze? What is the time frame? What indicators do you need to focus on?" And after the user provides the information, it can further initiate confirmation questions: "You want to analyze the Chinese smartphone market, with a time frame of 2022 to 2023, focusing on sales volume and market share, is that correct?" This ultimately completes the task requirement: analyze the sales volume and market share trends of the Chinese smartphone market from 2022 to 2023.
[0039] In intelligent information service platforms, users may submit vague search requests, requiring interaction to complete specific conditions. For example, when a user enters, "I want to find nearby restaurants," the aforementioned entity proactively asks, "What type of restaurant are you looking for? What is your budget range? What location do you need?" After the user provides this information, it can further confirm with questions such as, "You are looking for a Chinese restaurant with a budget of under 100 yuan per person, located in the city center, is that correct?" This results in the final task requirement being: Find Chinese restaurants in the city center with a budget of under 100 yuan per person.
[0040] To achieve the aforementioned effects, the implementing entity should first be able to parse the user's input (i.e., the original task requirements) using Natural Language Understanding (NLU) technology, identify key information within the task requirements, recognize the user's core intent, distinguish between "must-have" and "optional" parts of the task requirements, identify ambiguous or unclear parts of the task requirements, guide the user to clarify, and monitor the logical consistency of the task requirements in real time during the interaction, prompting the user to make further corrections. Simultaneously, contextual information should be preserved in multi-turn dialogues to avoid repeated questioning or information loss.
[0041] To further enhance the interactive experience, the interaction methods and content can be customized based on the user's historical preferences or behavioral habits. In addition, by combining domain knowledge bases, professional suggestions or options can be provided during the interaction process. It can even support multiple interaction methods (such as voice, text, images, etc.) to improve the user experience.
[0042] Step 202: Break down the complete task requirements into multiple subtasks, each containing at least one complex subtask;
[0043] Building upon step 201, this step aims to achieve efficient task execution by having the aforementioned executing entity break down the complete task requirements into multiple subtasks (including at least one complex subtask and a simple subtask), which are then handled by an execution intelligence unit (for handling complex subtasks) and sub-agents (for handling simple subtasks), respectively. A complex subtask refers to a subtask that requires at least two sub-agents to process it according to a collaborative process. In other words, by breaking down the complete task requirements into multiple subtasks, each subtask is ensured to be processed efficiently, while the introduction of a complex subtask processing mechanism reduces the difficulty of task decomposition and improves processing effectiveness.
[0044] The task decomposition process can first employ the following identification mechanism:
[0045] 1) Complex Subtask Identification: Identify the parts of a task requirement that require collaboration among multiple sub-agents. For example, the task requirement may contain steps with multiple dependencies, or require support from various professional capabilities.
[0046] 2) Simple subtask identification: Identify the parts of a task requirement that can be handled independently by a single sub-agent. For example, a single operation or explicit instruction in the task requirement.
[0047] 3) Task dependency analysis: Analyze the dependencies between subtasks to ensure that the decomposed subtasks can be executed in the correct order.
[0048] 4) Task Priority Assignment: Assign priority to subtasks based on the characteristics of task requirements to ensure that critical tasks are handled first.
[0049] After identifying the above components, complex subtasks need to be processed by an execution intelligence unit, which contains multiple sub-intelligent agents and executes them according to a collaborative process:
[0050] 1) Sub-agent division of labor: Each sub-agent is responsible for a specific step in a complex sub-task. For example, one sub-agent is responsible for data collection, and another sub-agent is responsible for data analysis.
[0051] 2) Collaboration Process: Multiple sub-agents collaborate according to a dynamically orchestrated processing flow that conforms to the processing logic, ensuring the efficiency and consistency of task processing. For example, after sub-agent A completes a task, it passes the result to sub-agent B, which then continues processing based on the result.
[0052] 3) Result integration: The execution intelligent unit integrates the processing results of multiple sub-intelligent agents to form the final output of complex sub-tasks.
[0053] In intelligent customer service scenarios, users may raise complex requests involving multiple steps, requiring task breakdown and collaborative processing. For example, a user inputting, "My order has a problem; I need a refund and to reorder," can be broken down into the following sub-tasks:
[0054] 1. Complex subtask: Processing refunds and reordering (requires collaboration of multiple sub-agents), specifically sub-agent A which verifies order information and processes refunds, and sub-agent B which reorders according to user needs;
[0055] 2. Simple subtask: Send a notification of the processing result to the user (handled by a single sub-agent).
[0056] In this scenario, the overall processing flow is as follows: After sub-agent A completes the refund, it transmits the result to sub-agent B, which then places a new order and notifies the user.
[0057] In intelligent assistant scenarios, users may request tasks involving multiple steps, requiring task breakdown and collaborative processing. For example, a user input like, "Help me arrange a business trip, including flights, hotels, and meeting arrangements," can be broken down into the following sub-tasks:
[0058] 1. Complex subtask: Arranging a business trip itinerary, specifically for sub-agent C which queries and books air tickets, sub-agent D which queries and books hotels, and sub-agent E which arranges meeting schedules.
[0059] 2. Simple subtask: Integrate the trip information and send it to the user.
[0060] In this scenario, the overall processing flow is as follows: after sub-agents C, D, and E complete their tasks, they pass the results to sub-agent E, which then integrates the information and notifies the user.
[0061] In decision support systems, users may raise complex requests involving multiple analytical steps, requiring task decomposition and collaborative processing. For example, a user input like, "Please analyze market trends and provide investment advice," can be broken down into the following sub-tasks:
[0062] 1. Complex subtask: Market trend analysis and investment advice generation, specifically for sub-agent F which collects market data, sub-agent G which analyzes market trends, and sub-agent H which generates investment advice.
[0063] 2. Simple subtask: Present the analysis results to the user.
[0064] In this scenario, the overall processing flow is as follows: after sub-agents F, G, and H complete their tasks, they pass the results to sub-agent H, and sub-agent D presents the results to the user.
[0065] In intelligent information service platforms, users may raise complex requests involving multiple query conditions, requiring task decomposition and collaborative processing. For example, a user input like, "I want to find a restaurant suitable for family gatherings, with children's facilities and offering vegetarian options," can be broken down into the following sub-tasks:
[0066] 1. Complex subtask: Restaurant query and filtering, specifically sub-agent I for querying restaurants suitable for family gatherings, sub-agent J for filtering restaurants with children's facilities, and sub-agent K for filtering restaurants that offer vegetarian options.
[0067] 2. Simple subtask: Present the filtering results to the user.
[0068] In this scenario, the overall processing flow is as follows: after sub-agents I, J, and K complete their tasks, they pass the results to sub-agent K, which then presents the results to the user.
[0069] Furthermore, to improve the effectiveness of task decomposition and processing, the decomposition method and number of subtasks can be dynamically adjusted according to the characteristics of task requirements, and subtasks can be intelligently allocated according to the capabilities and load of sub-agents. In addition, the collaborative process of complex subtasks can be optimized by analyzing historical task execution data.
[0070] Step 203: Distribute each subtask to the corresponding target execution intelligent unit or target sub-intelligent agent;
[0071] Building upon step 202, this step aims to ensure that complex subtasks can be efficiently processed by multiple sub-agents according to a collaborative process, while simple subtasks can be quickly completed by a single sub-agent, by assigning subtasks to the target execution intelligent unit or target sub-agent. Specifically, the target execution intelligent unit contains multiple sub-agents arranged according to a collaborative process for processing the decomposed complex subtasks. This target execution intelligent unit is an execution intelligent unit for processing complex subtasks split from the complete task requirements. The execution intelligent unit is an intelligent task unit for executing tasks. This intelligent task unit is a general-purpose task completion unit, designed to include at least two sub-agents and collaborative process information representing the collaborative process that should be followed between different sub-agents. The collaborative process information is determined based on the processing logic of the specific complex subtask being executed. To facilitate collaboration between different sub-agents, the intelligent task unit can also be designed to include an information storage module and a communication module for transmitting information between different sub-agents. Of course, if there is no specially designed information storage module and communication module, other mechanisms that can achieve similar information interaction and information storage can be selected, and no specific limitation is made here.
[0072] Therefore, the core objectives of the solution provided in this step are: 1) Task allocation: Based on the characteristics of the subtasks, complex subtasks are allocated to the execution intelligent unit, and simple subtasks are allocated to the sub-intelligent agents; 2) Collaborative processing: Through the collaborative process within the execution intelligent unit, multiple sub-intelligent agents can be ensured to efficiently collaborate in processing complex subtasks; and 3) Information transmission: Through the information storage module and the communication module, information sharing and transmission between sub-intelligent agents are realized.
[0073] The execution intelligence unit is a specialized intelligent task unit designed to handle complex subtasks, and its design includes the following core components:
[0074] 1) Multiple sub-agents: Each sub-agent is responsible for a specific step in a complex subtask. For example, sub-agent A is responsible for data collection, and sub-agent B is responsible for data analysis.
[0075] 2) Collaboration process information: This represents the collaboration process between sub-agents, ensuring the orderliness and efficiency of task processing. For example, after sub-agent A completes a task, it passes the result to sub-agent B, which then continues processing based on the result.
[0076] 3) Information storage module: Used to store intermediate results and task states between sub-agents. For example, it stores data collected by sub-agent A for use by sub-agent B.
[0077] 4) Communication module: Used to realize information transmission and synchronization between sub-agents. For example, after sub-agent A completes its task, it notifies sub-agent B to start processing through the communication module.
[0078] As for the allocation of subtasks, it can be handled flexibly according to the characteristics of the subtasks:
[0079] 1) Complex Subtask Allocation: Complex subtasks are allocated to the execution intelligence unit, where multiple sub-agents within it process the tasks according to a collaborative workflow dynamically orchestrated during the processing. For example, the task of "arranging a business trip" is allocated to the execution intelligence unit, with sub-agents A, B, and C handling flight tickets, hotel bookings, and meeting arrangements, respectively.
[0080] 2) Simple subtask assignment: Assign simple subtasks to individual sub-agents for independent completion. For example, assign the "send notification" task to sub-agent D for independent completion.
[0081] It can also monitor the execution status of tasks in real time to ensure that tasks are completed as expected. For example, monitoring the task progress of sub-agents A, B, and C can ensure that business trip arrangements are completed on time.
[0082] Step 204: Summarize the subtask execution results returned by each target execution intelligent unit and each target sub-intelligent agent, and present the summarized task processing results to the target user.
[0083] Building upon step 203, this step aims to have the aforementioned executing entity summarize the execution results of each sub-task and present the final processing result to the target user, ensuring that the user receives complete, accurate, and easily understandable task feedback. This primarily involves result integration (integrating the execution results of each sub-task into a complete task processing result), result optimization (optimizing the summarized result to better meet the user's needs and expectations), and result presentation (presenting the task processing result in a way that is easy for the user to understand, thereby improving the user experience).
[0084] Specifically, the process of summarizing and presenting results needs to be handled flexibly according to the characteristics of the task:
[0085] 1) Results Collection: Collect the execution results of sub-tasks from each target execution intelligent unit and target sub-agent. For example, collect the results of flight booking, hotel booking, and meeting arrangement from sub-agents A, B, and C, respectively.
[0086] 2) Results Integration: Integrate the collected sub-task results into a complete task processing result. For example, integrate the results of flight tickets, hotel bookings, and meeting arrangements into a complete business trip itinerary.
[0087] 3) Result Optimization: Optimize the integrated results, such as deduplication, sorting, and formatting. For example, arrange the trip information in chronological order and format it in a user-friendly format.
[0088] 4) Result Presentation: Present task processing results in a way that is easy for users to understand, such as text, tables, charts, etc. For example, present business trip itineraries in tabular form, along with detailed time and location information.
[0089] In intelligent customer service scenarios, users may raise complex requests involving multiple steps, requiring a complete solution through result aggregation and presentation. Taking the user input: "My order has a problem, I need a refund and to reorder," the results would be summarized as follows: Sub-agent A: Refund processing completed, refund amount returned to the original payment account; Sub-agent B: Reorder completed, new order number generated.
[0090] The result displayed is: "Your refund has been processed, and the refund amount has been returned to the original payment account. A new order has been successfully generated, with order number 123456."
[0091] In intelligent assistant scenarios, users may propose task requests involving multiple steps, requiring a complete solution through result aggregation and presentation. Taking the user input: "Help me arrange a business trip, including flights, hotels, and meeting arrangements," as an example, the results would be summarized as follows: Sub-agent A: Flight booked, flight number CA123, time: 3 PM on March 25th; Sub-agent B: Hotel booked, hotel name: XX Hotel, check-in time: March 25th; Sub-agent C: Meeting arrangements completed, time: 10 AM on March 26th, location: XX meeting room.
[0092] The result reads: "Your business trip itinerary is as follows: On March 25th at 3 PM, take flight CA123 to your destination and check into Hotel XX. On March 26th at 10 AM, attend a meeting in Conference Room XX."
[0093] In decision support systems, users may raise complex requests involving multiple analytical steps, requiring the provision of complete analytical results through result aggregation and presentation. Taking the user input: "Please help me analyze market trends and provide investment advice," the results would be summarized as follows: Sub-agent A: Collected market data includes sales volume and market share from 2022 to 2023; Sub-agent B: Analysis results show a market growth trend with an average annual growth rate of 5%; Sub-agent C: Investment advice is "It is recommended to increase investment in the XX industry."
[0094] The results show that: "Market analysis results indicate that the average annual growth rate of sales volume and market share in the XX industry is 5% from 2022 to 2023. It is recommended to increase investment in this industry."
[0095] In intelligent information service platforms, users may submit complex requests involving multiple query conditions, requiring the provision of complete query results through result aggregation and presentation. Taking the user input: "I want to find a restaurant suitable for family gatherings, with children's facilities and offering vegetarian options," as an example, the results would be aggregated as follows: Sub-agent A: Found 10 restaurants suitable for family gatherings; Sub-agent B: Filtered 5 restaurants with children's facilities; Sub-agent C: Filtered 3 restaurants offering vegetarian options.
[0096] The results show: "We recommend the following 3 restaurants: 1. XX Restaurant (with children's facilities and vegetarian options); 2. YY Restaurant (with children's facilities and vegetarian options); 3. ZZ Restaurant (with children's facilities and vegetarian options)."
[0097] The fully automated processing method for complex tasks based on multi-agent collaboration provided in this disclosure first interacts with the target user who initiated the original task request to determine the complete task requirement through information obtained during the interaction process. Then, the complete task requirement is broken down into multiple subtasks, each containing at least one complex subtask. The complex subtask is used to distinguish it from a simple subtask that can be completed by a single agent. It refers to a subtask that requires at least two agents to process it according to a collaborative process. Next, the complex subtask is assigned to a target execution intelligent unit, and the simple subtask is assigned to a target agent. The target execution intelligent unit contains at least two agents for processing the complex subtask. Finally, the subtask execution results returned by each target execution intelligent unit and each target agent are summarized. This solution first attempts to supplement any missing information in the original task information by interacting with the target user, thereby obtaining a more accurate and complete task requirement. Then, by introducing an execution intelligent unit specifically for handling complex sub-tasks, the multiple sub-agents within this execution intelligent unit, arranged according to a collaborative process, can handle the complex sub-tasks more centrally. This eliminates the need to break down the complete task requirement into the most granular simple sub-tasks, reducing the difficulty of task decomposition, while achieving better sub-task processing results through the execution intelligent unit integrating multiple sub-agents.
[0098] Please refer to Figure 3 , Figure 3 This disclosure provides a schematic diagram of a branching process that breaks down a complete task requirement using different fixed protocol procedures, illustrating the following two schemes:
[0099] Option 1: First, the complete task requirements are sent to a pre-defined planning intelligent unit. This planning intelligent unit is an intelligent task unit used to decompose the received complete task requirements and execute the planning. Then, the planning intelligent unit is controlled to decompose the task in sequence through the task understanding sub-intelligent and the task splitting sub-intelligent agents according to the pre-defined first collaboration process, resulting in multiple subtasks containing at least one complex subtask. The task splitting sub-intelligent agents have a higher priority in decomposing complex subtasks than in decomposing simple subtasks.
[0100] In this solution, the planning intelligent unit, which serves as a dedicated intelligent task unit for task decomposition and planning, is designed to include the following core components:
[0101] 1) Task Understanding Sub-Agent: Responsible for parsing the core intent and key information of the complete task requirements. For example, identifying the flight tickets, hotel, and meeting arrangements involved in a user's "arrange business trip" task.
[0102] 2) Task decomposition sub-agent: Responsible for breaking down the complete task requirements into multiple sub-tasks, and prioritizing the identification of complex sub-tasks. For example, the task of "arranging a business trip" can be broken down into three sub-tasks: "booking flights", "booking hotels", and "arranging meetings", and "arranging meetings" can be identified as the complex sub-task.
[0103] 3) First Collaboration Process: Define the order and rules for task understanding and decomposition to ensure the orderly and efficient decomposition of tasks. For example, first, the task understanding sub-agent parses the task requirements, and then the task decomposition sub-agent decomposes the task.
[0104] The task understanding sub-agent parses the core intent and key information of the complete task requirement. For example, it identifies the time frame, market type, and analytical indicators involved in the user-submitted task of "analyzing market trends." The task decomposition sub-agent breaks down the complete task requirement into multiple sub-tasks, prioritizing the identification of complex sub-tasks. For instance, it breaks down the "analyzing market trends" task into three sub-tasks: "collecting market data," "analyzing market trends," and "generating investment recommendations," identifying "analyzing market trends" as the complex sub-task. The task decomposition sub-agent prioritizes the decomposition of complex sub-tasks, ensuring they are processed first. For example, when decomposing the "arranging business trips" task, it prioritizes identifying "arranging meetings" as the complex sub-task.
[0105] Option 2: First, the complete task requirements are sent to a pre-defined planning intelligent unit. This planning intelligent unit is an intelligent task unit used to decompose and execute the planning of the received complete task requirements. Then, the planning intelligent unit is controlled to decompose and check the task in sequence through the task understanding sub-intelligent, task splitting sub-intelligent, and splitting result self-checking sub-intelligent according to the pre-defined second collaboration process. This results in multiple sub-tasks, each containing at least one complex sub-task. The task splitting sub-intelligent prioritizes decomposing complex sub-tasks over simple sub-tasks. The splitting result self-checking sub-intelligent is used to confirm the correctness and executability of each split sub-task.
[0106] Unlike Option 1, Option 2 uses a second collaborative process that, based on the first collaborative process, further introduces a self-checking sub-agent for splitting results, which brings the following advantages:
[0107] 1) Improve the accuracy of disassembly results
[0108] Improvements: By splitting the results into self-checking sub-agents, the correctness of each sub-task is checked; Advantages: It avoids subsequent execution failures or deviations caused by incorrect task decomposition, thus improving the reliability of task processing.
[0109] Example: In the "Arrange Business Trip" task, the self-checking sub-agent checks whether the three sub-tasks "Book Flights", "Book Hotels", and "Arrange Meetings" are complete and conflict-free.
[0110] 2) Ensure the executability of subtasks.
[0111] Improvements: By splitting the results, the sub-agents are self-checked to confirm whether each sub-task is executable; Advantages: This avoids interruptions in subsequent processing caused by the non-executability of sub-tasks, thus improving the success rate of task processing.
[0112] Example: In the "Analyze Market Trends" task, the self-checking sub-agent checks whether the three sub-tasks of "Collect Market Data", "Analyze Market Trends" and "Generate Investment Recommendations" are executable (such as whether the data source is available).
[0113] 3) Reduce the risks of subsequent processing
[0114] Improvements: Self-inspection allows for the early detection and correction of issues in the disassembly results. Advantages: Reduces risks during subsequent task execution and improves overall task processing efficiency and effectiveness.
[0115] Example: In the "Find a Restaurant" task, the self-checking sub-agent checks whether there are logical conflicts or unexecutable conditions in the three sub-tasks of "searching for restaurants suitable for family meals", "filtering restaurants with children's facilities", and "filtering restaurants that offer vegetarian options".
[0116] 4) Enhance the robustness of the system
[0117] Improvements: Enhance the system's ability to handle abnormal situations through self-checking; Advantages: Improve the system's robustness and ensure stable operation under complex task requirements.
[0118] Example: In the "Process Refund and Reorder" task, the self-checking sub-agent checks whether there are dependencies or execution conflicts between the two sub-tasks "Process Refund" and "Reorder".
[0119] Furthermore, based on the second collaborative process, machine learning technology can be used to improve the intelligence level of the self-checking sub-agent, enabling it to automatically identify and correct more complex decomposition problems. It can even dynamically adjust the self-checking rules according to the characteristics of task requirements to ensure that the self-checking results are more in line with actual needs. In addition, it can support joint inspection of the decomposition results of multiple related tasks to ensure that the dependencies and execution order between tasks are correct.
[0120] Based on the above embodiments which have clarified how to decompose tasks, the planning intelligent unit can be further controlled to distribute each sub-task to the corresponding target execution intelligent unit or target sub-intelligent body in sequence through the correlation matching sub-intelligent body and the distribution sub-intelligent body according to the preset third collaboration process. The correlation matching sub-intelligent body is used to determine the target execution intelligent unit or target sub-intelligent body that matches each sub-task based on the correlation between the task and the task processing capability of the sub-intelligent body. The correlation matching sub-intelligent body associates complex sub-tasks with target execution intelligent units that have matching task processing capabilities.
[0121] In this embodiment, the core step of task allocation aims to distribute each subtask to the matched target execution intelligent unit or target sub-intelligent agent according to the preset third collaborative process, through associative matching of sub-intelligent agents and distribution of sub-intelligent agents. This ensures that complex subtasks can be efficiently processed by execution units with corresponding capabilities, while simple subtasks can be quickly completed by a single sub-intelligent agent. Specifically, associating complex subtasks with target execution intelligent units with corresponding processing capabilities ensures that they can be collaboratively processed by multiple sub-intelligent agents.
[0122] The third collaboration process mentioned above is the core process of task allocation, and its design includes the following key components:
[0123] 1) Relevance-based matching of sub-agents: Based on the relevance between tasks and sub-agents, the most suitable execution unit or sub-agent is matched for each sub-task. For example, the sub-task of "analyzing market trends" is matched to a target execution intelligent unit with data analysis capabilities.
[0124] 2) Distribute to sub-agents: Distribute the matched sub-tasks to the corresponding execution units or sub-agents to ensure that the tasks can be processed efficiently. For example, the "book flight tickets" sub-task can be distributed to a sub-agent with booking service capabilities.
[0125] 3) Task processing capability library: Stores task processing capability information for each execution unit and sub-agent, providing support for correlation matching. For example, it records that a certain execution unit has capabilities such as "data analysis" and "image processing".
[0126] To further improve the effectiveness of task matching and distribution, the current load of each execution unit and sub-agent can be considered during the task matching process to ensure a balanced task allocation. For example, tasks can be prioritized for allocation to execution units with lighter loads to avoid resource overload. Furthermore, historical task execution data can be used to optimize the task matching strategy and improve matching accuracy. For instance, based on historical success rates, execution units with better processing performance can be prioritized for matching.
[0127] Based on the above embodiments, if there is no target execution intelligent unit with task processing capabilities that match the complex subtasks, the correlation matching sub-intelligent agent constituting the planning intelligent unit can be controlled to return a prompt message indicating that the association of the complex subtasks has failed to be completed to the task splitting sub-intelligent agent, so that the task splitting sub-intelligent agent can re-split the originally split complex subtasks into multiple simple subtasks according to the prompt message.
[0128] This embodiment addresses the anomaly where a complex subtask cannot be matched with a suitable target execution intelligent unit. It returns a failure message to the task splitting intelligent agent, indicating a failed association, and triggers a re-splitting of the complex subtask to ensure continued task processing. This process includes anomaly handling (identifying the anomaly where a complex subtask cannot be matched with a target execution intelligent unit), information feedback (returning a failure message to the task splitting intelligent agent, providing a basis for re-splitting), and task re-splitting (re-splitting the complex subtask into multiple simpler subtasks based on the message, ensuring continued task processing).
[0129] Specifically, when the correlation matching sub-agent attempts to match complex sub-tasks, it finds that no target execution intelligent unit with the corresponding task processing capabilities exists. For example, the complex sub-task "generating investment recommendations" requires data analysis, risk assessment, and prediction capabilities, but the current system lacks an execution unit with these capabilities. The correlation matching sub-agent then generates a correlation failure message and returns it to the task splitting sub-agent. The message includes the specific content of the complex sub-task and the reason for the matching failure (such as missing capabilities). For example, the message might be: "Complex sub-task 'generating investment recommendations' failed to match; reason: lack of data analysis and risk assessment capabilities."
[0130] In the task decomposition stage, the task decomposition agent can break down the original complex subtask into multiple simpler subtasks based on the prompts. For example, "generating investment advice" can be broken down into three simpler subtasks: "collecting market data," "analyzing market trends," and "assessing investment risks."
[0131] To further improve the effectiveness of task re-splitting, the task splitting rules can be optimized based on the matching failure reasons in the prompts. This ensures that the re-splitting simpler subtasks can be processed by existing sub-agents. For example, in response to the prompt "lack of data analysis and risk assessment capabilities," priority should be given to splitting into simpler subtasks that can be processed by existing sub-agents. Furthermore, the task splitting strategy can be dynamically adjusted based on the current system resource status to ensure that the re-splitting simpler subtasks can be processed efficiently. For instance, when system resources are scarce, complex subtasks can be split into smaller, simpler subtasks.
[0132] Taking the intelligent customer service scenario as an example, if the matching fails when dealing with the complex subtask "process refund and reorder", and it is confirmed that the reason is the lack of an execution task unit with order processing capabilities, then the complex subtask can be re-split into two simple subtasks, "process refund" and "reorder", through the task splitting sub-agent.
[0133] Based on any of the above embodiments, Figure 4 The flowchart of a method for ensuring the execution of a subtask by completing questions, provided in this embodiment of the disclosure, aims to describe how, after a subtask has entered the execution phase, missing information can be further supplemented by questions to better ensure the effective execution of the subtask. The process 400 includes the following steps:
[0134] Step 401: In response to the target execution intelligent unit or the target sub-intelligent agent discovering the lack of necessary information during the execution of the corresponding sub-task, control the supplementary query sub-intelligent agent in the target execution intelligent unit or the target sub-intelligent agent to initiate a supplementary question to the target user regarding the missing necessary information;
[0135] Step 402: The supplementary query sub-agent or the target sub-agent in the target execution intelligent unit continues to execute the corresponding sub-task based on the necessary information provided by the target user in the supplementary reply.
[0136] This embodiment addresses the anomaly where the target intelligent unit or target sub-intelligent agent discovers missing necessary information while executing a sub-task through steps 401-402. It ensures the task is completed completely and accurately by initiating a completion question to the target user and continuing execution based on the user's supplementary response. The main key technical points involved are as follows:
[0137] 1) Information completion: Identify necessary information missing during subtask execution and prompt the user with a completion question; 2) Task continuation: Continue executing the subtask based on the missing information provided by the user, ensuring that the task can be completed completely and accurately; 3) User experience optimization: By proactively prompting for completion, avoid task interruption or failure due to missing information, thereby improving the user experience.
[0138] The information completion process can be carried out according to a preset procedure:
[0139] 1) Missing Information Identification: When the target execution intelligent unit or target sub-intelligent agent is executing a sub-task, it discovers that necessary information is missing. For example, when executing the "book a flight ticket" sub-task, it is found that the departure date information is missing.
[0140] 2) Question Completion Initiation: The supplementary inquiry sub-agent or target sub-agent initiates a question to the target user to complete the missing necessary information. For example, asking the user, "When is your departure date?"
[0141] 3) User supplementary response: The target user supplements the missing necessary information based on the completion question. For example, the user replies "The departure date is March 25th".
[0142] 4) Task continues execution: The target execution intelligent unit or target sub-intelligent agent continues to execute the sub-task based on the missing information provided by the user. For example, based on the departure date "March 25th", the "book flight" sub-task continues to be executed.
[0143] To further improve the effectiveness of information completion, more precise questions can be asked by incorporating contextual information. For example, when asking about the departure date, the user could be prompted, "You mentioned you are available on March 25th. Would you like to depart on that day?" Multi-round completion questions can also be supported to gradually fill in missing necessary information; for example, the departure date could be asked in the first round, and the departure time in the second. Default values or recommended options can also be provided in the completion questions to simplify user input. For example, "March 25th" could be recommended as the default value when asking about the departure date. Furthermore, exception handling can be implemented when the user fails to provide necessary or invalid information; for example, if the user does not provide a departure date, the user could be prompted, "The departure date is required; please complete it."
[0144] Taking the intelligent customer service scenario as an example, when executing the subtask of "processing refunds", it is found that the necessary information is missing: the refund amount. Therefore, the supplementary inquiry sub-agent can be controlled to ask the user: "What is your refund amount?", and then the user replies: "The refund amount is 500 yuan", so that the target execution intelligent unit can continue to execute the "processing refunds" subtask based on the supplementary information.
[0145] Considering the core challenges in completing open and complex tasks, the following points need to be addressed: 1) How to flexibly and efficiently understand user needs and accurately generate task execution plans. Since user input is often vague and abstract, traditional simple queries or keyword matching are difficult to accurately understand and reasonably break down; 2) How to select and schedule suitable agents to collaborate in completing complex tasks. Different tasks have diversity, complexity, and uncertainty, requiring the system to accurately identify task requirements and reasonably select and efficiently schedule multiple agents to work together based on task characteristics; 3) How to effectively display the task execution process and results through interactive interaction, allowing users to fully understand and participate in task execution, thereby improving the quality of task completion and user satisfaction.
[0146] Therefore, based on the inventive concept provided in the above embodiments, this embodiment specifically constructs a unified intelligent task unit as the basic unit of intelligent interaction, realizing unified scheduling and interaction of intelligent agents with different levels of complexity.
[0147] like Figure 5-1 As shown, each intelligent task unit consists of a task execution module, a communication and interaction module, and a history storage module. The task execution module supports multiple implementation methods, including basic script or API (Application Programming Interface) calls, large language model-driven agents, agents based on a combination of large language models and tools, and multi-agent collaborative workflows. The task execution module can receive different types of user input data (such as text, files, and multimodal content), efficiently complete specific tasks through intelligent tools and collaborative workflows, and output results. The communication and interaction module is responsible for unified interaction management with users or other agents, including accepting user input, providing execution result feedback, and proactively obtaining further user feedback. The history storage module records the historical results of each task execution to support efficient invocation of subsequent tasks and avoid duplicate execution.
[0148] like Figure 5-2 As shown, based on the aforementioned unified intelligent task unit, this embodiment further proposes a task completion system based on multi-agent collaboration, which includes three core components: an interactive intelligent unit, a planning intelligent unit, and an execution intelligent unit.
[0149] The interactive intelligence unit, serving as the primary interface between the system and the user, is responsible for engaging in natural dialogue. Through semantic understanding technology, it deeply analyzes the user's expressed needs, accurately translating ambiguous requests into clear task requirements. When a user's needs are unclear or complex, the interactive intelligence unit proactively guides the user to provide further clarification. Furthermore, the interactive intelligence unit also supports user feedback on previous execution results, dynamically optimizing further task execution.
[0150] Upon receiving the explicit task requirement text from the interactive intelligent unit, the planning intelligent unit automatically analyzes the content and characteristics of the requirement text, dynamically generating a detailed task execution plan. This plan includes a clear description of the tasks at each execution stage and the corresponding selection strategy for the execution intelligent unit. Based on the characteristics of the task requirements and the capabilities of each agent, the planning intelligent unit intelligently matches and determines the most suitable execution intelligent unit and coordinates the execution order of the tasks to ensure task execution efficiency and effectiveness.
[0151] The execution intelligent unit is specifically responsible for task implementation. It receives the task requirements text explicitly conveyed by the planning intelligent unit, as well as the execution output from the previous stage, and invokes the corresponding intelligent agents, dedicated tools, or intelligent collaborative workflows to complete the task. When encountering difficulties or requiring additional user confirmation during task execution, the execution intelligent unit will proactively interact with the user to obtain feedback and optimize the execution process. After the task is completed, the execution intelligent unit will report the task execution results to the planning intelligent unit, so that the planning intelligent unit can integrate and evaluate the results, dynamically update subsequent task plans, or determine whether the task execution meets user requirements and enter the task summary stage.
[0152] like Figure 5-2 As shown in the figure, the complete collaboration process of the task completion system based on multi-agent cooperation proposed in this embodiment is as follows:
[0153] 1) Users submit task requirements, which are then clarified through natural language interaction using interactive intelligent units;
[0154] 2) The interactive intelligent unit analyzes and clarifies the task requirements, transforms them into explicit task requirement text, and transmits it to the planning intelligent unit;
[0155] 3) The planning intelligent unit automatically generates a task execution plan based on the task requirement text, including determining the phased execution strategy of the task and selecting the optimal execution intelligent unit;
[0156] 4) The execution intelligent unit executes the tasks at each stage based on the planning information transmitted by the planning intelligent unit, and calls on intelligent agents, tools or collaborative workflows to complete the tasks;
[0157] 5) The execution intelligent unit proactively interacts and provides feedback to the user to obtain user confirmation or additional input, ensuring task quality and user satisfaction;
[0158] 6) After the task phase is completed, the execution intelligent unit feeds back the execution results to the planning intelligent unit, which then determines whether to adjust the subsequent planning or enter the task summary phase.
[0159] 7) After all task phases are completed or the termination conditions are met, the planning intelligent unit will summarize and integrate the overall execution results and form the final output results, which will be displayed to the user through the interactive intelligent unit.
[0160] The solution provided in this embodiment has the following innovative features compared to existing solutions:
[0161] 1) A unified and universal intelligent task unit architecture is proposed, which realizes the construction of intelligent agents of different complexity and types and a unified task scheduling mechanism. This mechanism can efficiently manage the execution and interaction of intelligent agents, tools and workflows, and support intelligent task units to actively initiate interactions with users, significantly improving task execution efficiency and user engagement.
[0162] 2) A multi-agent collaborative system based on interactive intelligent units, planning intelligent units, and execution intelligent units was constructed, realizing intelligent task decomposition, agent selection, and dynamic automatic scheduling, enabling the efficient completion of complex tasks through collaborative efforts. This system can flexibly handle various unexpected situations and dynamic changes during task execution, ensuring the efficient completion of tasks.
[0163] 3) It provides an interactive visualization of the task execution process and results. Users can not only intuitively observe the execution status and achievements of each stage of the task, but also participate in the task execution process at any time, helping the system to further optimize the task execution strategy through real-time feedback. This interactive task completion mechanism significantly enhances users' understanding and control over the task completion process.
[0164] This invention is applicable to intelligent customer service, intelligent assistants, decision support systems, and various intelligent information service platforms (see the examples given in the above embodiments for relevant examples). It can significantly improve the automation and intelligent processing level of complex tasks and meet users' needs for a higher level of interactive experience and task completion effect.
[0165] Further reference Figure 6 As an implementation of the methods shown in the above figures, this disclosure provides an embodiment of a fully automated complex task processing device based on multi-agent collaboration. This device embodiment is similar to... Figure 2 Corresponding to the method embodiments shown, this device can be specifically applied to various electronic devices.
[0166] like Figure 6As shown, the fully automated complex task processing device 600 based on multi-agent collaboration in this embodiment may include: a complete task requirement acquisition unit 601, a complete task splitting unit 602, a sub-task allocation unit 603, and a sub-task result summary and presentation unit 604. The system includes: a complete task requirement acquisition unit 601, configured to obtain the complete task requirement by interacting with the target user who proposed the original task requirement; a complete task splitting unit 602, configured to split the complete task requirement into multiple subtasks, each containing at least one complex subtask; wherein a complex subtask refers to a subtask that requires at least two sub-agents to process according to a collaborative process, and a single sub-agent is used to process a simple subtask; a subtask allocation unit 603, configured to distribute the complex subtasks in each subtask to the corresponding target execution intelligent unit and distribute the simple subtasks in each subtask to the corresponding target sub-agents; wherein the target execution intelligent unit contains at least two sub-agents for processing the complex subtasks; and a subtask result summarization and presentation unit 604, configured to summarize the subtask execution results returned by each target execution intelligent unit and each target sub-agent, and present the summarized task processing results to the target user.
[0167] In this embodiment, the specific processing and technical effects of the complete task requirement acquisition unit 601, the complete task splitting unit 602, the subtask allocation unit 603, and the subtask result summarization and presentation unit 604 in the multi-agent collaborative complex task fully automated processing device 600 can be referred to respectively. Figure 2 The relevant descriptions of steps 201-204 in the corresponding embodiments will not be repeated here.
[0168] In some other optional implementations of this embodiment, the complete task requirement acquisition unit 601 can be further configured to:
[0169] Based on the original task requirements proposed by the target users, identify the missing task information.
[0170] Missing task information is obtained by completing the question through at least one round of missing information requests initiated to the target user;
[0171] In response to obtaining the missing task information corresponding to all missing task information items, the complete task requirements are determined based on all missing task information and the original task requirements.
[0172] In some other optional implementations of this embodiment, the target execution intelligent unit is an execution intelligent unit for processing complex subtasks split from the complete task requirements, and the execution intelligent unit is an intelligent task unit for executing tasks. The intelligent task unit is designed to include at least two sub-intelligent agents and collaboration process information for characterizing the collaboration process that should be followed between different sub-intelligent agents. The collaboration process information is determined based on the processing logic of the specific complex subtask being adapted and executed.
[0173] In some other optional implementations of this embodiment, the intelligent task unit is also designed to include: an information storage module and a communication module for transmitting information storage between different sub-intelligent agents.
[0174] In some other optional implementations of this embodiment, the complete task splitting unit 602 is further configured as follows:
[0175] The complete task requirements are sent to the pre-defined planning intelligent unit; the planning intelligent unit is an intelligent task unit used to break down the received complete task requirements and execute the plan.
[0176] The control and planning intelligent unit decomposes the task sequentially through the task understanding sub-intelligent agent and the task decomposition sub-intelligent agent according to the preset first collaborative process, and obtains multiple subtasks containing at least one complex subtask; among them, the task decomposition sub-intelligent agent has a higher priority in decomposing complex subtasks than in decomposing simple subtasks.
[0177] In some other optional implementations of this embodiment, the complete task splitting unit 602 is further configured as follows:
[0178] The complete task requirements are sent to the pre-defined planning intelligent unit; the planning intelligent unit is an intelligent task unit used to break down the received complete task requirements and execute the plan.
[0179] The control and planning intelligent unit sequentially performs task decomposition and verification through the task understanding sub-intelligence, task decomposition sub-intelligence, and decomposition result self-checking sub-intelligence according to the preset second collaborative process, resulting in multiple subtasks containing at least one complex subtask. Among them, the task decomposition sub-intelligence has a higher priority in decomposing complex subtasks than in decomposing simple subtasks, and the decomposition result self-checking sub-intelligence is used to verify the correctness and executability of each decomposed subtask.
[0180] In some other optional implementations of this embodiment, the subtask allocation unit 603 is further configured to:
[0181] The control and planning intelligent unit, following a pre-defined third collaborative process, sequentially distributes complex subtasks from each subtask to the corresponding target execution intelligent unit and simple subtasks from each subtask to the corresponding target subtask through the correlation matching sub-intelligent agent and the distribution sub-intelligent agent. The correlation matching sub-intelligent agent is used to determine the target execution intelligent unit or target sub-intelligent agent that matches each subtask based on the correlation between the task and the task processing capabilities of the sub-intelligent agent. The correlation matching sub-intelligent agent associates complex subtasks with target execution intelligent units that have matching task processing capabilities.
[0182] In some other optional implementations of this embodiment, the fully automated complex task processing device 600 based on multi-agent cooperation may further include:
[0183] In response to the absence of a target execution intelligent unit with task processing capabilities matching the complex subtask, the associativity matching sub-intelligent agent in the planning intelligent unit is controlled to return a prompt message indicating that the association of the complex subtask has failed to be completed to the task splitting sub-intelligent agent, so that the task splitting sub-intelligent agent can re-split the originally split complex subtask into multiple simple subtasks based on the prompt message.
[0184] In some other optional implementations of this embodiment, the fully automated complex task processing device 600 based on multi-agent cooperation may further include:
[0185] In response to the target execution intelligent unit or the target sub-intelligent agent discovering missing necessary information during the execution of the corresponding sub-task, the supplementary inquiry sub-intelligent agent in the target execution intelligent unit or the target sub-intelligent agent is controlled to initiate a supplementary question to the target user regarding the missing necessary information.
[0186] The supplementary query sub-agent or target sub-agent in the target execution intelligent unit continues to execute the corresponding sub-task based on the necessary information provided by the target user in the supplementary reply.
[0187] This embodiment exists as a device embodiment corresponding to the above method embodiment. The fully automatic processing device for complex tasks based on multi-agent collaboration provided in this embodiment first interacts with the target user who initiated the original task request, thereby determining the complete task request through the information obtained in the interaction process. Then, the complete task request is broken down into multiple subtasks, each containing at least one complex subtask. The complex subtask is used to distinguish it from the simple subtask that can be completed by a single agent. It refers to the subtask that requires at least two agents to process according to a collaborative process. Next, the complex subtask is assigned to the target execution intelligent unit, and the simple subtask is assigned to the target agent. The target execution intelligent unit contains at least two agents for processing the complex subtask. Finally, the subtask execution results returned by each target execution intelligent unit and each target agent are summarized. This solution first attempts to supplement any missing information in the original task information by interacting with the target user, thereby obtaining a more accurate and complete task requirement. Then, by introducing an execution intelligent unit specifically for handling complex sub-tasks, the multiple sub-agents within this execution intelligent unit, arranged according to a collaborative process, can handle the complex sub-tasks more centrally. This eliminates the need to break down the complete task requirement into the most granular simple sub-tasks, reducing the difficulty of task decomposition, while achieving better sub-task processing results through the execution intelligent unit integrating multiple sub-agents.
[0188] According to embodiments of this disclosure, this disclosure also provides an electronic device, which includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, which, when executed by the at least one processor, enable the at least one processor to implement the fully automated processing method for complex tasks based on multi-agent cooperation described in any of the above embodiments.
[0189] According to embodiments of this disclosure, this disclosure also provides a readable storage medium storing computer instructions that enable a computer to execute the fully automated processing method for complex tasks based on multi-agent cooperation as described in any of the above embodiments.
[0190] According to embodiments of this disclosure, this disclosure also provides a computer program product that, when executed by a processor, can implement the fully automated processing method for complex tasks based on multi-agent cooperation described in any of the above embodiments.
[0191] Figure 7A schematic block diagram of an example electronic device 700 that can be used to implement embodiments of the present disclosure is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device may also represent various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the present disclosure described and / or claimed herein.
[0192] like Figure 7 As shown, device 700 includes a computing unit 701, which can perform various appropriate actions and processes based on a computer program stored in read-only memory (ROM) 702 or a computer program loaded from storage unit 708 into random access memory (RAM) 703. RAM 703 may also store various programs and data required for the operation of device 700. The computing unit 701, ROM 702, and RAM 703 are interconnected via bus 704. Input / output (I / O) interface 705 is also connected to bus 704.
[0193] Multiple components in device 700 are connected to I / O interface 705, including: input unit 706, such as keyboard, mouse, etc.; output unit 707, such as various types of monitors, speakers, etc.; storage unit 708, such as disk, optical disk, etc.; and communication unit 709, such as network card, modem, wireless transceiver, etc. Communication unit 709 allows device 700 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.
[0194] The computing unit 701 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 701 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 701 performs the various methods and processes described above, such as a fully automated processing method for complex tasks based on multi-agent cooperation. For example, in some embodiments, the fully automated processing method for complex tasks based on multi-agent cooperation can be implemented as a computer software program tangibly contained in a machine-readable medium, such as storage unit 708. In some embodiments, part or all of the computer program can be loaded and / or installed on device 700 via ROM 702 and / or communication unit 709. When the computer program is loaded into RAM 703 and executed by the computing unit 701, one or more steps of the fully automated processing method for complex tasks based on multi-agent cooperation described above can be performed. Alternatively, in other embodiments, computing unit 701 may be configured by any other suitable means (e.g., by means of firmware) to perform a fully automated processing method for complex tasks based on multi-agent cooperation.
[0195] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SoCs), payload-programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.
[0196] The program code used to implement the methods of this disclosure may be written in any combination of one or more programming languages. This program code may be provided to a processor or controller of a general-purpose computer, special-purpose computer, or other programmable data processing apparatus, such that when executed by the processor or controller, the program code causes the functions / operations specified in the flowcharts and / or block diagrams to be implemented. The program code may be executed entirely on a machine, partially on a machine, as a standalone software package partially on a machine and partially on a remote machine, or entirely on a remote machine or server.
[0197] In the context of this disclosure, a machine-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.
[0198] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device for displaying information to the user (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor); and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the computer. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).
[0199] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as a data server), or computing systems that include middleware components (e.g., an application server), or computing systems that include frontend components (e.g., a user computer with a graphical user interface or web browser through which a user can interact with embodiments of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., a communication network). Examples of communication networks include local area networks (LANs), wide area networks (WANs), and the Internet.
[0200] Computer systems can include clients and servers. Clients and servers are generally located far apart and typically interact through communication networks. The client-server relationship is created by computer programs running on the respective computers and having a client-server relationship with each other. The server can be a cloud server, also known as a cloud computing server or cloud host, which is a hosting product within the cloud computing service system to address the shortcomings of traditional physical hosts and Virtual Private Server (VPS) services, such as high management difficulty and weak business scalability.
[0201] According to the technical solution of this disclosure, firstly, the complete task requirement is determined by interacting with the target user who initiated the original task requirement and obtaining information through the interaction process. Then, the complete task requirement is broken down into multiple subtasks, each containing at least one complex subtask. The complex subtask is used to distinguish it from a simple subtask that can be completed by a single intelligent agent. It refers to a subtask that requires at least two intelligent agents to process it according to a collaborative process. Next, the complex subtask is assigned to the target execution intelligent unit, and the simple subtask is assigned to the target intelligent agent. The target execution intelligent unit contains at least two intelligent agents for processing the complex subtask. Finally, the subtask execution results returned by each target execution intelligent unit and each target intelligent agent are summarized. This solution first attempts to supplement any missing information in the original task information by interacting with the target user, thereby obtaining a more accurate and complete task requirement. Then, by introducing an execution intelligent unit specifically for handling complex sub-tasks, the multiple sub-agents within this execution intelligent unit, arranged according to a collaborative process, can handle the complex sub-tasks more centrally. This eliminates the need to break down the complete task requirement into the most granular simple sub-tasks, reducing the difficulty of task decomposition, while achieving better sub-task processing results through the execution intelligent unit integrating multiple sub-agents.
[0202] It should be understood that the various forms of processes shown above can be used to rearrange, add, or delete steps. For example, the steps described in this disclosure can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution disclosed in this disclosure can be achieved, and this is not limited herein.
[0203] The specific embodiments described above do not constitute a limitation on the scope of protection of this disclosure. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this disclosure should be included within the scope of protection of this disclosure.
Claims
1. A fully automated method for processing complex tasks based on multi-agent cooperation, comprising: Complete task requirements are obtained by interacting with the target user who proposed the original task requirements; The complete task requirements are sent to a preset planning intelligent unit. The planning intelligent unit is an intelligent task unit used to decompose the received complete task requirements and execute the plan. The intelligent task unit is designed to include at least two sub-intelligent agents and collaboration process information used to characterize the collaboration process that should be followed between different sub-intelligent agents. The planning intelligent unit is controlled to decompose the task sequentially through the task understanding sub-intelligent agent and the task splitting sub-intelligent agent according to the preset first collaborative process, or the planning intelligent unit is controlled to decompose and check the task sequentially through the task understanding sub-intelligent agent, the task splitting sub-intelligent agent and the splitting result self-checking sub-intelligent agent according to the preset second collaborative process, so as to obtain multiple sub-tasks containing at least one complex sub-task. The complex subtask refers to a subtask that requires at least two sub-agents to process according to a collaborative process. A single sub-agent is used to process simple subtasks. The collaborative process information is determined based on the processing logic of the specific complex subtask being adapted and executed. The task splitting sub-agents prioritize the decomposition of complex subtasks. The complex subtasks in each of the subtasks are assigned to the corresponding target execution intelligent unit, and the simple subtasks in each of the subtasks are assigned to the corresponding target sub-intelligent agents. The target execution intelligent unit includes at least two sub-intelligent agents for processing the complex subtasks. The subtask execution results returned by each target execution intelligent unit and each target sub-intelligent agent are summarized, and the summarized task processing results are presented to the target user.
2. The method according to claim 1, wherein, The process of obtaining complete task requirements through interaction with the target user who proposed the original task requirements includes: Based on the original task requirements proposed by the target user, identify the missing task information items; Missing task information is obtained by completing the missing information request dialogue initiated to the target user in at least one round. In response to obtaining missing task information corresponding to all missing task information items, the complete task requirements are determined based on all missing task information and the original task requirements.
3. The method according to claim 1, wherein, The target execution intelligent unit is an execution intelligent unit used to process complex subtasks split from the complete task requirements, and the execution intelligent unit is an intelligent task unit used to execute tasks.
4. The method according to claim 3, wherein, The intelligent task unit is also designed to include an information storage module and a communication module for transmitting information storage between different sub-intelligent agents.
5. The method according to claim 3 or 4, wherein, The self-checking sub-agent of the split results is used to confirm the correctness and executability of each split sub-task.
6. The method according to claim 5, wherein, The step of distributing complex subtasks from each of the subtasks to the corresponding target execution intelligent units and simple subtasks from each of the subtasks to the corresponding target sub-intelligent agents includes: The planning intelligent unit controls the sequential distribution of complex subtasks from each subtask to the corresponding target execution intelligent unit and simple subtasks from each subtask to the corresponding target sub-intelligent unit through a preset third collaboration process via an association matching sub-intelligent unit and a distribution sub-intelligent unit. The association matching sub-intelligent unit determines the target execution intelligent unit or target sub-intelligent unit that matches each subtask based on the association between the task and the task processing capabilities of the sub-intelligent unit. The association matching sub-intelligent unit associates the complex subtask with a target execution intelligent unit that has the matching task processing capabilities.
7. The method according to claim 6, further comprising: In response to the absence of a target execution intelligent unit with task processing capabilities matching the complex subtask, the associativity matching sub-intelligent agent constituting the planning intelligent unit is controlled to return a prompt message indicating that the complex subtask association has failed to occur to the task splitting sub-intelligent agent, so that the task splitting sub-intelligent agent can re-split the originally split complex subtask into multiple simple subtasks according to the prompt message.
8. The method according to any one of claims 1-4, further comprising: In response to the target execution intelligent unit or the target sub-intelligent agent discovering missing necessary information during the execution of the corresponding sub-task, the supplementary inquiry sub-intelligent agent in the target execution intelligent unit or the target intelligent agent initiates a supplementary question to the target user regarding the missing necessary information. The target execution intelligent unit controls the supplementary query sub-intelligent agent or the target sub-intelligent agent to continue executing the corresponding sub-task based on the necessary information supplemented by the target user.
9. A fully automated processing device for complex tasks based on multi-agent cooperation, comprising: The complete task requirement acquisition unit is configured to obtain the complete task requirements by interacting with the target user who proposed the original task requirements. A complete task decomposition unit is configured to send the complete task requirements to a preset planning intelligent unit. The planning intelligent unit is an intelligent task unit used to decompose and execute planning for the received complete task requirements. The intelligent task unit is designed to include at least two sub-intelligent agents and collaboration process information to characterize the collaboration process that should be followed between different sub-intelligent agents. The planning intelligent unit is controlled to decompose the task according to a preset first collaboration process by sequentially passing through the task understanding sub-intelligent agent and the task decomposition sub-intelligent agent, or controlled to decompose and check the task according to a preset second collaboration process by sequentially passing through the task understanding sub-intelligent agent, the task decomposition sub-intelligent agent, and the decomposition result self-checking sub-intelligent agent, to obtain multiple sub-tasks containing at least one complex sub-task. The complex subtask refers to a subtask that requires at least two sub-agents to process according to a collaborative process. A single sub-agent is used to process simple subtasks. The collaborative process information is determined based on the processing logic of the specific complex subtask being adapted and executed. The task splitting sub-agents prioritize the decomposition of complex subtasks. The subtask allocation unit is configured to assign complex subtasks from each of the subtasks to the corresponding target execution intelligence unit and to assign simple subtasks from each of the subtasks to the corresponding target sub-intelligent agents; wherein the target execution intelligence unit includes at least two sub-intelligent agents for processing the complex subtasks; The subtask result summarization and presentation unit is configured to summarize the subtask execution results returned by each of the target execution intelligent units and each of the target sub-intelligent agents, and present the summarized task processing results to the target user.
10. An electronic device, comprising: At least one processor; as well as A memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor, which, when executed by the at least one processor, enables the at least one processor to perform the fully automated processing method for complex tasks based on multi-agent cooperation as described in any one of claims 1-8.
11. A non-transitory computer-readable storage medium storing computer instructions for causing the computer to perform the fully automated processing method for complex tasks based on multi-agent cooperation as described in any one of claims 1-8.
12. A computer program product comprising a computer program that, when executed by a processor, implements the steps of the fully automated processing method for complex tasks based on multi-agent cooperation as described in any one of claims 1-8.
Citation Information
Patent Citations
Task processing method, device and equipment based on large model agent arrangement, storage medium and program product
CN118819778A
Intelligent scheduling method, system, device, equipment and medium
CN119151244A