Method for processing explicit result acquisition type task based on multi-agent cooperation and related device

By decomposing and processing user tasks through a multi-agent collaborative approach, the problem of generative large language models outputting results that do not meet specific needs in complex tasks is solved, thus achieving a more efficient and personalized task solution.

CN119558340BActive Publication Date: 2025-11-04BEIJING BAIDU NETCOM SCI & TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411615690.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-11-12
Publication Date
2025-11-04
Estimated Expiration
2044-11-12

AI Technical Summary

Technical Problem

Existing technologies struggle to effectively utilize generative large language models to output results that meet specific needs, especially lacking flexibility and personalized solutions when dealing with complex user problems.

Method used

A multi-agent collaborative approach is adopted, in which the main agent decomposes the user task into multiple sub-target tasks, which are then processed by different types of sub-agents. The sub-processing results are combined with the user's personalized preferences to output the sub-processing results, which are finally summarized by the main agent into the target processing results.

Benefits of technology

It improves the flexibility and personalized solution capabilities of generative large language models in handling complex tasks, reduces overall costs, and achieves better task processing results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119558340B_ABST
    Figure CN119558340B_ABST
Patent Text Reader

Abstract

The present disclosure provides a multi-agent cooperation-based explicit result obtaining task processing method and related device, which relates to the fields of artificial intelligence technologies such as generative large language models and agents. The method comprises: determining a target task to be solved according to a natural language input of a user; decomposing the target task into a plurality of sub-target tasks including at least an explicit result obtaining task, and issuing each sub-target task to each sub-target agent including at least an explicit result obtaining agent, different sub-agents being used to process different types of sub-tasks; controlling each sub-target agent to output a sub-processing result corresponding to the sub-target task to which the sub-target agent belongs in combination with the personalized preferences of the user, the sub-processing result including an explicit result corresponding to the explicit result obtaining task to which the explicit result obtaining agent belongs and output by the explicit result obtaining agent; and aggregating each sub-processing result to obtain a target processing result corresponding to the target task. By applying the scheme, better comprehensive task processing effects can be achieved at a lower comprehensive cost.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present disclosure relates to the technical field of data processing, in particular to the technical field of artificial intelligence such as generative large language models and agents, and more particularly to a method and apparatus for processing tasks of obtaining explicit results based on multi-agent cooperation, an electronic device, a computer readable storage medium, and a computer program product. BACKGROUND

[0002] With the rapid development and iteration of generative large language models, the generative large language models have good understanding of user input and the ability to give corresponding results.

[0003] In order to make the results output by the generative large language model more in line with specific needs, an agent is constructed by using the generative large language model as a base model and combining a preset role parameter.

[0004] How to use the agent to solve complex problems raised by users is still a technical problem to be solved by the current technical personnel in the field. SUMMARY

[0005] Embodiments of the present disclosure provide a method and apparatus for processing tasks of obtaining explicit results based on multi-agent cooperation, an electronic device, a computer readable storage medium, and a computer program product.

[0006] In a first aspect, a method for processing tasks of obtaining explicit results based on multi-agent cooperation is provided, including: determining a target task to be solved according to a natural language input of a user; decomposing the target task into a plurality of sub-target tasks including at least a task of obtaining explicit results, and issuing each sub-target task to each sub-target agent including at least an explicit result obtaining agent; wherein different sub-agents are used to process different types of sub-tasks; controlling each sub-target agent to output a sub-processing result corresponding to the sub-target task to which the sub-target agent belongs in combination with the individualized preference of the user; wherein the sub-processing result includes an explicit result corresponding to the task of obtaining explicit results output by the explicit result obtaining agent; and aggregating each sub-processing result to obtain a target processing result corresponding to the target task.

[0007] In a second aspect, the embodiments of the present disclosure provide a device for processing a task of obtaining a clear result based on multi-agent cooperation, which comprises: a target task determination unit configured to determine a target task to be solved according to natural language input of a user; a task decomposition and corresponding issuing unit configured to decompose the target task into a plurality of sub-target tasks including at least a task of obtaining a clear result, and issue each sub-target task to each sub-target agent including at least an agent for obtaining a clear result; wherein different sub-agents are used to process different types of sub-tasks; a sub-target agent control output unit configured to control each sub-target agent to output a sub-processing result corresponding to the sub-target task to which the sub-target agent belongs in combination with individualized preferences of the user; wherein the sub-processing result includes a clear result output by the agent for obtaining a clear result corresponding to the task of obtaining a clear result to which the agent belongs; and a sub-processing result aggregation unit configured to aggregate each sub-processing result to obtain a target processing result corresponding to the target task.

[0008] In a third aspect, the embodiments of the present disclosure provide an electronic device, which comprises: at least one processor; and a memory connected with the at least one processor in communication; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to implement the method for processing a task of obtaining a clear result based on multi-agent cooperation as described in the first aspect.

[0009] In a fourth aspect, the embodiments of the present disclosure provide a non-transitory computer-readable storage medium storing computer instructions for enabling a computer to implement the method for processing a task of obtaining a clear result based on multi-agent cooperation as described in the first aspect.

[0010] In a fifth aspect, the embodiments of the present disclosure provide a computer program product comprising a computer program, which, when executed by a processor, enables the steps of the method for processing a task of obtaining a clear result based on multi-agent cooperation as described in the first aspect.

[0011] The processing scheme for a task of obtaining a clear result based on multi-agent cooperation provided by the present disclosure comprises: determining, by a master agent, a target task to be solved according to natural language input of a user; decomposing the target task into a plurality of sub-target tasks including at least a task of obtaining a clear result; issuing each sub-target task to each sub-target agent including at least an agent for obtaining a clear result; outputting, by each sub-target agent, a sub-processing result corresponding to the sub-target task to which the sub-target agent belongs in combination with individualized preferences of the user under arrangement of the master agent; and aggregating, by the master agent, each sub-processing result to obtain a target processing result corresponding to the target task.

[0012] That is, the present disclosure adopts an agent cluster formed by a pre-constructed main agent and multiple sub-agents to process the task demand proposed by a user, wherein the main agent is responsible for understanding the task demand of the user, decomposing the overall task demand into multiple sub-target tasks that can be respectively executed by different sub-agents, and each sub-agent processes the sub-task matched with itself according to the mobilization of the main agent. That is, through the cooperation between the main agent and each sub-agent, the different components of the complex task can be processed by each sub-agent, and since different sub-agents are pre-constructed to be dedicated to processing different types of tasks, the scheme of the present disclosure for the main agent and multiple sub-agents to cooperatively process the task demand has better processing effect on a single type of task than using a single and all-purpose agent, and the relatively small sub-agents are also convenient for flexible addition and modification of corresponding functions, which can bring better comprehensive task processing effect with lower comprehensive cost.

[0013] It should be understood that the content described in this part is not intended to identify the key or important features of the embodiments of the present disclosure, nor is it intended to limit the scope of the present disclosure. Other features of the present disclosure will become apparent through the following description. BRIEF DESCRIPTION OF DRAWINGS

[0014] Other features, objects, and advantages of the present disclosure will become more apparent through reading the following detailed description of non-limiting embodiments made with reference to the accompanying drawings:

[0015] Figure 1 is an exemplary system architecture to which the present disclosure can be applied;

[0016] Figure 2 A flowchart of a method for processing a clear result obtaining type task based on multi-agent cooperation provided by an embodiment of the present disclosure;

[0017] Figure 3 A flowchart of a method for determining a target task according to natural language input provided by an embodiment of the present disclosure;

[0018] Figure 4 A flowchart of a method for decomposing a target task into multiple sub-target tasks including a clear result obtaining type task according to task elements provided by an embodiment of the present disclosure;

[0019] Figure 5 A flowchart of a method for creating a sub-agent provided by an embodiment of the present disclosure;

[0020] Figure 6 A flowchart of a method for controlling each sub-target agent to execute a corresponding sub-target task according to a determined execution order provided by an embodiment of the present disclosure;

[0021] Figure 7 A two-branch schematic diagram for controlling a sub-target intelligent agent to output a corresponding sub-processing result is provided for an embodiment of the present disclosure.

[0022] Figures 8-1 to 8-10 All are example diagrams of the method for processing the explicit result acquisition type task based on multi-agent cooperation in an application scenario provided by the embodiments of the present disclosure.

[0023] Figure 9 A structural block diagram of a device for processing the explicit result acquisition type task based on multi-agent cooperation is provided for an embodiment of the present disclosure.

[0024] Figure 10 A structural schematic diagram of an electronic device suitable for executing the method for processing the explicit result acquisition type task based on multi-agent cooperation is provided for an embodiment of the present disclosure. DETAILED DESCRIPTION

[0025] Exemplary embodiments of the present disclosure are described below with reference to the accompanying drawings, which include various details of the embodiments of the present disclosure to assist in understanding, and should be considered as merely exemplary. Therefore, those of ordinary skill in the art should recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of the present disclosure. Also, in the following description, descriptions of well-known functions and structures are omitted for clarity and conciseness. It should be noted that the embodiments in the present disclosure and the features in the embodiments can be combined with each other without conflict.

[0026] In the technical solutions of the present disclosure, the collection, storage, use, processing, transmission, provision and disclosure of user personal information involved in the technical solutions comply with relevant laws and regulations and do not violate public order and good customs.

[0027] Figure 1 An exemplary system architecture 100 that can apply the embodiments of the method, device, electronic device and computer readable storage medium for processing the explicit result acquisition type task based on multi-agent cooperation of the present disclosure is shown.

[0028] As shown in Figure 1 The system architecture 100 can include terminal devices 101, 102, 103, a network 104 and a server 105. The network 104 is used to provide a communication link medium between the terminal devices 101, 102, 103 and the server 105. The network 104 can include various connection types, such as wired, wireless communication links or optical fiber cables, etc.

[0029] The user can use the terminal device 101, 102, 103 to interact with the server 105 through the network 104 to receive or send messages, etc. The terminal device 101, 102, 103 and the server 105 can be installed with various applications for realizing information communication between the two, such as complex task processing applications, browser applications, instant messaging applications, etc.

[0030] The terminal device 101, 102, 103 and the server 105 can be hardware or software. When the terminal device 101, 102, 103 is hardware, it can be various electronic devices with display screens, including but not limited to smartphones, tablets, laptop computers, desktop computers, etc.; when the terminal device 101, 102, 103 is software, it can be installed in the above-mentioned electronic devices, which can be implemented as multiple software or software modules, or as a single software or software module, which is not specifically limited here. When the server 105 is hardware, it can be implemented as a distributed server cluster composed of multiple servers, or as a single server; when the server is software, it can be implemented as multiple software or software modules, or as a single software or software module, which is not specifically limited here.

[0031] The server 105 can provide various services through various built-in applications. Taking a complex task processing application that can provide one-key processing services for complex tasks as an example, the server 105 can achieve the following effects when running the complex task processing application: first, receiving a natural language input transmitted by the user through the terminal device 101, 102, 103 through the network 104, and then determining a target task to be solved according to the natural language input; then, decomposing the target task into multiple sub-target tasks containing at least explicit result obtaining tasks, and assigning each sub-target task to a corresponding sub-target agent containing at least an explicit result obtaining agent, different sub-agents being used to process different types of sub-tasks; next, controlling each sub-target agent to output a sub-processing result corresponding to the sub-target task to which it belongs in combination with the user's individualized preferences, the sub-processing result including an explicit result corresponding to the explicit result obtaining task to which the explicit result obtaining agent belongs; finally, aggregating each sub-processing result to obtain a target processing result corresponding to the target task.

[0032] Further, the server 105 can also return the target processing result to the terminal device 101, 102, 103 through the network 104, so that the terminal device 101, 102, 103 displays the received target processing result to the user.

[0033] It should be noted that the natural language input can be pre-stored locally on the server 105 in various ways in addition to being obtained from the terminal devices 101, 102, 103 through the network 104. Therefore, when the server 105 detects that the data has been stored locally (for example, a pending task left before starting processing), the data can be obtained directly from the local, and in this case, the exemplary system architecture 100 can also not include the terminal devices 101, 102, 103 and the network 104.

[0034] Since processing complex tasks requires more computing resources and stronger computing power, the multi-agent collaboration-based explicit result acquisition task processing method provided in subsequent embodiments of the present disclosure is generally executed by the server 105 with stronger computing power and more computing resources, and accordingly, the multi-agent collaboration-based explicit result acquisition task processing device is generally also provided in the server 105. However, it should also be noted that when the terminal devices 101, 102, 103 also have computing power and computing resources that meet the requirements, the terminal devices 101, 102, 103 can also complete the above-mentioned operations by installing a complex task processing application thereon, and then output the same results as the server 105. Especially in the case where there are multiple terminal devices with different computing power, but the complex task processing application determines that the terminal device has strong computing power and has more remaining computing resources, the terminal device can be allowed to execute the above-mentioned operations, thereby appropriately reducing the computing pressure of the server 105, and accordingly, the multi-agent collaboration-based explicit result acquisition task processing device can also be provided in the terminal devices 101, 102, 103. In this case, the exemplary system architecture 100 can also not include the server 105 and the network 104.

[0035] It should be noted that the main agent and each sub-target agent can be installed on the server 105, and each sub-target agent controlled by the main agent can also be installed on other servers or terminal devices different from the server 105, which is not specifically limited here.

[0036] It should be understood that Figure 1 The number of terminal devices, networks and servers in the exemplary system architecture 100 is only illustrative. According to the needs of implementation, there can be any number of terminal devices, networks and servers.

[0037] Please refer to Figure 2 , Figure 2 A flowchart of a multi-agent collaboration-based explicit result acquisition task processing method provided by an embodiment of the present disclosure, wherein the flowchart 200 includes the following steps:

[0038] Step 201: Determine the target task to be solved based on the user's natural language input;

[0039] This step aims to obtain the executor of the task-processing method (e.g., based on the explicit results of multi-agent collaboration) from the desired outcome. Figure 1 The server 105 shown, which carries the main intelligent agent, fully understands and analyzes the natural language input received from the user (e.g., through...). Figure 1 The network 104 shown receives the natural language input by the user through the terminal devices 101, 102, and 103, and then determines the target task that the user wishes to solve, that is, the target task corresponds to a task requirement.

[0040] Specifically, this step is typically a multi-step integrated processing procedure involving complex intelligent reasoning processes such as language understanding, context analysis, and task modeling. Examples include input reception, preprocessing and normalization, semantic understanding and intent recognition, and task modeling and reasoning.

[0041] In this process, the main intelligent agent, acting as the execution subject, first receives the user's natural language input, then preprocesses and normalizes the received natural language input to ensure that it can be effectively understood and processed. Next, it understands the user's intent reflected in the preprocessed and normalized input information based on the context, thereby determining the target task.

[0042] The input of natural language can take the form of speech, text, images, video, or other forms of language data. To facilitate processing, conversion techniques can be used to convert non-textual natural language input into natural language text that is easy to recognize and process. Considering that natural language input can present request information in various forms (e.g., direct requests, indirect requests, and multi-step requests), preprocessing and normalization processes, including word segmentation and annotation, noise removal, spell correction, and synonym processing, can be used to effectively understand the user's actual task requirements. In the intent recognition stage, keyword extraction, entity recognition, and the combination of contextual information can be used to accurately identify the user's intent expressed through natural language input. To convert the identified intent into a specific task, task modeling operations, including task abstraction (mapping the user's needs to a system-recognizable task framework) and target task definition, can be used to finally obtain the target task corresponding to the natural language input.

[0043] Step 202: Decompose the target task into multiple sub-target tasks that include at least a task with a clear result acquisition, and distribute each sub-target task to the corresponding sub-target agent that includes at least a clear result acquisition agent;

[0044] On the basis of step 201, this step aims to decompose the target task into a plurality of sub-target tasks by the above-mentioned execution subject, and correspondingly issue each sub-target task to each sub-target agent. Among them, the plurality of sub-target tasks decomposed by this step at least contain an explicit result acquisition type task, and correspondingly, each sub-target agent should at least contain an explicit result acquisition agent specially used for processing the explicit result acquisition type task.

[0045] Among them, different sub-agents are used to process different types of sub-tasks, that is, sub-agents specially used for single type tasks are respectively created in advance according to different task types. The task type division method is various, for example, according to the processing steps of the task, the processing method of the task, the complexity of the task, the data type involved in the task, etc. Here, it is not limited, that is, different sub-agents are used to process different atomic tasks (that is, the smallest unit of task, here "smallest" is a relative concept, not an absolute concept, which is the smallest task unit that can be decomposed at present).

[0046] Among them, the explicit result acquisition type task refers to the existence of a clear result for this type of task. The result can be unique or not unique, but the result should be clear and not suspected. For example, four arithmetic operation tasks, fixed parameter query tasks, matching tasks with fixed corresponding relationship, etc. Taking the task of "finding a travel plan from A to B" as an example, it is known that the travel plan to be obtained is from A to B. Without missing the necessary information of the travel plan, a plurality of feasible travel plans can be determined. Of course, the number of travel plans will decrease as the task restrictions increase, for example, adding "departure at X o'clock in the afternoon", "fastest travel plan", "travel plan with time not exceeding Y hours", "non-transfer, direct travel plan" and other task restrictions, but it will not affect the clear result. In other words, this explicit result acquisition type task is a type of task that only needs to determine the clear existing result expected by the user and feed it back to the user. Therefore, the execution focus of this type of task is to obtain the clear existing result that meets the task requirements.

[0047] It should be noted that the plurality of sub-target tasks decomposed by the target task in this embodiment should at least contain an explicit result acquisition type task. In addition, the task type of other sub-target tasks decomposed is not limited, which can be any type of task. Correspondingly, a sub-target agent with corresponding task processing ability should be used.

[0048] Step 203: Control each sub-target agent to output a sub-processing result corresponding to the sub-target task respectively according to the user's individualized preference.

[0049] On the basis of step 202, the present step aims to control each sub-target agent to output a sub-processing result corresponding to the sub-target task to which the sub-target agent belongs in combination with the personalized preference of the user by the above-mentioned execution subject. Among the multiple sub-processing results, at least the explicit result corresponding to the explicit result acquisition type task to which the explicit result acquisition agent belongs is output by the explicit result acquisition agent, and the remaining sub-processing results can also be the results obtained by other agents when processing the corresponding type of task. The sub-processing result can be obtained in dependence on other sub-processing results, or can be determined by the sub-target task description information without dependence on other sub-processing results.

[0050] That is, each sub-target agent outputs a sub-processing result corresponding to the sub-target task issued under the calling or arrangement of the main agent in combination with the personalized preference of the user. Specifically, the main agent can inform the execution order or execution trigger signal when issuing each sub-target task to each sub-target agent, so that each sub-target agent executes the corresponding sub-target task according to the informed execution order or executes the corresponding sub-target task when judging that the current state meets the requirements of the execution trigger signal, that is, each sub-target agent actively executes the corresponding sub-target task at the appropriate time; the main agent can also actively call each sub-target agent to execute the corresponding sub-target task according to the requirements when determining that different sub-target agents meet the execution time, that is, each sub-target agent only passively executes the corresponding sub-target task according to the calling instruction.

[0051] Whether it is the active execution mechanism mentioned above or the passive execution mechanism, it belongs to different collaborative forms of the main agent and multiple sub-agents in processing the task cooperatively. The specific choice can be made according to the actual needs of the actual application scene, which is not limited here.

[0052] Step 204: aggregating each sub-processing result to obtain a target processing result corresponding to the target task.

[0053] On the basis of step 203, the present step aims to obtain a target processing result corresponding to the target task by aggregating each sub-processing result by the above-mentioned execution subject.

[0054] Specifically, when aggregating the sub-processing result, the decomposition manner of decomposing the target task into multiple sub-target tasks and multiple processing manners including deduplication, adjusting expression, adjusting sentence and enriching expression manner can be fully referred to, so that the finally obtained target processing result can not only meet the task demand of the user, but also highlight the key demand result information as much as possible and improve the information recognition degree.

[0055] The method for processing a task of obtaining a clear result based on multi-agent cooperation provided by the embodiments of the present disclosure adopts an agent cluster formed by a pre-constructed main agent and multiple sub-agents to process a task demand proposed by a user, wherein the main agent is responsible for understanding the task demand of the user, decomposing the overall task demand into multiple sub-target tasks that can be executed by different sub-agents respectively, and each sub-agent processes the sub-task matched with itself according to the mobilization of the main agent. That is, through the cooperation between the main agent and each sub-agent, different components of a complex task can be processed by each sub-agent in its own position. Moreover, since different sub-agents are pre-constructed to be dedicated to processing different types of tasks, the scheme of the main agent and multiple sub-agents for processing a task demand cooperatively adopted by the present disclosure has a better processing effect on a single type of task than a single and all-purpose agent, and the relatively small sub-agents are also convenient for flexible addition and modification of corresponding functions, so that the scheme can bring better comprehensive task processing effect with lower comprehensive cost.

[0056] For a better understanding of how to determine the target task, please also refer to Figure 3 , Figure 3 A flowchart of a method for determining a target task according to a natural language input provided by the embodiments of the present disclosure, wherein the flow 300 includes the following steps:

[0057] Step 301: converting the natural language input of the user into a natural language text;

[0058] This step aims to convert various forms of natural language input into natural language text that is convenient for recognition and output by the above-mentioned execution subject, that is, to uniformly convert various forms into a text form, for example, a voice-to-text technology can be used to convert a voice form of natural language signal into a natural language text.

[0059] Step 302: performing intent recognition on the natural language text to obtain an intent recognition result containing an intent of obtaining a clear result;

[0060] On the basis of step 301, this step aims to perform intent recognition on the natural language text by the above-mentioned execution subject to obtain an intent recognition result containing an intent of obtaining a clear result. That is, the intent recognition result at least contains the intent of obtaining a clear result, and can also contain other intents, which are not limited here.

[0061] An implementation mode including but not limited to the above can include the following specific steps:

[0062] First, the semantic understanding of the natural language text is performed by the execution subject to obtain a semantic understanding result; then, the execution subject determines a preliminary intent (or suspected intent) according to the semantic understanding result, and in order to improve the accuracy of intent determination, the execution subject can also use the personalized preferences of the user to modify the preliminary intent to obtain an intent recognition result including the explicit result obtaining intent.

[0063] That is, the implementation mode first determines the preliminary intent corresponding to the natural language text by using semantic understanding, and then uses the personalized preferences of the user to modify the preliminary intent, such as screening, excluding, supplementing implicit or missing information, and the like, so that the finally determined intent recognition result is more in line with the actual expectations of the user.

[0064] In addition to the implementation mode, the various technologies involved in the multiple processing links mentioned in step 201 can also be used to construct another implementation mode of determining the intent recognition result according to the actual application scenario, which will not be listed one by one here.

[0065] Step 303: determining a target task to be solved according to the intent recognition result.

[0066] On the basis of step 302, this step aims to determine a target task to be solved by the execution subject according to the intent recognition result. The target task includes a task for obtaining an explicitly existing result.

[0067] Specifically, to convert the recognized intent into a specific task, task modeling operations including task abstraction (mapping user's demand to a system recognizable task framework) and target task definition can also be used to finally obtain a target task corresponding to the natural language input.

[0068] The embodiment provides an implementation mode of determining a target task according to a natural language input through steps 301-303, which includes the key processing steps of form conversion, intent recognition, and determining a target task based on the intent recognition result, so as to provide an implementation scheme with high feasibility and more accurate determined target task.

[0069] To deepen the understanding of how to decompose the target task into multiple sub-target tasks, please refer to Figure 4 , Figure 4 A flowchart of a method for decomposing a target task into multiple sub-target tasks including an explicit result obtaining task according to a task element provided by the embodiment of the disclosure, the flowchart 400 includes the following steps:

[0070] Step 401: determining multiple task elements constituting the target task;

[0071] This step aims to determine the task elements that constitute the target task by the above execution subject, and different task elements correspond to different types of tasks.

[0072] Among them, task elements refer to the basic units or components that constitute the target task. Each target task may be composed of several independent elements, which are the key steps, sub-tasks or resources necessary to achieve the target task. Each task element may correspond to an independent task, and they work together to promote the completion of the target task.

[0073] Suppose the target task is "developing a new software product", the target task can be decomposed into multiple task elements, each of which corresponds to a different type of task:

[0074] 1) Requirement analysis (management task): determine user requirements, market demand;

[0075] 2) System design (technical task): architecture design, database design, interface design;

[0076] 3) Coding development (technical task): programmers write code and implement functions;

[0077] 4) Testing and quality control (technical and management tasks): unit testing, integration testing, quality assurance;

[0078] 5) Marketing and promotion (management and communication tasks): develop a marketing plan and promote the product.

[0079] Step 402: According to the single task element that specifically determines the explicitly existing result matching the question raised by the user, decompose the target task into an explicit result acquisition type task for obtaining an explicit result;

[0080] This step aims to decompose the target task into an explicit result acquisition type task for obtaining an explicit result by the above execution subject, specifically, relying on the single task element that specifically determines the explicitly existing result matching the question raised by the user.

[0081] Step 403: Decompose the target task into multiple sub-target tasks containing only a single task element.

[0082] This step aims to decompose the target task into multiple sub-target tasks containing only a single task element by the above execution subject, referring to the way of decomposing the explicit result acquisition type task by task element in step 402, that is, the sub-target task at least contains the explicit result acquisition type task.

[0083] The embodiment provides a task decomposition manner according to task elements constituting a target task and decomposing each task element into each sub-target task, and in combination with predefinition of the task elements (for example, according to a manner of constructing different sub-agents), the decomposition of the task can be flexibly completed according to actual conditions.

[0084] Further, even if the same task element, there is an additional condition attached, so that when the task is decomposed, it needs to be decomposed into a sub-target task carrying the additional condition, so as to subsequently select a suitable sub-agent as a sub-target agent in combination with the additional condition when the sub-target task is issued to the sub-target agent.

[0085] An implementation manner, which includes but is not limited to, can be: in response to the existence of an additional condition corresponding to a task element, decomposing the target task into a plurality of sub-target tasks containing only a single task element and a corresponding additional condition, the additional condition can include: a specified task processing manner and / or a specified result display form. Different task processing manners are abstracted from different processing logics of the same type of task (for example, the same type of task can have multiple different processing logics, but finally all can get the correct answer); the result display form can include at least one of: a text type, a table type, an image type, a video type, and an interactive card type.

[0086] To deepen the understanding of how to pre-create different sub-agents, please refer to Figure 5 , Figure 5 A flowchart of a method for creating a sub-agent provided by the embodiment of the present disclosure, the flow 500 includes the following steps:

[0087] Step 501: determining a task type set for creating a sub-agent;

[0088] Step 502: for each type of task in the task type set, constructing a sub-agent having a task processing capability of the corresponding type of task.

[0089] The embodiment provides a scheme for constructing a sub-agent having a task processing capability of each type of task for the determined multiple types of tasks. Further, the number of sub-agents corresponding to each type of task can be only one, or can be multiple, and the multiple can be different or completely the same, whether to be different can be determined according to whether the type of task usually contains further additional conditions.

[0090] Specifically, for a first target type task having at least two task processing modes, a sub-agent corresponding to each task processing mode can be constructed for the first target type task; and for a second target type task having at least two result presentation forms, a sub-agent corresponding to each result presentation form can be constructed for the second target type task.

[0091] For a better understanding of how the multiple sub-target agents are executed in an orderly manner and how the final sub-processing result is ensured, please refer to Figure 6 , Figure 6 A flowchart of a method for controlling the sub-target agents to execute the corresponding sub-target tasks according to the determined execution order is provided in an embodiment of the present disclosure, and the flowchart 600 includes the following steps:

[0092] Step 601: determining the execution order between different sub-target tasks with a dependency relationship as serial.

[0093] Step 602: determining the execution order between different sub-target tasks without a dependency relationship as parallel.

[0094] That is, steps 601-602 determine the serial or parallel execution order by determining whether there is a dependency relationship between different sub-target tasks.

[0095] Step 603: controlling each sub-target agent executed in a serial manner to output the sub-processing result corresponding to the sub-target task in turn in combination with the personalized preference.

[0096] The sub-processing result output by the preceding sub-target agent executed in a serial manner will be used as another part of the input of the following sub-target agent.

[0097] Step 604: controlling each sub-target agent executed in a parallel manner to output the sub-processing result corresponding to the sub-target task in combination with the personalized preference.

[0098] Steps 603 and 604 provide an implementation scheme for controlling the sub-target agents to execute in the serial or parallel execution order determined above by the execution subject.

[0099] On the basis of any of the above embodiments, considering that the user can modify the content of the part of the sub-processing result output or issue new restriction information at any time during the entire process in which the sub-target agents execute the corresponding sub-target tasks to output the sub-processing result, please refer to Figure 7 , Figure 7 A two-branch schematic diagram for controlling the sub-target agents to output the corresponding sub-processing result is provided in an embodiment of the present disclosure, and the flowchart 700 includes the following steps:

[0100] Step 701: in response to the instruction input box being in the selected state, controlling the currently executed sub-target agent to pause output of the corresponding sub-processing result;

[0101] The instruction input box is used for user to input instructions, and the instruction input box is in the unselected state in the process of outputting the corresponding sub-processing result by the sub-target agent. That is, under the scheme of the embodiments provided in the disclosure, once the instruction input box is in the selected state, it can be represented that the user needs to input some new instructions or needs to interrupt the original sub-processing result output process.

[0102] Step 702: in response to no new instructions being generated in the process of the instruction input box recovering from the selected state to the unselected state, controlling the currently executed sub-target agent to continue outputting the corresponding sub-processing result;

[0103] This step corresponds to a branch case that no new instructions are generated in the process of the instruction input box recovering from the selected state to the unselected state, that is, the user does not input new instructions, and then the sub-processing result output result interrupted by selecting the instruction input box will continue.

[0104] Step 703: in response to new instructions being generated in the process of the instruction input box recovering from the selected state to the unselected state, extracting correction information from the new instructions;

[0105] This step corresponds to another branch case that new instructions are generated in the process of the instruction input box recovering from the selected state to the unselected state, and then the above-mentioned execution agent needs to extract correction information from the new instructions. The correction information includes correction of original information and addition of new information.

[0106] Step 704: determining the sub-target agent affected by the correction information;

[0107] Step 705: controlling the affected sub-target agent to output the corresponding sub-processing result in combination with the correction information.

[0108] Steps 704 and 705 are that the above-mentioned execution agent first determines the sub-target agent affected by the correction information, and then controls the affected sub-target agent to output the corresponding sub-processing result in combination with the correction information. Specifically, if all sub-target agents are affected, the execution can also be started from the process of outputting the sub-processing result by the first executed sub-target agent based on the correction information.

[0109] On the basis of any of the above embodiments, in order to deepen the understanding of how to aggregate the sub-processing results to obtain the target processing result, a specific implementation manner is further provided, including the following steps:

[0110] First, the sub-processing results are spliced in a reverse manner according to a decomposition manner of decomposing each sub-target task from the target task to obtain a reverse splicing result; then, repeated redundant parts in the reverse splicing result are removed, and the expression style of the reverse splicing result is adjusted to obtain a target processing result corresponding to the target task. Further, the expression style of the reverse splicing result can be adjusted to a personalized preference requirement expression style, or the expression style can be adjusted to other required expression styles.

[0111] To deepen understanding, the present disclosure also gives a specific implementation scheme for trying to eliminate the prior art defects and overcome the problems of the prior art in combination with the actual existing prior art defects in specific application scenarios.

[0112] The existing demand satisfaction mode mainly uses a search engine to perform recall, sorting, and heterogeneous result mixing of related web pages through a multi-layer system funnel. Each strategy funnel performs web page sorting based on basic relevance, user feedback behavior, authority, and other information and truncates the output to the next layer. The disadvantage of this system is that it can only match relevance from the content level based on relevance and cannot understand and solve problems for users from the task level. In addition, using web pages can only use fixed content to meet user needs, and when users express personalized needs and want to further meet through multiple rounds, it cannot achieve high-quality and coherent satisfaction effect. Thirdly, the entire matching process is not interpretable to users, and users can only perform final information screening.

[0113] Related large model-based assistant products usually try to solve this problem through a large model-based all-purpose assistant. However, a single large model is difficult to achieve high-quality satisfaction capability in different open fields and different tasks. In professional fields and special scenarios, a single assistant is difficult to establish trust between users and achieve the expected satisfaction effect.

[0114] That is, the related art has the following difficulties in solving the task completion problem in the open field:

[0115] 1) How to efficiently understand and decompose the user's expressed needs as a key step;

[0116] 2) How to find the best satisfaction method for each key step;

[0117] 3) How to output the execution process and integrated results in a user-friendly and complete manner.

[0118] To solve the above difficulties, the present embodiment proposes a user demand end-to-end task completion solution based on a multi-agent collaboration method. The characteristics of this scheme are as follows:

[0119] Before search, it is single-point, one query (query word, query sentence, search word, search sentence) completes one requirement, each requirement needs the user to think to define by himself, which is too high for the user; for a complex task, the user often does not know how to disassemble the task and how to conceive multiple queries to search. The paradigm proposed in this embodiment aims to completely help the user complete the task, and the task completion paradigm can help the user to complete the problem to be solved or the task to be completed directly in one step. Compared with the single query satisfaction of search, the task satisfaction paradigm deeply understands the complete task of the user and disassembles and completes it, which is more in line with the needs of the user.

[0120] This embodiment is based on the cooperation between multiple agents to realize the disassembly and multi-step satisfaction of user demand, and finally automatically integrates into a complete solution that can satisfy the entire task. This solution can complete the following functions:

[0121] 1) Understand the key steps of the requirement, build the main agent, understand the intention of the requirement expressed by the user and disassemble it into multiple sub-task agents that can complete the key steps;

[0122] 2) Sub-task agent generation, based on the understanding of key steps and agent capabilities, the main agent is responsible for generating multiple sub-task agent candidates required to complete the current task;

[0123] 3) Multi-agent collaborative task completion, multiple sub-task agents complete the entire task according to the requirements of the task and the key steps, and the user preference information input, through the scheduling and cooperation of multiple agents, to generate the step-by-step results of task completion;

[0124] 4) Integration and output of the completion scheme, the main agent integrates the task disassembly process and the results of the cooperation of multiple sub-task agents, selects the best presentation form, and presents the final complete process and results.

[0125] That is, this embodiment schedules multiple sub-task agents for cooperation by building a main agent, and completes the entire task through the cooperation between the main agent and multiple sub-task agents, and the cooperation between multiple sub-task agents. Compared with search or intelligent assistant products, it can use subfield agents with better effects to complete and meet better effects;

[0126] And by providing a scheduling and distribution mechanism based on an end-to-end generative large model, the expression ability of the large model is stronger, which can overcome the problems of consistency and global optimality of traditional multi-layer sorting mechanism, and can realize task-oriented global optimal combination scheduling optimization. From the perspective of agent understanding and task understanding, the model can be trained end-to-end.

[0127] At the same time, since the generation of the entire task is personalized, from task decomposition to the completion of subtasks, it is generated according to user input and user preferences, and different solutions are generated for different users, decomposition of subtasks, intelligent agent completion of solutions and presentation effect, to improve user satisfaction.

[0128] The implementation block diagram of the scheme is as follows Figure 8-1 As shown in the figure, the input of the system is: user input requirements and distributable agent set (including the basic settings of these agents, mounted plug-ins, workflow, etc.), and the output is the multi-level generated content of the completed task. The entire system is constructed based on a multi-agent collaborative mode, mainly including two types of agents: 1) Main Agent: overall understanding, decomposition and connection of the task, and giving the final result. 2) Task Agent: the main agent dispatches the task agent to cooperatively complete the task assigned by the main agent.

[0129] The scheme provided by the embodiment can be widely applied to various information satisfaction application programs or independent products. The following takes the scene of the Baidu X application program as an example to show the specific application form of the embodiment:

[0130] 1) The user can switch to the satisfaction form in the invention by one key (the "AI button" in the bottom bar in the following figure) while searching traditionally.

[0131] The user inputs "How to arrange a 2-day tour in Harbin" (for reference, see Figure 8-2 ), and the results of web search and intelligent answer are displayed. By switching to the "Harbin 2-day tour - intelligent private customization" page through the AI button, the personalized results generated for the user are displayed, as follows Figure 8-3 The task completion page displayed contains personalized requirements such as "starting from Beijing", "single white-collar", "special forces travel", and "high cost performance".

[0132] In the task completion page, the working results of multiple task completion agents are displayed, including the host agent "tour planner" and the subtask agents "Harbin travel assistant", "what to wear in Harbin", "Harbin scenic spot linking master", and "Harbin hotel recommendation assistant".

[0133] The results of each task completion agent are displayed through structured, personalized and user-friendly generated UI. The user can download and share the results, as shown in Figure 8-4 , Figure 8-5 and Figure 8-6 .

[0134] The system presents completely personalized solutions for different users, as shown in Figure 8-7 and Figure 8-8Show the effect of "2-day tour of Harbin" for the elderly user.

[0135] Revisit mode, when the user returns to the Baidu X application, even if there is no new query input in the search bar, you can return to the task list that can display history by clicking the entry button, at this time you can still click to enter the corresponding result page, such as Figure 8-9 and Figure 8-10 Other applications can also have the same independent entry (not shown separately).

[0136] From the above examples, it can be seen more specifically that the present embodiment has the following improvements and technical effects relative to the prior art:

[0137] 1) Task satisfaction paradigm, the user's request is no longer treated as an information retrieval, but upgraded to a complete multi-agent task completion satisfaction mode

[0138] The current search engine can do related matching information retrieval for the request, and the intelligent assistant can perform multi-step decomposition and step-by-step search satisfaction for a task, but the multi-agent collaboration for task satisfaction decomposition and satisfaction of the user's complete demand is a new user demand satisfaction mode. And proposed decomposition and multi-agent capability alignment training, which can achieve high consistency between task decomposition and task completion, greatly improving the completion effect of the task.

[0139] 2) Multi-agent candidate generation based on multi-agent collaboration

[0140] According to the results of task decomposition, the target collaborative agent set is generated through the end-to-end generation of the large model, which is a new approach different from the traditional retrieval recall ranking system for intelligent agent sorting. Multi-agent collaboration tasks require multi-agent combination on the requirements of task completion, requiring the large model to deeply understand the task decomposition and understand the capabilities of each agent to accurately depict the difficulty boundary of the problem, and finally obtain the optimal agent set under the current target in a combination optimization way. The difficulties here include:

[0141] Difficulty 1: Deep understanding of agent capabilities, understanding of agent capability boundaries and agent expertise, and fine depiction of subtle differences between capabilities, such as different creative styles of painting agents and subtle differences in photo editing capabilities.

[0142] Difficulty 2: Understanding the matching relationship between tasks and agent capabilities, not semantic similarity matching, but deep matching in task completion capabilities, requiring a deep understanding of tasks, task types and task boundaries, and matching on the open set of agent capabilities and task requirements is a great challenge.

[0143] Difficulty 3: In the case of a given multi-agent collaboration task, obtaining the optimal combination is a combinatorial optimization problem that requires combinatorial optimization.

[0144] 3) The task completion process and results are generated end-to-end by the agent. Currently, it is difficult for search engines to show the retrieval process to users, while intelligent assistants show the thinking results and single-step execution results, and the UI is basically preset in style. Thanks to the multi-agent collaboration task decomposition, combination and multi-subtask completion mode, the embodiment can display the results in an end-to-end generated manner, including the steps, style and results of each step. And the results of each step, the subtask results completed by each subtask agent can be further interacted. This experience is a comprehensive innovation.

[0145] Further reference Figure 9 , as an implementation of the method shown in the above figures, the present disclosure provides an embodiment of a multi-agent collaboration-based explicit result acquisition type task processing apparatus. The device embodiment corresponds to the method embodiment shown in Figure 2 , and the device can be applied to various electronic devices.

[0146] As shown in Figure 9 , the multi-agent collaboration-based explicit result acquisition type task processing apparatus 900 of the embodiment can include: a target task determination unit 901, a task decomposition and corresponding issuing unit 902, a sub-target agent control output unit 903, and a sub-processing result aggregation unit 904. Wherein, the target task determination unit 901 is configured to determine the target task to be solved according to the natural language input of the user; the task decomposition and corresponding issuing unit 902 is configured to decompose the target task into at least one sub-target task containing an explicit result acquisition type task, and issue each sub-target task to at least one sub-target agent containing an explicit result acquisition agent; wherein different sub-agents are used to process different types of sub-tasks; the sub-target agent control output unit 903 is configured to control each sub-target agent to output a sub-processing result corresponding to the belonging sub-target task respectively combining the individualized preferences of the user; wherein the sub-processing result includes an explicit result output by the explicit result acquisition agent corresponding to the belonging explicit result acquisition type task; the sub-processing result aggregation unit 904 is configured to aggregate each sub-processing result to obtain a target processing result corresponding to the target task.

[0147] In the embodiment, in the multi-agent collaboration-based explicit result acquisition type task processing apparatus 900: the specific processing of the target task determination unit 901, the task decomposition and corresponding issuing unit 902, the sub-target agent control output unit 903, and the sub-processing result aggregation unit 904 and the technical effects brought by them can be respectively referred to Figure 2The related descriptions of steps 201-204 in the corresponding embodiments will not be repeated here.

[0148] In some optional implementations of the present embodiment, the target task determination unit 901 can include:

[0149] a conversion subunit configured to convert the natural language input of the user into natural language text;

[0150] an intent recognition subunit configured to perform intent recognition on the natural language text to obtain an intent recognition result including the explicit result obtaining intent;

[0151] a target task determination subunit configured to determine a target task to be solved according to the intent recognition result; wherein the target task includes a task for obtaining an explicitly existing result.

[0152] In some optional implementations of the present embodiment, the intent recognition subunit can be further configured to:

[0153] perform semantic understanding on the natural language text to obtain a semantic understanding result;

[0154] determine a preliminary intent according to the semantic understanding result, and perform personalized correction on the preliminary intent using the personalized preference of the user to obtain the intent recognition result including the explicit result obtaining intent.

[0155] In some optional implementations of the present embodiment, the task decomposition and corresponding issuing unit 902 can include a task decomposition subunit configured to decompose the target task into a plurality of sub-target tasks, which can include:

[0156] a task element determination module configured to determine a plurality of task elements constituting the target task; wherein different task elements correspond to different types of tasks;

[0157] an explicit result obtaining type task decomposition module configured to decompose, from the target task, an explicit result obtaining type task for obtaining an explicitly existing result according to a single task element specifically determined to match the question raised by the user;

[0158] a sub-target task decomposition module configured to decompose the target task into a plurality of sub-target tasks each containing only a single task element; wherein the sub-target task at least contains the explicit result obtaining type task.

[0159] In some optional implementations of the present embodiment, the sub-target task decomposition module can be further configured to:

[0160] In response to the existence of an additional condition corresponding to the task element, the target task is decomposed into a plurality of sub-target tasks each containing only a single task element and a corresponding additional condition; wherein the additional condition includes a specified task processing mode and / or a specified result presentation form.

[0161] In some optional implementations of the embodiment, the multi-agent cooperation-based explicit result acquisition type task processing apparatus 900 can further include a sub-agent construction unit configured to construct different sub-agents, and the sub-agent construction unit can include:

[0162] A task type set determination sub-unit configured to determine a task type set for creating sub-agents;

[0163] A construction sub-unit configured to construct, for each type of task in the task type set, a sub-agent having a task processing capability of the corresponding type of task.

[0164] In some optional implementations of the embodiment, the construction sub-unit can be further configured to:

[0165] In response to the first target type task having at least two task processing modes, constructing, for the first target type task, a sub-agent corresponding to each task processing mode; wherein different task processing modes are abstracted from different processing logics of the same type of task.

[0166] In some optional implementations of the embodiment, the construction sub-unit can be further configured to:

[0167] In response to the second target type task having at least two result presentation forms, constructing, for the second target type task, a sub-agent corresponding to each result presentation form; wherein the result presentation form includes at least one of a text type, a table type, an image type, a video type, and an interactive card type.

[0168] In some optional implementations of the embodiment, the multi-agent cooperation-based explicit result acquisition type task processing apparatus 900 can further include:

[0169] A serial order determination unit configured to determine the execution order between different sub-target tasks having a dependency relationship as serial;

[0170] A parallel order determination unit configured to determine the execution order between different sub-target tasks not having a dependency relationship as parallel;

[0171] Correspondingly, the sub-target agent control output unit 903 can be further configured to:

[0172] The sub-target intelligent agents executed in a serial manner each output a sub-processing result corresponding to the sub-target task to which the sub-target intelligent agent belongs in combination with the personalized preference; wherein the sub-processing result output by a preceding sub-target intelligent agent executed in a serial manner serves as another part of input of a following sub-target intelligent agent.

[0173] The sub-target intelligent agents executed in a parallel manner each output a sub-processing result corresponding to the sub-target task to which the sub-target intelligent agent belongs in combination with the personalized preference.

[0174] In some optional implementations of the embodiment, the multi-intelligent agent cooperation-based explicit result acquisition type task processing apparatus 900 can further include:

[0175] The pause output control unit is configured to control the sub-target intelligent agent currently executed to pause output of the corresponding sub-processing result in response to the instruction input box being in the selected state; wherein the instruction input box is used for user input of instructions, and the instruction input box is in an unselected state in a process in which the sub-target intelligent agent outputs the corresponding sub-processing result.

[0176] The continue output control unit is configured to control the sub-target intelligent agent currently executed to continue output of the corresponding sub-processing result in response to no new instruction being generated in a process in which the instruction input box recovers from the selected state to the unselected state.

[0177] In some optional implementations of the embodiment, the multi-intelligent agent cooperation-based explicit result acquisition type task processing apparatus 900 can further include:

[0178] The correction information extraction unit is configured to extract correction information from the new instruction in response to the new instruction being generated in the process in which the instruction input box recovers from the selected state to the unselected state.

[0179] The affected sub-target intelligent agent determination unit is configured to determine the sub-target intelligent agent affected by the correction information.

[0180] The re-output control unit is configured to control the affected sub-target intelligent agent to output the corresponding sub-processing result in combination with the correction information.

[0181] In some optional implementations of the embodiment, the sub-processing result aggregation unit 904 can include:

[0182] The reverse splicing sub-unit is configured to splice the sub-processing results in a reverse manner according to a decomposition manner by which the target task is decomposed into the sub-target tasks, to obtain a reverse splicing result.

[0183] The optimization processing sub-unit is configured to remove repeated redundant parts in the reverse splicing result, and adjust a presentation style of the reverse splicing result, to obtain a target processing result corresponding to the target task.

[0184] In some optional implementation forms of the embodiment, the optimization processing subunit can comprise a representation style adjustment module configured to adjust the representation style of the inverse splicing result, and the representation style adjustment module can be further configured to:

[0185] adjust the representation style of the inverse splicing result to a representation style required by the personalization preference.

[0186] The embodiment as a device embodiment corresponding to the above-mentioned method embodiment exists, and the device for processing explicit result obtaining type task based on multi-agent cooperation provided by the embodiment adopts an agent cluster formed by a pre-constructed main agent and multiple sub-agents to process the task demand proposed by a user, wherein the main agent is responsible for understanding the task demand of the user, decomposing the overall task demand into multiple sub-target tasks that can be executed by different sub-agents respectively, and each sub-agent processes the sub-task matched with itself according to the mobilization processing of the main agent. That is, through the cooperation between the main agent and each sub-agent, different components of a complex task can be processed by each sub-agent in its own position. Moreover, since different sub-agents are pre-constructed to be respectively dedicated to processing different types of tasks, the scheme of processing task demand by the main agent and multiple sub-agents cooperatively adopted by the disclosure has better processing effect on a single type of task than using a single and all-purpose agent, and the relatively small sub-agents are also convenient for flexible increase and change of corresponding functions, so that the scheme can bring better comprehensive task processing effect with lower comprehensive cost.

[0187] According to the embodiments of the disclosure, the disclosure further provides an electronic device, comprising: at least one processor; and a memory connected with the at least one processor in communication; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to implement the multi-agent cooperation based explicit result obtaining type task processing method described in any of the above embodiments.

[0188] According to the embodiments of the disclosure, the disclosure further provides a readable storage medium, which stores computer instructions for enabling a computer to implement the multi-agent cooperation based explicit result obtaining type task processing method described in any of the above embodiments when the computer executes the computer instructions.

[0189] According to the embodiments of the disclosure, the disclosure further provides a computer program product, which can implement the multi-agent cooperation based explicit result obtaining type task processing method described in any of the above embodiments when the processor executes the computer program.

[0190] Figure 10A schematic block diagram of an example electronic device 1000 that can be used to implement embodiments of the present disclosure is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device may also represent various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the present disclosure described and / or claimed herein.

[0191] like Figure 10 As shown, device 1000 includes a computing unit 1001, which can perform various appropriate actions and processes according to a computer program stored in read-only memory (ROM) 1002 or a computer program loaded from storage unit 1008 into random access memory (RAM) 1003. The RAM 1003 may also store various programs and data required for the operation of device 1000. The computing unit 1001, ROM 1002, and RAM 1003 are interconnected via bus 1004. Input / output (I / O) interface 1005 is also connected to bus 1004.

[0192] Multiple components in device 1000 are connected to I / O interface 1005, including: input unit 1006, such as keyboard, mouse, etc.; output unit 1007, such as various types of monitors, speakers, etc.; storage unit 1008, such as disk, optical disk, etc.; and communication unit 1009, such as network card, modem, wireless transceiver, etc. Communication unit 1009 allows device 1000 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.

[0193] The computing unit 1001 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 1001 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 1001 performs the various methods and processes described above, such as the explicit result acquisition task processing method based on multi-agent cooperation. For example, in some embodiments, the explicit result acquisition task processing method based on multi-agent cooperation can be implemented as a computer software program tangibly contained in a machine-readable medium, such as storage unit 1008. In some embodiments, part or all of the computer program can be loaded and / or installed on device 1000 via ROM 1002 and / or communication unit 1009. When the computer program is loaded into RAM 1003 and executed by the computing unit 1001, one or more steps of the explicit result acquisition task processing method based on multi-agent cooperation described above can be performed. Alternatively, in other embodiments, computing unit 1001 may be configured by any other suitable means (e.g., by means of firmware) to perform explicit result-acquisition task processing methods based on multi-agent cooperation.

[0194] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SoCs), payload-programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.

[0195] Program code for carrying out methods of the present disclosure can be written in any combination of one or more programming languages. The program code can be provided to a processor or controller of a general purpose computer, special purpose computer, or other programmable data processing apparatus to produce a machine, such that the program code, when executed by the processor or controller, produces the functions / operations specified in the flowcharts and / or block diagrams. The program code can be executed entirely on a machine, partially on a machine, partially on a machine as a standalone software package, or entirely on a remote machine or server.

[0196] In the context of the present disclosure, a machine-readable medium can be a tangible medium that contains or stores a program for use by or in connection with an instruction execution system, apparatus, or device. The machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include but is not limited to an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination of the foregoing. More specific examples of the machine-readable storage medium will include one or more lines of electrical connections, portable computer disks, hard disk drives, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or Flash memory), optical fibers, portable compact disc read-only memories (CD-ROMs), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.

[0197] To provide for interaction with a user, the systems and techniques described here can be implemented on a computer having a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user and a keyboard and a pointing device (e.g., a mouse or a trackball) by which the user can provide input to the computer. Other kinds of devices can be used to provide for interaction with a user as well; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form, including acoustic, speech, or tactile input.

[0198] The systems and techniques described here can be implemented in a computing system that includes a back end component (e.g., as a data server), or that includes a middleware component (e.g., an application server), or that includes a front end component (e.g., a user computer having a graphical user interface or a Web browser through which a user can interact with an implementation of the systems and techniques described here), or any combination of such back end, middleware, or front end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include a local area network (LAN), a wide area network (WAN), and the Internet.

[0199] The computer system can include clients and servers. A client and server are generally remote from each other and typically interact through a communication network. The relationship of client and server arises by virtue of computer programs running on the respective computers and having a client-server relationship to each other. A server can be a cloud server, also known as cloud computing server or cloud host, which is a host product in the cloud computing service system to solve the defects of large management difficulty and weak business scalability in traditional physical host and virtual private server (VPS, Virtual Private Server) services.

[0200] The technical scheme of the embodiment of the present disclosure adopts an agent cluster formed by a pre-constructed master agent and multiple sub-agents to process task requirements proposed by a user, wherein the master agent is responsible for understanding the task requirements of the user, decomposing the overall task requirements into multiple sub-target tasks that can be executed by different sub-agents respectively, and each sub-agent processes the sub-task matched with itself according to the mobilization of the master agent. That is, through the cooperation between the master agent and each sub-agent, different components of a complex task can be processed by each sub-agent according to its own function, and because different sub-agents are pre-constructed to be dedicated to processing different types of tasks, the scheme of the master agent and multiple sub-agents cooperating to process task requirements adopted by the present disclosure has better processing effect on a single type of task than using a single and all-purpose agent, and the relatively small sub-agents are also convenient for flexible addition and modification of corresponding functions, so that the scheme can bring better comprehensive task processing effect with lower comprehensive cost.

[0201] It should be understood that various forms of flow shown above can be used to reorder, add or delete steps. For example, each step described in the present disclosure can be executed in parallel, sequentially or in different order, as long as the desired results of the technical scheme of the present disclosure can be achieved, and the present disclosure is not limited herein.

[0202] The above detailed description does not limit the scope of the disclosure. Various modifications, combinations, sub-combinations and alternatives can be made to the detailed description. Any modification, equivalent replacement and improvement etc. made within the spirit and principle of the disclosure shall be included in the scope of the disclosure.

Claims

1. A method for handling explicit result acquisition tasks based on multi-agent collaboration, applied to the main agent, comprising: The target task to be solved is determined based on the user's natural language input; The target task is decomposed into multiple sub-target tasks, each of which includes at least a task with a clear result acquisition. Each sub-target task is then assigned to a sub-target agent that includes at least a clear result acquisition agent. Different sub-agents are used to handle different types of sub-tasks. Each of the sub-target intelligent agents is controlled to output sub-processing results corresponding to its respective sub-target task in combination with the user's personalized preferences. The sub-processing results include explicit results output by the explicit result acquisition intelligent agent that correspond to the explicit result acquisition class task to which it belongs. By summing up the results of each sub-processing step, the target processing result corresponding to the target task is obtained. In response to the instruction input box for the user input instruction being selected, the currently executing sub-target agent is controlled to pause the output of the corresponding sub-processing result, and the instruction input box is in an unselected state during the process of the sub-target agent outputting the corresponding sub-processing result; If no new instruction is generated during the process of the instruction input box returning from the selected state to the unselected state, the currently executing sub-target agent is controlled to continue outputting the corresponding sub-processing result; In response to a new instruction being generated during the process of the instruction input box returning from the selected state to the unselected state, correction information is extracted from the new instruction; the sub-target intelligent agent affected by the correction information is determined; and the affected sub-target intelligent agent is controlled to re-output the corresponding sub-processing result in conjunction with the correction information.

2. The method according to claim 1, wherein, The step of determining the target task to be solved based on the user's natural language input includes: Convert the user's natural language input into natural language text; Intent recognition is performed on the natural language text to obtain intent recognition results that include the explicit intent to obtain the result; The target task to be solved is determined based on the intent recognition result; wherein, the target task includes a task for obtaining a clearly existing result.

3. The method according to claim 2, wherein, The process of performing intent recognition on the natural language text to obtain intent recognition results includes: Perform semantic understanding on the natural language text to obtain the semantic understanding result; The initial intent is determined based on the semantic understanding results, and the initial intent is then personalized and modified using the user's personalized preferences to obtain an intent recognition result that includes the intent to obtain a specific result.

4. The method according to claim 1, wherein, The step of decomposing the target task into multiple sub-target tasks includes: Identify multiple task elements that constitute the target task; wherein, different task elements correspond to different types of tasks; Based on a single task element that specifically determines a clearly existing result that matches the question raised by the user, the target task is decomposed into clearly defined result acquisition tasks for obtaining the clearly defined result. The target task is decomposed into multiple sub-target tasks, each containing only a single task element; wherein, each sub-target task contains at least the explicit result acquisition task.

5. The method according to claim 4, wherein, The step of decomposing the target task into multiple sub-target tasks, each containing only a single task element, includes: In response to the existence of additional conditions corresponding to the task element, the target task is decomposed into multiple sub-target tasks, each containing only a single task element and a corresponding additional condition; wherein the additional conditions include: a specified task processing method and / or a specified result presentation format.

6. The method according to claim 1, wherein, The process of constructing different sub-agents includes: Determine the set of task types to be used to create the sub-agent; For each type of task in the set of task types, a sub-agent with the task processing capabilities of the corresponding type of task is constructed.

7. The method according to claim 6, wherein, For each type of task in the set of task types, a sub-agent with the corresponding task processing capabilities is constructed, including: In response to a first target type task having at least two task processing methods, a sub-agent corresponding to each of the task processing methods is constructed for the first target type task; wherein, the different task processing methods are abstracted from different processing logics for the same type of task.

8. The method according to claim 6, wherein, For each type of task in the set of task types, a sub-agent with the corresponding task processing capabilities is constructed, including: In response to a second target type task having at least two result display formats, a sub-agent is constructed for the second target type task corresponding to each of the result display formats; wherein, the result display formats include at least one of the following: text, table, image, video, and interactive card.

9. The method according to claim 1, further comprising: The execution order of different sub-target tasks with dependencies is determined to be sequential; The execution order of different sub-target tasks that have no dependencies is determined to be parallel; Correspondingly, the control of each sub-target intelligent agent, in conjunction with the user's personalized preferences, outputs sub-processing results corresponding to its respective sub-target task, including: Each sub-target agent executing in the serial manner outputs a sub-processing result corresponding to its respective sub-target task in sequence, taking into account the personalized preferences; wherein, the sub-processing result output by the preceding sub-target agent executing in the serial manner will serve as another part of the input for the subsequent sub-target agent. Each sub-target agent, which is controlled to execute in the parallel manner, outputs a sub-processing result corresponding to its respective sub-target task, based on the personalized preferences.

10. The method according to any one of claims 1-9, wherein, The process of summarizing the results of each sub-processing step to obtain the target processing result corresponding to the target task includes: According to the decomposition method of the sub-target tasks decomposed from the target task, the sub-processing results are spliced ​​together in reverse to obtain the reverse splicing result; By removing duplicate and redundant parts from the reverse splicing result and adjusting the expression style of the reverse splicing result, the target processing result corresponding to the target task is obtained.

11. The method according to claim 10, wherein, The adjustment of the presentation style of the reverse splicing results includes: Adjust the expression style of the reverse splicing result to the expression style required by the personalized preference.

12. A task processing device for explicit result acquisition based on multi-agent cooperation, applied to a main agent, comprising: The target task determination unit is configured to determine the target task to be solved based on the user's natural language input. The task decomposition and corresponding assignment unit is configured to decompose the target task into multiple sub-target tasks, which at least include tasks of obtaining clear results, and to assign each sub-target task to a sub-target agent that at least includes an agent of obtaining clear results; wherein, different sub-agents are used to handle different types of sub-tasks. The sub-target agent control output unit is configured to control each of the sub-target agents to output a sub-processing result corresponding to its respective sub-target task, based on the user's personalized preferences; wherein, the sub-processing result includes the explicit result output by the explicit result acquisition agent corresponding to its respective explicit result acquisition class task; The sub-processing result summarization unit is configured to summarize the sub-processing results to obtain the target processing result corresponding to the target task. The pause output control unit is configured to control the currently executing sub-target agent to pause the output of the corresponding sub-processing result in response to the instruction input box for the user input instruction being in a selected state, wherein the instruction input box is in an unselected state during the process of the sub-target agent outputting the corresponding sub-processing result; The output control unit is configured to control the currently executing sub-target agent to continue outputting the corresponding sub-processing result in response to the fact that no new instruction is generated during the process of the instruction input box returning from the selected state to the unselected state. The correction information extraction unit is configured to extract correction information from the new instruction generated during the process of the instruction input box returning from the selected state to the unselected state; the affected sub-target agent determination unit is configured to determine the sub-target agent affected by the correction information; and the re-output control unit is configured to control the affected sub-target agent to re-output the corresponding sub-processing result in combination with the correction information.

13. The apparatus according to claim 12, wherein, The target task determination unit includes: The conversion subunit is configured to convert the user's natural language input into natural language text; The intent recognition subunit is configured to perform intent recognition on the natural language text to obtain an intent recognition result that includes the explicit intent to obtain the result. The target task determination subunit is configured to determine the target task to be solved based on the intent recognition result; wherein the target task includes a task for obtaining a definite result.

14. The apparatus according to claim 13, wherein, The intent recognition subunit is further configured to: Perform semantic understanding on the natural language text to obtain the semantic understanding result; The initial intent is determined based on the semantic understanding results, and the initial intent is then personalized and modified using the user's personalized preferences to obtain an intent recognition result that includes the intent to obtain a specific result.

15. The apparatus according to claim 12, wherein, The task decomposition and corresponding distribution unit includes a task decomposition subunit configured to decompose the target task into multiple sub-target tasks, the task decomposition subunit comprising: The task element determination module is configured to determine multiple task elements that constitute the target task; wherein, different task elements correspond to different types of tasks. The explicit result acquisition task decomposition module is configured to decompose explicit result acquisition tasks from the target task based on a single task element that specifically determines an explicitly existing result that matches the question raised by the user. The sub-target task decomposition module is configured to decompose the target task into multiple sub-target tasks that each contain only a single task element; wherein, the sub-target task contains at least the explicit result acquisition task.

16. The apparatus according to claim 15, wherein, The sub-target task decomposition module is further configured to: In response to the existence of additional conditions corresponding to the task element, the target task is decomposed into multiple sub-target tasks, each containing only a single task element and a corresponding additional condition; wherein the additional conditions include: a specified task processing method and / or a specified result presentation format.

17. The apparatus of claim 12, further comprising: Sub-agent building units are configured to construct different sub-agents, the sub-agent building units comprising: The task type set determines the sub-unit, which is configured to determine the set of task types used to create the sub-agent; Each sub-unit is configured to be a task of a certain type in the set of task types, and each sub-agent is constructed with the task processing capabilities of the corresponding task type.

18. The apparatus according to claim 17, wherein, The building subunit is further configured to: In response to a first target type task having at least two task processing methods, a sub-agent corresponding to each of the task processing methods is constructed for the first target type task; wherein, the different task processing methods are abstracted from different processing logics for the same type of task.

19. The apparatus according to claim 18, wherein, The building subunit is further configured to: In response to a second target type task having at least two result display formats, a sub-agent is constructed for the second target type task, corresponding to each of the result display formats; wherein, the result display formats include at least one of: text, table, image, video, and interactive card.

20. The apparatus of claim 12, further comprising: The serial order determination unit is configured to determine the execution order of different sub-target tasks that have dependencies as serial; The parallel order determination unit is configured to determine the execution order of different sub-target tasks that have no dependency relationship as parallel; Correspondingly, the sub-target intelligent agent control output unit is further configured to: Each sub-target agent executing in the serial manner outputs a sub-processing result corresponding to its respective sub-target task in sequence, taking into account the personalized preferences; wherein, the sub-processing result output by the preceding sub-target agent executing in the serial manner will serve as another part of the input for the subsequent sub-target agent. Each sub-target agent, which is controlled to execute in the parallel manner, outputs a sub-processing result corresponding to its respective sub-target task, based on the personalized preferences.

21. The apparatus according to any one of claims 12-20, wherein, The sub-processing result aggregation unit includes: The reverse splicing subunit is configured to splice the sub-processing results in a reverse manner according to the decomposition method of the sub-target tasks decomposed from the target task, so as to obtain the reverse splicing result; The optimization processing subunit is configured to remove duplicate and redundant parts from the reverse splicing result and adjust the expression style of the reverse splicing result to obtain the target processing result corresponding to the target task.

22. The apparatus according to claim 21, wherein, The optimization processing subunit includes a description style adjustment module configured to adjust the description style of the reverse splicing result, and the description style adjustment module is further configured to: Adjust the expression style of the reverse splicing result to the expression style required by the personalized preference.

23. An electronic device, comprising: At least one processor; as well as A memory communicatively connected to the at least one processor; wherein, The memory stores instructions executable by the at least one processor, which, when executed by the at least one processor, enables the at least one processor to perform the explicit result acquisition task processing method based on multi-agent cooperation as described in any one of claims 1-11.

24. A non-transitory computer-readable storage medium storing computer instructions for causing the computer to perform the explicit result-acquiring task processing method based on multi-agent cooperation as described in any one of claims 1-11.

25. A computer program product comprising a computer program that, when executed by a processor, implements the steps of the method for processing explicit result acquisition-type tasks based on multi-agent cooperation according to any one of claims 1-11.

Citation Information

Patent Citations

  • Information recommendation method based on generative large language model and related device

    CN116932733A

  • Aerial document analysis and test case generation system driven by large model agent

    CN117909243A