Interaction method and device based on agent cooperation, agent and storage medium

Through the interactive method of intelligent agent collaboration, the intention of user input information is understood and the pattern is detected, the matching interactive elements are called to interact with the user, and the intelligent agent is controlled to perform tasks. This solves the problem of low accuracy of reply content in the existing technology and realizes more efficient and accurate personalized content generation.

CN120654731AActive Publication Date: 2025-09-16BEIJING BAIDU NETCOM SCI & TECH CO LTD

Patent Information

Application Number
CN202510865498.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-25
Publication Date
2025-09-16
Estimated Expiration
2045-06-25

AI Technical Summary

Technical Problem

It is difficult for existing technologies to accurately generate reply content that meets the user's personalized needs based on user input information, resulting in low reply accuracy.

Method used

Through an interactive method based on agent collaboration, the initial input information is understood in terms of intent, the first agent is used to detect the required attributes of the interaction mode, the matching interactive elements are called to interact with the target object, and the second agent is controlled to perform the target task and push the execution results to meet user needs.

Benefits of technology

It improves the interaction efficiency and the accuracy of reply content, and meets the personalized needs of users.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120654731A_ABST
    Figure CN120654731A_ABST
Patent Text Reader

Abstract

The invention provides an interaction method and device based on agent cooperation, an agent and a storage medium, belongs to the technical field of artificial intelligence, particularly relates to the technical fields of deep learning, large models, agents, AIGC and the like, and is applied to application scenes such as wisdom education, video production, wisdom medical treatment and the like. The intelligent agent cooperation-based interaction method comprises the following steps: performing intention understanding on initial input information for a target object to obtain an initial intention understanding result; performing interaction mode demand detection on the initial intention understanding result by using a first agent to obtain an interaction mode demand attribute; an interaction element matched with the interaction mode demand attribute is called to interact with the target object, interaction information representing the demand intention of the target object is obtained, and the interaction element is used for responding to the interaction behavior; and controlling the second agent to execute the target task based on the interaction information to obtain an execution result, and pushing the execution result to the target object.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of artificial intelligence technology, in particular to technical fields such as deep learning, large models, intelligent agents, and AIGC (Artificial Intelligence Generated Content), and is applied to application scenarios such as smart education, video production, and smart medical care. Background Art

[0002] With the rapid development of artificial intelligence technology, it is possible to interact with users based on large models with generative capabilities, such as multimodal large models, so that the needs of users can be met based on the generated content of the large models. Summary of the Invention

[0003] The present disclosure provides an interactive method, device, agent and storage medium based on agent collaboration.

[0004] According to one aspect of the present disclosure, an interaction method based on agent collaboration is provided, including: understanding the intent of initial input information for a target object to obtain an initial intent understanding result; using a first agent to detect interaction mode requirements on the initial intent understanding result to obtain interaction mode requirement attributes; calling an interaction element that matches the interaction mode requirement attributes to interact with the target object to obtain interaction information that characterizes the target object's requirement intention, and the interaction element is used to respond to the interaction behavior; controlling a second agent to execute a target task based on the interaction information to obtain an execution result, and pushing the execution result to the target object.

[0005] According to another aspect of the present disclosure, an interaction device based on agent collaboration is provided, including: an intention understanding module, which is used to understand the intention of initial input information for a target object and obtain an initial intention understanding result; a detection module, which is used to use a first agent to perform interaction mode requirement detection on the initial intention understanding result and obtain interaction mode requirement attributes; a calling module, which is used to call an interaction element that matches the interaction mode requirement attributes to interact with the target object and obtain interaction information that characterizes the target object's requirement intention, and the interaction element is used to respond to the interaction behavior; an execution module, which is used to control the second agent to perform the target task based on the interaction information, obtain the execution result, and push the execution result to the target object.

[0006] According to another aspect of the present disclosure, an artificial intelligence agent is provided, configured to execute the method provided according to the embodiment of the present disclosure.

[0007] According to another aspect of the present disclosure, an electronic device is provided, comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the method provided according to an embodiment of the present disclosure.

[0008] According to another aspect of the present disclosure, a non-transitory computer-readable storage medium storing computer instructions is provided, wherein the computer instructions are used to enable the computer to execute the method provided according to an embodiment of the present disclosure.

[0009] According to another aspect of the present disclosure, a computer program product is provided, including a computer program, which implements the method provided according to the embodiment of the present disclosure when executed by a processor.

[0010] It should be understood that the contents described in this section are not intended to identify the key or important features of the embodiments of the present disclosure, nor are they intended to limit the scope of the present disclosure. Other features of the present disclosure will become readily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0011] The accompanying drawings are provided to facilitate a better understanding of the present invention and do not constitute a limitation of the present disclosure.

[0012] Figure 1 Schematically illustrates an exemplary system architecture to which an interactive method and apparatus based on agent collaboration according to an embodiment of the present disclosure can be applied;

[0013] Figure 2 The flowchart of the interaction method based on intelligent agent collaboration according to an embodiment of the present disclosure is schematically shown;

[0014] Figure 3 Schematically illustrates an interaction scene diagram of an interaction method based on agent collaboration according to an embodiment of the present disclosure;

[0015] Figure 4 Schematically illustrates an interaction scene diagram of an interaction method based on agent collaboration according to another embodiment of the present disclosure;

[0016] Figure 5 Schematically illustrates an interaction scene diagram of an interaction method based on agent collaboration according to yet another embodiment of the present disclosure;

[0017] Figure 6 The schematic diagram schematically shows the principle of intelligent agent collaboration according to an embodiment of the present disclosure;

[0018] Figure 7 Schematically shows a schematic diagram of the principle based on intelligent agent collaboration according to another embodiment of the present disclosure;

[0019] Figure 8 Schematically shows a block diagram of an interactive device based on agent collaboration according to an embodiment of the present disclosure;

[0020] Figure 9 A block diagram schematically illustrates a structure of an artificial intelligence agent according to an embodiment of the present disclosure; and

[0021] Figure 10 A schematic block diagram of an example electronic device that can be used to implement the agent collaboration-based interaction method according to an embodiment of the present disclosure is shown. DETAILED DESCRIPTION

[0022] The following description of exemplary embodiments of the present disclosure is made in conjunction with the accompanying drawings, including various details of the embodiments of the present disclosure to facilitate understanding. These details should be considered as merely exemplary. Therefore, those skilled in the art will recognize that various changes and modifications may be made to the embodiments described herein without departing from the scope and spirit of the present disclosure. Similarly, for the sake of clarity and conciseness, descriptions of well-known functions and structures are omitted in the following description.

[0023] In the technical solution disclosed herein, the acquisition, storage and application of user personal information involved comply with the provisions of relevant laws and regulations, take necessary confidentiality measures, and do not violate public order and good morals.

[0024] The inventors discovered that with the rapid development of artificial intelligence technology, relevant internet platforms can process user input information based on large models with generative capabilities and meet user needs by outputting generative content such as text and video. However, users' actual needs are often highly personalized and highly differentiated, making it difficult for generative models to accurately generate responses based on user input, resulting in low response accuracy.

[0025] The embodiments of the present disclosure provide an interactive method, device, agent, and storage medium based on agent collaboration. The interactive method based on agent collaboration includes: performing intent understanding on initial input information for a target object to obtain an initial intent understanding result; using a first agent to perform interaction mode requirement detection on the initial intent understanding result to obtain interaction mode requirement attributes; calling an interactive element that matches the interaction mode requirement attributes to interact with the target object to obtain interaction information representing the target object's requirement intention, and the interactive element is used to respond to the interactive behavior; and controlling a second agent to execute a target task based on the interaction information to obtain an execution result, and pushing the execution result to the target object.

[0026] According to the embodiments of the present disclosure, by understanding the intention of the initial input information, and then using the first intelligent agent to detect based on the initial intention understanding result to obtain the interaction mode requirement attributes that can meet the actual demand scenario of the target object, the interaction mode that matches the scenario of the target object's needs can be more accurately determined through multi-level intention understanding. In this way, the matching interaction elements are called to interact with the target object based on the interaction mode requirement attributes of the target object, so that the target object can perform interactive behavior input based on the interactive elements that are adapted to the required scenario and can accurately represent the actual demand intention, so that the second intelligent agent can be controlled to perform the target task based on the interactive information to obtain the execution result that meets the demand intention. By pushing the execution result to the target object, the interaction efficiency is improved, and the accuracy of the reply content is improved to meet the personalized needs of the target object.

[0027] Figure 1 An exemplary system architecture to which an interaction method and apparatus based on agent collaboration can be applied according to an embodiment of the present disclosure is schematically shown.

[0028] It should be noted that Figure 1 The examples shown are merely examples of system architectures to which the embodiments of the present disclosure may be applied, to help those skilled in the art understand the technical content of the present disclosure, but do not mean that the embodiments of the present disclosure may not be applied to other devices, systems, environments, or scenarios. For example, in another embodiment, an exemplary system architecture to which the agent-based collaborative interaction method and apparatus may be applied may include a terminal device, but the terminal device may implement the agent-based collaborative interaction method and apparatus provided by the embodiments of the present disclosure without interacting with a server.

[0029] like Figure 1 As shown, the system architecture 100 according to this embodiment may include terminal devices 101, 102, 103, a network 104, and a server 105. The network 104 is used as a medium for providing communication links between the terminal devices 101, 102, 103 and the server 105. The network 104 may include various connection types, such as wired and / or wireless communication links, etc.

[0030] Users can use terminal devices 101, 102, and 103 to interact with server 105 via network 104 to receive or send messages, etc. Various communication client applications can be installed on terminal devices 101, 102, and 103, such as knowledge reading applications, web browser applications, search applications, instant messaging tools, email clients, and / or social platform software (for example only).

[0031] The terminal devices 101 , 102 , and 103 may be various electronic devices having a display screen and supporting web browsing, including but not limited to smart phones, tablet computers, laptop computers, and desktop computers.

[0032] Server 105 may be a server that provides various services, such as a background management server (for example only) that supports content browsed by users using terminal devices 101, 102, and 103. The background management server may analyze and process received data such as user requests, and feed back processing results (e.g., web pages, information, or data obtained or generated based on user requests) to the terminal device.

[0033] Server 105 can be a cloud server, also known as a cloud computing server or cloud host. A cloud server is a host product within the cloud computing service system. It addresses the management difficulties and poor scalability of traditional physical hosts and VPS services ("Virtual Private Servers" or "VPS"). Server 105 can also be a server for a distributed system or a server integrated with blockchain.

[0034] The server 105 can obtain the first agent or the second agent by deploying the generative model. It should be noted that the first agent or the second agent can be deployed in the same server, or the first agent or the second agent can also be deployed in different servers.

[0035] It should be noted that the agent-cooperation-based interaction method provided in the embodiment of the present disclosure can generally be executed by the terminal device 101, 102, or 103. Accordingly, the agent-cooperation-based interaction device provided in the embodiment of the present disclosure can also be set in the terminal device 101, 102, or 103.

[0036] Alternatively, the interactive method based on intelligent agent collaboration provided by the embodiment of the present disclosure may also be generally executed by the server 105. Accordingly, the interactive device based on intelligent agent collaboration provided by the embodiment of the present disclosure may generally be set in the server 105. The interactive method based on intelligent agent collaboration provided by the embodiment of the present disclosure may also be executed by a server or server cluster that is different from the server 105 and can communicate with the terminal devices 101, 102, 103 and / or the server 105. Accordingly, the interactive device based on intelligent agent collaboration provided by the embodiment of the present disclosure may also be set in a server or server cluster that is different from the server 105 and can communicate with the terminal devices 101, 102, 103 and / or the server 105.

[0037] It should be understood that Figure 1The number of terminal devices, networks and servers in the embodiment is merely illustrative. Any number of terminal devices, networks and servers may be provided as required.

[0038] Figure 2 The flowchart of the interaction method based on intelligent agent collaboration according to an embodiment of the present disclosure is schematically shown.

[0039] like Figure 2 As shown, the interaction method based on agent collaboration includes operations S210~S240.

[0040] In operation S210 , intent understanding is performed on initial input information for a target object to obtain an initial intent understanding result.

[0041] In operation S220, the first agent is used to perform interaction mode requirement detection on the initial intention understanding result to obtain interaction mode requirement attributes.

[0042] In operation S230 , an interactive element matching the interaction mode requirement attribute is called to interact with the target object to obtain interaction information representing the target object's requirement intention.

[0043] In operation S240, the second agent is controlled to execute the target task based on the interaction information, an execution result is obtained, and the execution result is pushed to the target object.

[0044] According to embodiments of the present disclosure, the first agent or the second agent can be constructed based on a large model with generative functions, such as a generative model. The first agent or the second agent can be deployed in a server, or can be deployed separately in multiple servers that are communicatively connected. The agents involved in embodiments of the present disclosure can perform data analysis based on the model algorithm or model parameters of the large model to obtain processing results.

[0045] In some embodiments, Figure 2 In the interactive method of the illustrated embodiment, operations S210 to S240 may be performed by the first agent as the execution subject, which may be understood as deploying one or more servers of the first agent to execute the interactive method provided by the embodiment of the present disclosure.

[0046] In some embodiments, Figure 2 In the interactive method of the illustrated embodiment, operation S220 is performed by the first agent, and operations S210, S230, and S240 can be performed by other servers or computing devices other than the first agent.

[0047] It should be noted that, in order to facilitate the explanation of the interaction method based on agent collaboration provided by the embodiment of the present disclosure, the embodiment of the present disclosure takes the first agent as the execution subject of the interaction method, and is not used to limit the interaction method to only be executed by the first agent.

[0048] According to an embodiment of the present disclosure, the initial input information of the target object can be data in any modality, such as voice data, text data, image data, etc. The embodiment of the present disclosure does not limit the specific data modality type of the initial input information.

[0049] According to embodiments of the present disclosure, the initial input information can be processed by invoking a large language model to obtain an initial intent understanding result. Alternatively, the initial input information can be subjected to intent understanding through other methods such as keyword extraction to obtain keywords representing the initial intent understanding result. The embodiments of the present disclosure do not limit the specific method of obtaining the initial intent understanding result.

[0050] According to embodiments of the present disclosure, the initial intent understanding result may be the initial level of demand intent represented by the initial input information. For example, the initial intent understanding result may represent the theme intent of the target object's demand intent, such as travel planning or health consultation. Furthermore, the initial intent understanding result may also include associated demand intents related to the theme intent.

[0051] For example, if the initial input information is "Plan a travel guide for City A," the initial intent understanding results may include thematic intent related to the theme of travel planning, destination intent related to the destination "City A," and "Plan a travel guide" representing the task requirement intent. Task requirement intent and destination intent can be related requirement intents.

[0052] According to an embodiment of the present disclosure, interactive elements are used to respond to interactive behaviors. For example, dialog interactive elements can be used to respond to user text input behaviors, and text input information is determined as interactive information. By utilizing the first intelligent agent to perform interactive mode demand detection on the initial intent understanding result, the interactive mode or interactive element represented by the interactive mode demand attribute can be matched with the subject intent and associated demand intention represented by the initial intent understanding result, so that the target object can naturally express the actual demand information by performing interactive behaviors on the interactive elements. In this way, it is possible to interact with the target object by calling interactive elements that match the interactive model demand attributes, and obtain interactive information that can represent the deep needs of the target object by responding to the interactive behavior of the target object through the interactive elements, thereby realizing demand intention mining of the target object, so that the interactive information can more accurately represent the multi-dimensional demand attributes of the target object.

[0053] It should be noted that invoking an interactive element to interact with a target object may include displaying the interactive element in a display interface to provide an interactive behavior object for the target object, and obtaining interactive information by responding to the target object's interactive behavior, such as text input, click, or voice input, through the interactive element. An interactive element can be a component that can respond to interactive behaviors performed by the target object. The interactive element can be displayed in a display interface, for example, based on element types such as icons, dialog boxes, and page links. The embodiments of this disclosure do not limit the specific display method of the interactive element.

[0054] According to an embodiment of the present disclosure, controlling the second agent to execute a target task based on interaction information may include determining, through the second agent, a target task that satisfies the actual demand intention of the target object based on the interaction information, and obtaining an execution result that satisfies the demand of the target object by executing the target task. The execution result may be, for example, a travel guide copy, a travel planning page containing images, text, and connections to attraction ticket reservations, etc. By invoking an interactive element that matches the interaction mode demand attribute of the target object through the first agent to interact with the target object to obtain interaction information, and by controlling the second agent to execute a target task that matches the demand intention represented by the interaction information, demand intention mining and task execution can be achieved based on a multi-agent collaborative approach, thereby improving the degree of match between the execution result and the actual demand intention of the target object, and further improving the satisfaction of the reply content.

[0055] In some embodiments, the target task includes at least one of the following: a travel strategy planning task, a text generation task, a graphic element generation task, and an image editing task.

[0056] According to an embodiment of the present disclosure, a copywriting generation task may be, for example, using a large language model to perform semantic understanding on interactive information to generate any type of copywriting information, such as lecture copywriting, tourism promotion copywriting, etc.

[0057] According to embodiments of the present disclosure, graphic element generation tasks can be used to generate graphic elements such as icons and cartoon characters displayed in videos or images. Graphic element generation tasks can include specific code generation tasks. For example, based on the text content described in interactive information, such as "Generate a red ball and make it move randomly in the image," a second agent can be controlled to use a large language model to understand the text content and generate a specific code script, which is then executed to control the random movement of the red ball in the image.

[0058] In some embodiments, image editing tasks may include performing editing operations on a specified image, such as cropping, compressing, and adjusting the color tone. By controlling the second agent to invoke a specified image editing tool to perform the image editing task, the image can be precisely edited according to the requirements represented by the interaction information, thereby improving the accuracy of the image editing results as the execution results.

[0059] In some embodiments, the travel strategy planning task may include multiple subtasks, such as scenic spot location query, travel route planning, and travel trajectory generation. By controlling the second agent to execute the travel strategy planning task based on the interaction information, the second agent can invoke various functional tools to execute different subtasks. This allows the travel strategy information determined based on the execution results of the multiple subtasks to more accurately meet the user's actual needs and enhance the user's travel experience.

[0060] It should be noted that the number of target tasks involved in the embodiment of the present disclosure can be one or more, and the second intelligent agent can be controlled to execute multiple target tasks to obtain multiple intermediate execution results, and generate execution results for pushing to the target object based on the multiple intermediate execution results.

[0061] For example, an image editing task for image editing includes multiple target tasks, namely, a specified image query task, an object detection task, and an image cropping task. The image query task can be used to query multiple images in a specified directory. The object detection task is used to identify images representing "kittens" from multiple images and determine the image regions within the images that represent "kittens." The image cropping task can crop the image regions representing "kittens" to obtain a variety of image patches representing "kittens."

[0062] According to an embodiment of the present disclosure, performing intent understanding on initial input information for a target object to obtain an initial intent understanding result may include: calling a large language model to perform structured semantic understanding on the initial input information to obtain multiple intent understanding information with structured attributes.

[0063] In some embodiments, the plurality of intent understanding information of the structured attributes may be information content corresponding to the intent attribute type.

[0064] For example, if the intent attribute type is the theme intent type, the information content corresponding to the theme intent type can be "travel." For another example, if the intent attribute type is the associated demand intent type "destination," the information content corresponding to the associated demand intent type "destination" can be "City A."

[0065] In some embodiments, the first agent can be used to call the large language model to perform semantic understanding on the initial input information based on the prompt word to obtain multiple intention understanding information with structured attributes. The prompt word is used to control the large language model to perform semantic understanding on the initial input information based on multiple intention attribute types represented by the structured attributes as targets, so that the multiple intention understanding information can correspond to multiple intention attribute types respectively. In this way, the multiple intention understanding information with structured attributes can clearly represent the subject intention and associated demand intention of the target object according to the intention attribute type. In this way, the first agent can detect the interaction scene or interaction mode required by the target object through multiple intention understanding information with structured attributes, and accurately determine the interaction mode requirement attributes.

[0066] According to an embodiment of the present disclosure, the intent understanding information includes an intent attribute type and an intent field corresponding to the intent attribute type. The intent field may represent field data such as a text field, an identification field, etc. corresponding to the intent attribute type.

[0067] In some embodiments, the large language model can perform semantic understanding on multiple intention understanding information with structured attributes, and based on the control of prompt words, enable the intention understanding information to indicate that the intent field corresponding to the preset requirement attribute type indicated by the prompt word in the initial input information has demand intention representation defects such as semantic ambiguity and text missing. Therefore, the first intelligent agent can determine the need to provide the missing information content of the preset requirement attribute type in the interaction mode requirement attribute by processing the demand intention representation defects represented by the intention understanding information, thereby improving the accuracy and completeness of the interaction information's actual demand intention representation for the target object, and thus ensuring that the execution result of the target task can meet the actual needs of the user.

[0068] For example, the initial input information is "travel planning", and the multiple intent understanding information only contains the theme intent "travel", and the intent field corresponding to the associated demand intent type "destination" is missing information content. Therefore, the first intelligent agent can perform interaction mode demand detection based on multiple intent understanding information under the control of the prompt word, and determine that the interaction mode demand attribute represents multiple interaction elements for the travel planning scenario, and the multiple interaction elements include destination interaction option elements representing "City A" and "City B". Therefore, the target object can interact with the interaction option elements to determine the intent field corresponding to the target object's destination demand, thereby improving the execution result of the tourism strategy planning task to meet the target object's actual travel needs.

[0069] In some embodiments, multiple intent understanding information may have a hierarchical structure. For example, a topic intent type and a topic intent field corresponding to the topic intent type may have a first hierarchy, and an associated intent type and an intent field corresponding to the associated intent type may have a second hierarchy.

[0070] In some embodiments, multiple intent fields corresponding to associated intent types may also have different structural levels. For example, the intent fields corresponding to the travel duration intent type include "summer vacation" and "one week." The intent field "summer vacation" has the second level, and the intent field "one week" has the third level.

[0071] According to the embodiments of the present disclosure, by representing the structured attributes of multiple intention understanding information through structural hierarchical relationships, the first intelligent agent can more accurately and deeply understand the current degree of clarity of the target object's demand for the subject intention according to the structural hierarchical relationship, and can call matching interactive elements to interact with the target object by outputting the interaction mode demand attributes, so that the target object can interact according to the interaction mode that is adapted to the degree of clarity of the demand, and enable multiple interaction information to progressively represent the actual needs of the target object according to the structural hierarchical relationship, thereby realizing demand intention mining for the target object, so that subsequent execution results can meet the potential demand intentions expressed by the target object through the interaction information, and also avoid the second intelligent agent from performing tasks by processing information with unclear or missing demand intentions, which affects the accuracy of the reply content, thereby improving the degree of matching between the execution results and the actual needs of the target object.

[0072] According to an embodiment of the present disclosure, the interaction method based on agent collaboration further includes: determining multiple types of interaction elements based on interaction mode requirement attributes, and the order in which the multiple types of interaction elements are called.

[0073] For example, the interaction mode requirement attribute represents a scenario mode for interacting with an intent field whose subject intent type is "travel", and indicates that intent fields such as "destination", "travel duration", and "number of tourists" are missing. The multiple interaction elements corresponding to the interaction mode requirement attribute may include prompt icons for prompting intent attribute types such as "destination", "travel duration", and "number of tourists", and multiple prompt icons are displayed in the order of "destination", "travel duration", and "number of tourists" to guide the target object to input interaction information according to the target object's way of thinking by targeting the prompt icons, so as to improve the accuracy and adequacy of the expression of the interaction information for the target object's actual needs and intentions, thereby improving the degree of match between the target task and execution results and the target object's actual needs and intentions, and improving the target object's interaction experience and interaction efficiency.

[0074] In some embodiments, the interaction mode requirement attribute also characterizes that the semantic accuracy requirement of the interaction information needs to meet a preset accuracy condition. For example, in interaction scenarios corresponding to subject intentions such as health consultation and precision instrument repair consultation, a specific interaction mode is needed to control the interaction information input by the target object to meet the preset accuracy condition, so as to avoid errors in the creation and execution of the target task due to the semantic accuracy of the interaction information not meeting the requirements. Therefore, for the interaction mode requirement attribute that characterizes that the semantic accuracy requirement of the interaction information needs to meet the preset accuracy condition, matching interaction elements can be determined to make the interaction information meet the preset accuracy condition.

[0075] In some embodiments, calling an interactive element that matches the interaction mode requirement attribute to interact with the target object may include: calling an interactive option element representing preset option data to interact with the target object.

[0076] According to an embodiment of the present disclosure, the interactive option element is used to determine target option data in preset option data in response to the interactive behavior, and the interactive information may include the target option data.

[0077] For example, multiple interactive option elements could represent preset option data corresponding to specific medication names, such as "Drug A," "Drug B," and "Drug C." The target subject can interact with the interactive option element representing "Drug A" to determine that the target option data is "Drug A." This prevents the target subject from entering unclear or incorrect medication names, which could lead to deviations in the second agent's performance of the health consultation task. This ensures that the accuracy of the health consultation results meets the target subject's actual needs.

[0078] Figure 3 The interactive scene diagram of the interactive method based on intelligent agent collaboration according to an embodiment of the present disclosure is schematically shown.

[0079] like Figure 3As shown, the application scenario includes a first interactive interface 300, in which the initial input information input by the target object is displayed, "Is there any way to relieve frequent headaches?" The first agent understands the intent of the initial input information by calling a large language model, and the initial understanding result obtained includes multiple intention understanding information with structured attributes. The first agent is used to perform interaction mode requirement detection on multiple intention understanding information, and the obtained interaction mode requirement attributes can represent the semantic clarity of the interaction information related to the health consultation theme intention, which needs to meet the preset accuracy conditions. In this way, multiple interaction option elements can be determined as interaction elements that match the interaction mode requirement attributes. Among them, the multiple interaction option elements can be multiple option icons displayed in the target option box 310, and the text content in the option icon represents the preset option data. For example, the target option box 310 may include option icons representing the preset option data "Drug A" and "Drug B", respectively.

[0080] The target subject can interact with the option icons in the target option box 310. In response to the interaction, the option icons can determine the target option data, namely, "Attack duration: 10 minutes to 20 minutes," "Headache frequency: 2 to 3 times per week," "Taking the following medication: Medication B," and "Pain location: Top of head." The second agent can then execute the target task based on the interaction information including the target option data, resulting in the following execution result: "Based on the information you entered, the preliminary conclusion is that the headache is caused by overwork. It is recommended to rest early and exercise more daily, and to visit a relevant specialist hospital for examination and consultation."

[0081] According to an embodiment of the present disclosure, the interaction mode requirement attribute can also represent whether the emotional requirement attribute of the initial intent understanding result meets the preset emotional requirement condition. For example, the specified theme requirement intentions such as tourism theme requirement intention and problem-solving theme requirement intention can meet the preset emotional requirement condition.

[0082] In some embodiments, the first agent can use the initial intention understanding results and historical interaction preferences of the target object to perform semantic understanding to determine whether the target object needs to interact with the target object in an interaction mode that meets preset emotional demand conditions.

[0083] For example, the first intelligent agent determines that the target object has a high frequency of inputting demand information for interactive scenarios such as videos and games by understanding historical interaction preferences. Therefore, the preset emotional demand conditions can be met by outputting the emotional demand attributes of the initial intention understanding results to represent the emotional demand attributes.

[0084] In some embodiments, calling an interactive element that matches the interaction mode requirement attribute to interact with the target object may further include: calling a virtual object interactive element to interact with the target object.

[0085] According to an embodiment of the present disclosure, a virtual object interaction element is used to interact with a target object by driving a virtual object to perform oral broadcasting. For example, an authorized cartoon character can be driven to explain a tourist destination, so as to enable the target object to interact more immersively through the cartoon character's oral broadcasting video element, and guide the target object to input the interaction information of the actual demand intention while fully exploring the target object's demand intention. In this way, the target object can meet the immersive emotional needs in the immersive interaction process provided by the virtual object interaction element, and input the interaction information under the condition of high interaction efficiency. This will further improve the accuracy and comprehensiveness of the second intelligent agent's understanding of the interaction information, improve the accuracy of the execution results, and enhance the interaction experience of the target object.

[0086] Figure 4 An interaction scene diagram of an interaction method based on agent collaboration according to another embodiment of the present disclosure is schematically shown.

[0087] like Figure 4 As shown, the application scenario includes a second interactive interface 400, in which a virtual object 420 corresponding to the emotional demand attribute represented by the interactive mode demand attribute is displayed. The virtual object 420 can make corresponding yoga movements during the oral explanation of the yoga course, so that the target object can interact with the first guidance element 411 and the second guidance element 412 indicated by the virtual object 420 in an immersive interactive scene to input the course items to be experienced or the course time to be booked. In addition, the target object can also input voice or text information into the input box 431 for the virtual object 420 to provide interactive information. It should be understood that the virtual object interactive elements may include the virtual object 420, the first guidance element 411, and the second guidance element 412. The input box 431 can be introduced into the current interactive mode as a dialogue interactive element to provide a rich interactive input entry.

[0088] The second agent can execute the target task based on the interaction information determined by the target object's interaction behavior with respect to the virtual object interaction element to generate a course experience list for the target object as an execution result.

[0089] According to an embodiment of the present disclosure, the interaction mode requirement attribute further represents whether the semantic ambiguity of the initial intention understanding result satisfies a preset ambiguity condition.

[0090] For example, if the number of missing intent fields corresponding to multiple related demand intents in the "travel" theme intent in the initial intent understanding result meets the preset quantity threshold, it can be determined that the initial intent understanding result meets the preset fuzziness condition.

[0091] In some embodiments, the first agent can also be controlled to perform semantic understanding of the initial input information based on a preset prompt word representing a preset ambiguity condition, so that the output initial intention understanding result includes "semantic ambiguity of the initial input information."

[0092] For example, if the initial input is "solving a problem," the preset prompts could include "need to understand the semantics of the problem, determine the subject and knowledge points." The initial semantic understanding result could then indicate "the semantics of the problem-solving requirements are ambiguous and require further questioning."

[0093] In some embodiments, calling an interactive element that matches the interaction mode requirement attribute to interact with the target object may further include: calling a dialog interaction element to interact with the target object.

[0094] According to embodiments of the present disclosure, a dialog interaction element is used to guide the target object's demand intent. For example, the dialog interaction element may include question information generated by the first agent based on the demand intent type indicated by the interaction mode demand attribute. By displaying the question information, the target object is prompted to perform an interactive action such as text input or voice input according to the text information to obtain information content corresponding to the demand intent type, thereby improving the ability of the interactive information to fully represent the actual demand intent.

[0095] For example, a dialog interaction element might include a question box displaying "What's your question?". If the target user enters the interactive information "The question is 2+2×3=", the next dialog interaction element's corresponding question box might also be displayed, "Do you need a detailed explanation?" This allows the dialog interaction element to be used to further inquire about the target user's intended needs, guiding them to enter interactive information that reflects their intended needs through interaction.

[0096] Figure 5 An interaction scene diagram of an interaction method based on agent collaboration according to yet another embodiment of the present disclosure is schematically shown.

[0097] like Figure 5As shown, this application scenario includes a third interactive interface 500. The third interactive interface 500 may include the user's initial input information, "Need to plan a trip." The initial intent understanding result may only include the subject intent, "Travel," with the semantic ambiguity level meeting a preset ambiguity level condition. Consequently, a first dialogue interaction element 511 may be displayed in the third interactive interface 500 to inquire about the target subject's associated demand intent, guiding the target subject to enter the interactive information for the first dialogue interaction element 511, "City B, three-day trip," through dialog box 501. Based on the current interaction information and the initial intent understanding result, the first agent may further invoke a second dialogue interaction element 512, "How many people are traveling? What types of attractions do you prefer?", to continue guiding the target subject's demand intent, allowing the target subject to enter subsequent interactive information through interactive operations. This allows the second agent to execute the travel strategy planning task based on the user's input interaction information, resulting in a travel planning solution displayed as an execution result on the third interactive interface 500, thereby pushing the execution result to the target subject.

[0098] According to an embodiment of the present disclosure, controlling the second intelligent agent to perform a target task based on interaction information may include: using the first intelligent agent to perform semantic understanding of the interaction information to obtain task requirement information; controlling the second intelligent agent to perform task arrangement based on the task requirements to obtain the target task, and controlling the second intelligent agent to perform the target task.

[0099] In some embodiments, the first agent can be used to perform semantic understanding of the interaction information, the initial intent understanding results, and the initial input information to determine the task requirement information. The task requirement information can, for example, include data in any modality, such as text or images, that describes the execution conditions, execution parameters, and other task elements of the target task. In this way, the demand intent represented by the interaction information and the initial input information can be converted into relevant information corresponding to the target task, thereby improving the efficiency of the second agent in creating and executing the target task and improving the accuracy of the execution results.

[0100] According to an embodiment of the present disclosure, the second agent may have task orchestration capabilities. The second agent may determine one or more target tasks that can meet the task requirement information by semantically understanding the task requirement information. This can be achieved by dynamically identifying the target object's needs by leveraging the first agent's intention understanding and interaction pattern matching capabilities, and by leveraging the second agent's task orchestration and task execution capabilities. This allows for multi-agent collaboration to determine execution results that meet the target object's diverse needs, thereby improving interaction satisfaction.

[0101] In some embodiments, controlling the second agent to perform task scheduling based on the task requirements to obtain the target task includes controlling the planning agent to perform task scheduling based on the task requirement information to obtain multiple target tasks and dependencies between the multiple target tasks. The dependencies may represent the execution order of the multiple target tasks.

[0102] According to an embodiment of the present disclosure, the second agent includes a planning agent, which sends multiple target tasks and dependencies between the multiple target tasks to the execution agent. The execution agent obtains execution results by executing the multiple target tasks according to the dependencies.

[0103] For example, the planning agent can arrange tasks based on the task requirement information related to the intention of the lecture video, and obtain multiple target tasks such as the problem-solving task, the text generation task, the subtitle generation task, and the oral script generation task. At the same time, the execution order of multiple target tasks is determined as a dependency relationship. By calling the execution agent to execute multiple target tasks according to the dependency relationship, a lecture video can be generated as an execution result. The text generation task is executed based on the problem-solving execution result of the problem-solving task, and the lecture text is obtained as an intermediate execution result. The oral script generation task is executed based on the lecture text, and the voice driving data and action driving data used to drive the virtual object to perform oral broadcasting are obtained as the intermediate execution result of the oral script generation task. The subtitle generation task is executed based on the voice driving data, and the subtitle file is determined as the intermediate specified result of the subtitle generation task. Video synthesis is performed through multiple intermediate execution results to generate a lecture video.

[0104] According to an embodiment of the present disclosure, the control planning agent performs task scheduling based on task requirement information, which may include: the control planning agent performs task scheduling on the task requirement information and the functional description information for the candidate execution agents, determines multiple target tasks, and the execution agents among the candidate execution agents for performing the target tasks.

[0105] According to an embodiment of the present disclosure, the dependency relationship can also characterize the calling order of multiple execution agents. The functional description information for the candidate execution agent can represent the task type, task parameters, execution result accuracy, and other information that can be performed by the candidate execution agent to describe the task execution capability boundary of the candidate execution agent. By controlling the planning agent to process the task requirement information and the functional description information for the candidate execution agent to perform task scheduling, the functional description information can be used as a prompt to control the planning agent to fully understand the capabilities of the candidate execution agent, and assign multiple target tasks that meet the task requirement information to the execution agent whose execution capability boundary can meet the task requirement conditions of the target task, thereby improving the accuracy of task scheduling, avoiding the mismatch between the execution capability boundary of the execution agent and the requirement conditions of the target task, resulting in errors in the intermediate execution results, thereby improving the accuracy of the execution results, and further improving the interactive satisfaction of the target object.

[0106] According to the embodiments of the present disclosure, a planning agent is used to perform task scheduling based on functional description information and task requirement information obtained by dynamically understanding the multi-level intent of the target object. The planning agent selects the optimal combination of execution agents to work together to execute multiple target tasks through task scheduling methods such as task decomposition and stage division, while comprehensively evaluating the execution capability boundaries, task adaptability, and current resource status of each candidate execution agent. In addition, the planning agent also clearly represents the execution order of the target tasks and the interactive collaboration process between multiple execution agents by outputting dependencies, so as to ensure the smoothness of information transmission between the execution agents, avoid conflicts in the execution of target tasks and redundant use of computing resources or storage resources, and effectively improve the efficiency and quality of task completion.

[0107] Figure 6 The schematic diagram schematically shows the principle of agent collaboration according to an embodiment of the present disclosure.

[0108] like Figure 6 As shown, target object 601 interacts with first agent 610 by operating a computer device. First agent 610 uses a large language model to perform semantic understanding on the initial input information provided by target object 601, obtaining an initial intent understanding result. The initial intent understanding result includes multiple pieces of intent understanding information with structured attributes. For example, the multiple pieces of intent understanding information may include "Theme intent type: travel; Destination: City A; Time: summer vacation."

[0109] The first agent 610 performs interaction mode requirement detection on the initial intent understanding results and determines whether the interaction mode requirement attributes meet the emotional requirement conditions. The first agent 610 then invokes the virtual object interaction element to interact with the target object 601 to determine the interaction information. The first agent 610 generates task requirement information by processing the interaction information and the initial intent understanding results and transmits the task requirement information to the second agent terminal 620.

[0110] The second intelligent agent terminal 620 includes multiple second intelligent agents, namely a planning intelligent agent 6211 and multiple candidate execution intelligent agents, wherein the multiple candidate execution intelligent agents are the first candidate execution intelligent agent 6221, the second candidate execution intelligent agent 6222...the nth candidate execution intelligent agent 622n. The planning intelligent agent 6211 arranges tasks based on the task requirement information and obtains two target tasks, namely the scenic spot ticket reservation task and the travel route planning task. The planning intelligent agent 6211 calls the first candidate execution intelligent agent 6221 and the second candidate execution intelligent agent 6222 as two execution intelligent agents to respectively execute the scenic spot ticket reservation task and the travel route planning task, thereby obtaining two intermediate execution results, namely the ticket reservation QR code and the travel route map. The first intelligent agent 610 synthesizes the ticket reservation QR code and the travel route map to obtain the target page as the execution result and pushes it to the target object 601.

[0111] According to an embodiment of the present disclosure, the interaction method based on agent collaboration may also include: responding to an execution feedback request for an execution result, controlling the second agent to update the target task based on the feedback information carried in the execution feedback request, and executing the updated target task to obtain a feedback execution result.

[0112] In some embodiments, the execution feedback request may be determined by an interactive operation performed by the target object on the execution result. For example, the feedback information carried in the execution feedback request may include "modify the itinerary and replace attraction A with attraction B".

[0113] In some embodiments, the first agent can process feedback information and interaction information to generate updated task requirement information, and the first agent can control the second agent to perform task scheduling based on the updated task requirement information to obtain an updated target task. The second agent can then execute the updated target task to obtain a new execution result as the feedback execution result.

[0114] For example, the feedback execution result may be a travel strategy planning document that replaces attraction A with attraction B.

[0115] In some embodiments, controlling the second agent to update the target task based on the feedback information carried in the execution feedback request may include: using the first agent to perform semantic understanding of the execution results and the target task based on the feedback information to obtain feedback task requirement information; and controlling the second agent to perform task scheduling based on the feedback task requirement information to obtain an updated target task.

[0116] In one embodiment, the first agent can process the execution results, target tasks, and interaction information using the feedback information as a prompt, outputting updated task requirement information as feedback task requirement information. This allows the task requirement information to be updated based on the feedback prompt semantics of the execution results and target tasks, enabling the feedback task requirement information to more accurately represent the target object's current needs. Furthermore, the execution agent can be controlled to execute the updated target task, resulting in a feedback execution result that satisfies the target object's current needs.

[0117] Figure 7 A schematic diagram of the principle of agent collaboration according to another embodiment of the present disclosure is schematically shown.

[0118] like Figure 7 As shown, the first target object 701 inputs initial input information and interaction information by interacting with the first agent 711. The first agent 711 determines the execution result by executing the method provided in the embodiment of the present disclosure, and pushes the execution result to the first target object 701. The first target object 701 controls the first agent to generate a feedback execution result based on the adjustment information content of the feedback information by transmitting feedback information to the first agent 711, and pushes the feedback execution result to the first target object 701. The first agent 711 can also store the feedback information and feedback execution result in the knowledge base to implement the update of the task execution strategy. For example, the parameters of the large model used for the first agent 711 can be fine-tuned to obtain the updated first agent 712.

[0119] The updated first agent 712 can interact with other target objects 702 based on the optimized capabilities to push the optimized execution results to them.

[0120] According to an embodiment of the present disclosure, the interactive behavior of the target object can be responded to through interactive elements such as a graphical interface and a natural language interactive window to determine feedback information. Thus, the first agent and the second agent can collaborate to process the feedback information and execute the updated target task, so as to timely adjust or modify the target object's demand intention for the execution result. In addition, the feedback information can also be stored in the knowledge base for the target object so that the first agent and the second agent can update the model parameters and task execution strategy by learning the feedback information, and adjust the capability boundary through the agent's automated backtracking, thereby continuously optimizing the interaction process with the target object.

[0121] Figure 8 A block diagram of an interactive device based on agent collaboration according to an embodiment of the present disclosure is schematically shown.

[0122] like Figure 8 As shown, the interactive device 800 based on intelligent agent collaboration includes: an intention understanding module 810, a detection module 820, a calling module 830 and an execution module 840.

[0123] The intention understanding module 810 is used to understand the intention of the initial input information for the target object and obtain an initial intention understanding result.

[0124] The detection module 820 is used to use the first agent to perform interaction mode requirement detection on the initial intention understanding result to obtain the interaction mode requirement attribute.

[0125] The calling module 830 is used to call the interactive element that matches the interactive mode requirement attribute to interact with the target object and obtain the interactive information representing the target object's requirement intention. The interactive element is used to respond to the interactive behavior.

[0126] The execution module 840 is used to control the second agent to execute the target task based on the interaction information, obtain the execution result, and push the execution result to the target object.

[0127] According to an embodiment of the present disclosure, the interaction mode requirement attribute represents that the semantic accuracy requirement of the interaction information satisfies a preset accuracy condition.

[0128] The first calling unit is configured to call an interactive option element representing preset option data to interact with a target object, wherein the interactive option element is configured to determine target option data in the preset option data in response to an interactive behavior, and the interactive information includes the target option data.

[0129] According to an embodiment of the present disclosure, the interaction mode requirement attribute further indicates that the semantic ambiguity of the initial intention understanding result meets a preset ambiguity condition.

[0130] The second calling unit is used to call the dialogue interaction element to interact with the target object, and the dialogue interaction element is used to guide the target object's demand intention.

[0131] According to an embodiment of the present disclosure, the interaction mode requirement attribute also represents that the emotional requirement attribute of the initial intention understanding result meets the preset emotional requirement condition; wherein, the calling module further includes a third calling unit.

[0132] The third calling unit is used to call the virtual object interaction element to interact with the target object, wherein the virtual object interaction element is used to interact with the target object by driving the virtual object to perform oral broadcasting.

[0133] According to an embodiment of the present disclosure, the interactive device based on agent collaboration further includes a first determination module.

[0134] The first determining module is used to determine multiple types of interactive elements and a calling order of the multiple types of interactive elements based on the interaction mode requirement attribute.

[0135] According to an embodiment of the present disclosure, the intention understanding module includes a first obtaining unit.

[0136] The first acquisition unit is used to call the large language model to perform structured semantic understanding on the initial input information to obtain multiple intention understanding information with structured attributes.

[0137] According to an embodiment of the present disclosure, a plurality of intent understanding information has a structural hierarchical relationship, and the intent understanding information includes an intent attribute type and an intent field corresponding to the intent attribute type.

[0138] According to an embodiment of the present disclosure, the execution module includes: a task requirement information obtaining unit and a control unit.

[0139] The task requirement information obtaining unit is used to use the first agent to perform semantic understanding on the interaction information to obtain task requirement information.

[0140] The control unit is used to control the second agent to perform task arrangement based on the task requirement information, obtain the target task, and control the second agent to execute the target task.

[0141] According to an embodiment of the present disclosure, the control unit includes a control subunit.

[0142] The control subunit is used to control the planning agent to perform task scheduling based on task requirement information to obtain multiple target tasks and the dependency relationships between the multiple target tasks; wherein, the second agent includes a planning agent, and the planning agent sends multiple target tasks and the dependency relationships between the multiple target tasks to the execution agent, and the execution agent obtains the execution results by executing the multiple target tasks according to the dependency relationships.

[0143] According to an embodiment of the present disclosure, the control subunit is further configured to: control the planning agent to perform task scheduling on the task requirement information and the functional description information for the candidate execution agents, determine multiple target tasks, and the execution agents among the candidate execution agents for executing the target tasks, wherein the dependency relationship represents the calling order of the multiple execution agents.

[0144] According to an embodiment of the present disclosure, the interactive device based on agent collaboration further includes an updating module.

[0145] The update module is used to respond to the execution feedback request for the execution result, control the second agent to update the target task based on the feedback information carried in the execution feedback request, and execute the updated target task to obtain the feedback execution result.

[0146] According to an embodiment of the present disclosure, the updating module includes: a semantic understanding unit and a second obtaining unit.

[0147] The semantic understanding unit is used to use the first agent to perform semantic understanding of the execution results and target tasks based on the feedback information to obtain feedback task requirement information.

[0148] The second obtaining unit is used to control the second intelligent agent to perform task scheduling based on the feedback task requirement information to obtain an updated target task.

[0149] According to an embodiment of the present disclosure, the target task includes at least one of the following: a travel strategy planning task, a text generation task, a graphic element generation task, and an image editing task.

[0150] Figure 9 The structural block diagram of an artificial intelligence agent according to an embodiment of the present disclosure is schematically shown.

[0151] In the embodiments of the present disclosure, Figure 9 As shown, the AI ​​agent 900 may include an input module 910 , a processing module 920 and an output module 930 .

[0152] Input module 910, for receiving initial input information and interaction information;

[0153] Processing module 920 is used to determine the initial intention understanding result based on the initial input information received by the input module, use the first agent to detect the interaction mode requirement of the initial intention understanding result, and obtain the interaction mode requirement attribute; call the interaction element that matches the interaction mode requirement attribute to interact with the target object to obtain interaction information that represents the target object's requirement intention; control the second agent to execute the target task based on the interaction information, and obtain the execution result as output information

[0154] The output module 930 is used to output the output information obtained by the processing module.

[0155] According to an embodiment of the present disclosure, the input module 910 is responsible for receiving or perceiving information such as queries, requests, instructions, signals, or data from the outside world (e.g., a user or the external environment) and converting it into a format that can be understood and processed by the AI ​​agent 900. The input module 910 is the primary link for the AI ​​agent 900 to interact with the outside world. It enables the AI ​​agent 900 to efficiently and accurately obtain the necessary "sensory" information from the outside world and respond to this information.

[0156] In an example, the input module 910 may input the initial input information and interaction information described above.

[0157] In the example, the processing module 920 is the core support for the AI ​​agent 900 to handle complex tasks. The processing module 920 can execute the above-described interaction method based on agent collaboration.

[0158] In this example, the performance of processing module 920 may be closely related to the large model underlying AI agent 900. To fully leverage the capabilities of the large model, the internal structure of processing module 920 may be designed to be highly configurable and extensible to handle a variety of different tasks and requirements in real-world scenarios.

[0159] Understandably, while large language models possess excellent language understanding and generation capabilities, like humans, they are limited in the tasks they can perform without tools. However, when AI Agent 900 is empowered with tool-based capabilities, it can perform tasks such as mathematical calculations using a calculator, data analysis using Python, and weather forecasting using search engines.

[0160] In an example, the output module 930 may output the execution result described above.

[0161] The AI ​​agent 900 according to the embodiment of the present disclosure can simply and effectively improve the level of intelligence, and enhance flexibility and versatility.

[0162] According to an embodiment of the present disclosure, the present disclosure also provides an electronic device, a readable storage medium, and a computer program product.

[0163] According to an embodiment of the present disclosure, an electronic device includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the method described above.

[0164] According to an embodiment of the present disclosure, a non-transitory computer-readable storage medium storing computer instructions is provided, wherein the computer instructions are used to cause a computer to execute the method described above.

[0165] According to an embodiment of the present disclosure, a computer program product includes a computer program, and when the computer program is executed by a processor, the computer program implements the method described above.

[0166] Figure 10 A schematic block diagram of an example electronic device of an agent-based collaborative interaction method that can be used to implement an embodiment of the present disclosure is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital assistants, cellular phones, smart phones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present disclosure described and / or claimed herein.

[0167] like Figure 10 As shown, device 1000 includes a computing unit 1001, which can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 1002 or a computer program loaded from a storage unit 1008 into a random access memory (RAM) 1003. RAM 1003 may also store various programs and data required for the operation of device 1000. Computing unit 1001, ROM 1002, and RAM 1003 are connected to each other via a bus 1004. An input / output (I / O) interface 1005 is also connected to bus 1004.

[0168] Various components in device 1000 are connected to I / O interface 1005, including an input unit 1006, such as a keyboard, mouse, etc.; an output unit 1007, such as various types of displays, speakers, etc.; a storage unit 1008, such as a magnetic disk, optical disk, etc.; and a communication unit 1009, such as a network card, modem, wireless communication transceiver, etc. The communication unit 1009 allows device 1000 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.

[0169] The computing unit 1001 can be any general-purpose and / or specialized processing component with processing and computing capabilities. Some examples of the computing unit 1001 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various specialized artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 1001 performs the various methods and processes described above, such as the agent-based interaction method. For example, in some embodiments, the agent-based interaction method can be implemented as a computer software program tangibly embodied in a machine-readable medium, such as the storage unit 1008. In some embodiments, part or all of the computer program can be loaded and / or installed onto the device 1000 via the ROM 1002 and / or the communication unit 1009. When the computer program is loaded into the RAM 1003 and executed by the computing unit 1001, one or more steps of the agent-based interaction method described above can be performed. Alternatively, in other embodiments, the computing unit 1001 may be configured to execute the agent collaboration-based interaction method in any other appropriate manner (eg, by means of firmware).

[0170] Various embodiments of the systems and techniques described above can be implemented in digital electronic circuit systems, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), system-on-chip systems (SOCs), programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include being implemented in one or more computer programs that are executable and / or interpreted on a programmable system that includes at least one programmable processor, which can be a special purpose or general purpose programmable processor that can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit data and instructions to the storage system, the at least one input device, and the at least one output device.

[0171] The program code for implementing the method of the present disclosure can be written in any combination of one or more programming languages. These program codes can be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing device so that when the program code is executed by the processor or controller, the functions / operations specified in the flow chart and / or block diagram are implemented. The program code can be executed entirely on the machine, partially on the machine, as a stand-alone software package, partially on the machine and partially on a remote machine, or entirely on a remote machine or server.

[0172] In the context of the present disclosure, a machine-readable medium may be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium may be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium may include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination of the foregoing. More specific examples of machine-readable storage media may include an electrical connection based on one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), optical fibers, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0173] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user can provide input to the computer. Other types of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).

[0174] The systems and techniques described herein can be implemented in a computing system that includes back-end components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes front-end components (e.g., a user computer with a graphical user interface or a web browser through which a user can interact with implementations of the systems and techniques described herein), or a computing system that includes any combination of such back-end components, middleware components, or front-end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include a local area network (LAN), a wide area network (WAN), and the Internet.

[0175] A computer system may include a client and a server. The client and server are generally remote from each other and typically interact through a communication network. The client-server relationship arises by virtue of computer programs running on the respective computers and having a client-server relationship to each other. The server may be a cloud server, a server in a distributed system, or a server integrated with a blockchain.

[0176] It should be understood that the various forms of the processes shown above can be used to reorder, add, or delete steps. For example, the steps described in this disclosure can be performed in parallel, sequentially, or in a different order, as long as the desired results of the technical solutions disclosed in this disclosure can be achieved. This is not a limitation herein.

[0177] The above specific embodiments do not constitute a limitation on the scope of protection of this disclosure. Those skilled in the art will appreciate that various modifications, combinations, sub-combinations, and substitutions may be made based on design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this disclosure shall be included within the scope of protection of this disclosure.

Claims

1. An interactive method based on agent collaboration, comprising: Performing intention understanding on the initial input information for the target object to obtain an initial intention understanding result; Using the first agent to perform interaction mode requirement detection on the initial intention understanding result to obtain interaction mode requirement attributes; Invoking an interactive element that matches the interaction mode requirement attribute to interact with the target object to obtain interaction information representing the target object's requirement intention, wherein the interactive element is used to respond to the interaction behavior; Based on the interaction information, the second agent is controlled to execute the target task, obtain the execution result, and push the execution result to the target object.

2. The method according to claim 1, wherein The interaction mode requirement attribute represents that the semantic accuracy requirement of the interaction information satisfies a preset accuracy condition; The calling of the interactive element matching the interaction mode requirement attribute to interact with the target object includes: An interactive option element representing preset option data is called to interact with the target object, wherein the interactive option element is used to determine target option data in the preset option data in response to the interactive behavior, and the interactive information includes the target option data.

3. The method according to claim 1 or 2, wherein: The interaction mode requirement attribute further represents that the semantic ambiguity of the initial intention understanding result satisfies a preset ambiguity condition; The calling of the interactive element matching the interaction mode requirement attribute to interact with the target object includes: A dialogue interaction element is called to interact with the target object, where the dialogue interaction element is used to guide the target object's demand intention.

4. The method according to claim 1 or 2, wherein: The interaction mode requirement attribute also represents that the emotional requirement attribute of the initial intention understanding result meets the preset emotional requirement condition; The calling of the interactive element matching the interaction mode requirement attribute to interact with the target object includes: Call a virtual object interaction element to interact with the target object, wherein the virtual object interaction element is used to interact with the target object by driving the virtual object to perform oral broadcasting.

5. The method according to claim 1, wherein The method further comprises: Based on the interaction mode requirement attributes, multiple types of interaction elements and a calling order of the multiple types of interaction elements are determined.

6. The method according to claim 1, wherein The performing intention understanding on the initial input information for the target object to obtain an initial intention understanding result includes: The large language model is called to perform structured semantic understanding on the initial input information to obtain multiple intention understanding information with structured attributes.

7. The method according to claim 6, wherein: There is a structural hierarchical relationship between the multiple intention understanding information, and the intention understanding information includes an intention attribute type and an intention field corresponding to the intention attribute type.

8. The method according to claim 1, wherein The controlling the second agent to perform the target task based on the interaction information includes: Using the first agent to perform semantic understanding on the interaction information to obtain task requirement information; The second agent is controlled to perform task arrangement based on the task requirement information to obtain the target task, and the second agent is controlled to execute the target task.

9. The method according to claim 8, wherein The controlling the second agent to perform task arrangement based on the task requirement information to obtain the target task includes: The control planning agent performs task arrangement based on the task requirement information to obtain a plurality of target tasks and dependencies between the plurality of target tasks; Among them, the second intelligent agent includes the planning intelligent agent, and the planning intelligent agent sends multiple target tasks and the dependency relationships between the multiple target tasks to the execution intelligent agent. The execution intelligent agent obtains the execution result by executing the multiple target tasks according to the dependency relationships.

10. The method according to claim 9, wherein: The control planning agent performs task scheduling based on the task requirement information, including: Control the planning agent to perform task scheduling on the task requirement information and the functional description information for the candidate execution agents, determine a plurality of the target tasks, and an execution agent among the candidate execution agents for executing the target tasks, wherein the dependency relationship represents the calling order of the plurality of the execution agents.

11. The method according to claim 1, wherein The method further comprises: In response to an execution feedback request for the execution result, the second agent is controlled to update the target task based on the feedback information carried in the execution feedback request, and the updated target task is executed to obtain a feedback execution result.

12. The method according to claim 11, wherein The controlling the second agent to update the target task based on the feedback information carried in the execution feedback request includes: Using the first agent to perform semantic understanding on the execution result and the target task based on the feedback information to obtain feedback task requirement information; and The second agent is controlled to perform task scheduling based on the feedback task requirement information to obtain the updated target task.

13. The method according to claim 1, wherein The target tasks include at least one of the following: Travel strategy planning tasks, copywriting generation tasks, graphic element generation tasks, and image editing tasks.

14. An interactive device based on intelligent agent collaboration, comprising: An intention understanding module is used to understand the intention of the initial input information for the target object and obtain an initial intention understanding result; a detection module, configured to use the first agent to perform interaction mode requirement detection on the initial intent understanding result to obtain interaction mode requirement attributes; A calling module, configured to call an interactive element that matches the interaction mode requirement attribute to interact with the target object, and obtain interaction information representing the target object's requirement intention, wherein the interactive element is used to respond to the interaction behavior; An execution module is used to control the second agent to execute the target task based on the interaction information, obtain the execution result, and push the execution result to the target object.

15. The device according to claim 14, wherein The interaction mode requirement attribute represents that the semantic accuracy requirement of the interaction information satisfies a preset accuracy condition; Wherein, the calling module includes: The first calling unit is configured to call an interactive option element representing preset option data to interact with the target object, wherein the interactive option element is configured to determine target option data in the preset option data in response to the interactive behavior, and the interaction information includes the target option data.

16. The device according to claim 14 or 15, wherein The interaction mode requirement attribute further represents that the semantic ambiguity of the initial intention understanding result satisfies a preset ambiguity condition; Wherein, the calling module includes: The second calling unit is used to call a dialogue interaction element to interact with the target object, and the dialogue interaction element is used to guide the target object's demand intention.

17. An artificial intelligence agent configured to perform the method according to any one of claims 1 to 13.

18. An electronic device comprising: at least one processor; as well as a memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method according to any one of claims 1 to 13.

19. A non-transitory computer-readable storage medium storing computer instructions, wherein: The computer instructions are used to cause the computer to execute the method according to any one of claims 1 to 13.

20. A computer program product comprising a computer program which, when executed by a processor, implements the method according to any one of claims 1 to 13.

Citation Information

Patent Citations

  • Multi-agent hierarchical autonomous decision-making method and system in complex environment

    CN115840892A

  • Information interaction method and device based on large language model and electronic equipment

    CN118093801A

  • Interaction method and device and computer readable storage medium

    CN118964688A

  • Intelligent agent-based recommendation method and device, electronic equipment and intelligent agent

    CN119129724A

  • Intelligent agent interaction method and device based on task self-feedback

    CN119476344A

Cited By

  • Object processing method, agent, service gateway and capability opening system

    CN121125846A

  • Interaction method and device based on intelligent agent, intelligent agent and electronic equipment

    CN121434452A

  • Interaction method and device based on intelligent agent, intelligent agent and storage medium

    CN121434453A

  • Interaction method and device based on large model, intelligent agent and storage medium

    CN121501956A