Interaction method and device based on agent cooperation, agent, and storage medium

CN120654731BActive Publication Date: 2026-08-07BEIJING BAIDU NETCOM SCI & TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
BEIJING BAIDU NETCOM SCI & TECH CO LTD
Filing Date
2025-06-25
Publication Date
2026-08-07

AI Technical Summary

Benefits of technology

[0010] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of this disclosure, nor is it intended to limit the scope of this disclosure. Other features of this disclosure will become readily apparent from the following description.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120654731B_ABST
    Figure CN120654731B_ABST
Patent Text Reader

Abstract

The present disclosure provides an interaction method and device based on agent cooperation, an agent and a storage medium, in the field of artificial intelligence technology, especially in the fields of deep learning, large model, agent, AIGC and the like, and is applied to application scenarios such as smart education, video production, smart medical treatment and the like. The interaction method based on agent cooperation comprises: performing intention understanding on initial input information for a target object to obtain an initial intention understanding result; performing interaction mode demand detection on the initial intention understanding result by using a first agent to obtain an interaction mode demand attribute; calling an interaction element matched with the interaction mode demand attribute to interact with the target object to obtain interaction information representing a demand intention of the target object, the interaction element being used to respond to an interaction behavior; controlling a second agent to perform a target task based on the interaction information to obtain an execution result, and pushing the execution result to the target object.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to the field of artificial intelligence technology, particularly to the fields of deep learning, large models, intelligent agents, and AIGC (Artificial Intelligence Generated Content), and is applied to application scenarios such as smart education, video production, and smart healthcare. Background Technology

[0002] With the rapid development of artificial intelligence technology, large models with generative capabilities, such as multimodal large models, can be used to interact with users, thereby meeting user needs by generating content based on these large models. Summary of the Invention

[0003] This disclosure provides an interaction method, apparatus, agent, and storage medium based on agent collaboration.

[0004] According to one aspect of this disclosure, an interaction method based on intelligent agent collaboration is provided, comprising: performing intent understanding on initial input information for a target object to obtain an initial intent understanding result; using a first intelligent agent to perform interaction mode requirement detection on the initial intent understanding result to obtain an interaction mode requirement attribute; invoking an interaction element that matches the interaction mode requirement attribute to interact with the target object to obtain interaction information representing the target object's requirement intent, wherein the interaction element is used to respond to the interaction behavior; controlling a second intelligent agent to execute a target task based on the interaction information to obtain an execution result, and pushing the execution result to the target object.

[0005] According to another aspect of this disclosure, an interaction device based on intelligent agent collaboration is provided, comprising: an intent understanding module, used to understand the intent of initial input information for a target object and obtain an initial intent understanding result; a detection module, used to use a first intelligent agent to detect interaction mode requirements based on the initial intent understanding result and obtain interaction mode requirement attributes; an invocation module, used to invoke an interaction element that matches the interaction mode requirement attributes to interact with the target object and obtain interaction information representing the target object's requirement intent, wherein the interaction element is used to respond to the interaction behavior; and an execution module, used to control a second intelligent agent to execute a target task based on the interaction information, obtain an execution result, and push the execution result to the target object.

[0006] According to another aspect of this disclosure, an artificial intelligence agent is provided, configured to perform a method provided according to embodiments of this disclosure.

[0007] According to another aspect of this disclosure, an electronic device is provided, comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, the instructions being executed by the at least one processor to enable the at least one processor to perform a method provided according to an embodiment of this disclosure.

[0008] According to another aspect of this disclosure, a non-transitory computer-readable storage medium is provided storing computer instructions, wherein the computer instructions are used to cause the computer to perform a method provided according to an embodiment of this disclosure.

[0009] According to another aspect of this disclosure, a computer program product is provided, including a computer program that, when executed by a processor, implements the method provided according to embodiments of this disclosure.

[0010] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of this disclosure, nor is it intended to limit the scope of this disclosure. Other features of this disclosure will become readily apparent from the following description. Attached Figure Description

[0011] The accompanying drawings are provided to better understand this solution and do not constitute a limitation of this disclosure. Wherein:

[0012] Figure 1 The illustration schematically shows an exemplary system architecture for applying agent-based interaction methods and apparatus according to embodiments of the present disclosure;

[0013] Figure 2 A flowchart illustrating an interaction method based on agent collaboration according to an embodiment of the present disclosure is shown schematically.

[0014] Figure 3 The diagram illustrates an interaction scenario of an agent-based collaborative interaction method according to an embodiment of the present disclosure.

[0015] Figure 4 The diagram illustrates an interaction scenario of an agent-based collaborative interaction method according to another embodiment of the present disclosure.

[0016] Figure 5 The diagram illustrates an interaction scenario of an agent-based collaborative interaction method according to yet another embodiment of the present disclosure.

[0017] Figure 6 The illustration shows a schematic diagram of the principle of agent-based collaboration according to an embodiment of the present disclosure;

[0018] Figure 7 The illustration shows a schematic diagram of the principle of agent-based collaboration according to another embodiment of the present disclosure;

[0019] Figure 8 A block diagram schematically illustrates an interactive device based on agent collaboration according to an embodiment of the present disclosure;

[0020] Figure 9 A schematic diagram illustrating the structure of an intelligent agent of artificial intelligence according to embodiments of the present disclosure; and

[0021] Figure 10 A schematic block diagram of an example electronic device for implementing embodiments of the present disclosure based on agent-based collaborative interaction methods is shown. Detailed Implementation

[0022] The exemplary embodiments of this disclosure are described below with reference to the accompanying drawings, including various details of the embodiments to aid understanding, and should be considered merely exemplary. Therefore, those skilled in the art will recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of this disclosure. Similarly, for clarity and brevity, descriptions of well-known functions and structures are omitted in the following description.

[0023] In the technical solution disclosed herein, the acquisition, storage, and application of user personal information comply with the provisions of relevant laws and regulations, necessary confidentiality measures have been taken, and there is no violation of public order and good morals.

[0024] The inventors discovered that with the rapid development of artificial intelligence technology, relevant internet platforms can process user input information based on large models with generative capabilities and meet user needs by outputting generative content such as language text and videos. However, users' actual needs are usually highly personalized and diverse, making it difficult for generative models to accurately generate responses that meet actual needs based on user input information, resulting in low response accuracy.

[0025] Embodiments of this disclosure provide an interaction method, apparatus, agent, and storage medium based on agent collaboration. The interaction method based on agent collaboration includes: performing intent understanding on initial input information for a target object to obtain an initial intent understanding result; using a first agent to perform interaction mode requirement detection on the initial intent understanding result to obtain interaction mode requirement attributes; invoking an interaction element matching the interaction mode requirement attributes to interact with the target object, obtaining interaction information representing the target object's requirement intent, the interaction element being used to respond to the interaction behavior; controlling a second agent to execute a target task based on the interaction information, obtaining an execution result, and pushing the execution result to the target object.

[0026] According to embodiments of this disclosure, by performing intent understanding on the initial input information, and then using a first intelligent agent to detect and obtain interaction mode requirement attributes that can meet the actual needs of the target object based on the initial intent understanding results, the interaction mode matching the target object's needs can be determined more accurately through multi-level intent understanding. Therefore, by calling matching interaction elements to interact with the target object based on the interaction mode requirement attributes, the target object can perform interactive behaviors on interaction elements adapted to the required scenario. The input of interaction information can accurately represent the actual need intent, thereby controlling a second intelligent agent to execute the target task based on the interaction information to obtain the execution result that meets the need intent. By pushing the execution result to the target object, the interaction efficiency is improved, the accuracy of the response content is enhanced, and the personalized needs of the target object are met.

[0027] Figure 1 The illustration schematically depicts an exemplary system architecture for applying agent-based collaborative interaction methods and apparatus according to embodiments of the present disclosure.

[0028] It is important to note that Figure 1 The examples shown are merely examples of system architectures that can be applied to embodiments of this disclosure, to help those skilled in the art understand the technical content of this disclosure, but do not imply that embodiments of this disclosure cannot be used in other devices, systems, environments, or scenarios. For example, in another embodiment, an exemplary system architecture for applying the agent-based collaborative interaction method and apparatus may include a terminal device, but the terminal device may implement the agent-based collaborative interaction method and apparatus provided by embodiments of this disclosure without interacting with a server.

[0029] like Figure 1 As shown, the system architecture 100 according to this embodiment may include terminal devices 101, 102, and 103, a network 104, and a server 105. The network 104 serves as a medium for providing a communication link between the terminal devices 101, 102, and 103 and the server 105. The network 104 may include various connection types, such as wired and / or wireless communication links, etc.

[0030] Users can use terminal devices 101, 102, and 103 to interact with server 105 via network 104 to receive or send messages, etc. Various communication client applications can be installed on terminal devices 101, 102, and 103, such as knowledge reading applications, web browser applications, search applications, instant messaging tools, email clients, and / or social platform software, etc. (for example only).

[0031] Terminal devices 101, 102, and 103 can be various electronic devices with displays and web browsing capabilities, including but not limited to smartphones, tablets, laptops, and desktop computers.

[0032] Server 105 can be a server that provides various services, such as a backend management server that supports the content browsed by users using terminal devices 101, 102, and 103 (for example only). The backend management server can analyze and process data such as received user requests, and feed back the processing results (such as web pages, information, or data obtained or generated according to user requests) to the terminal devices.

[0033] Server 105 can be a cloud server, also known as a cloud computing server or cloud host. It is a host product in the cloud computing service system, which solves the shortcomings of traditional physical hosts and VPS services ("Virtual Private Server", or "VPS" for short) in terms of high management difficulty and weak business scalability. Server 105 can also be a server for a distributed system or a server combined with blockchain.

[0034] Server 105 can obtain a first agent or a second agent by deploying a generative model. It should be noted that the first agent or the second agent can be deployed on the same server, or they can be deployed on different servers.

[0035] It should be noted that the agent-based interaction method provided in this disclosure can generally be executed by terminal devices 101, 102, or 103. Accordingly, the agent-based interaction device provided in this disclosure can also be disposed in terminal devices 101, 102, or 103.

[0036] Alternatively, the agent-based interaction method provided in this embodiment can generally be executed by server 105. Correspondingly, the agent-based interaction device provided in this embodiment can generally be located in server 105. The agent-based interaction method provided in this embodiment can also be executed by a server or server cluster that is different from server 105 and capable of communicating with terminal devices 101, 102, 103 and / or server 105. Correspondingly, the agent-based interaction device provided in this embodiment can also be located in a server or server cluster that is different from server 105 and capable of communicating with terminal devices 101, 102, 103 and / or server 105.

[0037] It should be understood that Figure 1The number of terminal devices, networks, and servers shown is merely illustrative. Depending on implementation needs, any number of terminal devices, networks, and servers can be included.

[0038] Figure 2 A flowchart illustrating an agent-based collaborative interaction method according to an embodiment of the present disclosure is shown.

[0039] like Figure 2 As shown, the interaction method based on agent collaboration includes operations S210~S240.

[0040] In operation S210, the initial input information for the target object is subjected to intent understanding to obtain the initial intent understanding result.

[0041] In operation S220, the first intelligent agent performs interaction mode requirement detection on the initial intent understanding result to obtain the interaction mode requirement attribute.

[0042] In operation S230, the interactive element that matches the interaction mode requirement attribute is invoked to interact with the target object and obtain interactive information that represents the target object's requirement intention.

[0043] In operation S240, the second intelligent agent is controlled to execute the target task based on interactive information, the execution result is obtained, and the execution result is pushed to the target object.

[0044] According to embodiments of this disclosure, the first or second intelligent agent can be constructed based on a large model with generative capabilities, such as a generative model. The first or second intelligent agent can be deployed on a server, or it can be deployed on multiple servers connected in communication. The intelligent agents involved in the embodiments of this disclosure can perform data analysis based on the model algorithms or model parameters of the large model to obtain processing results.

[0045] In some embodiments, Figure 2 In the interaction method of the illustrated embodiment, operations S210~S240 can be executed by a first intelligent agent as the execution subject. This can be understood as one or more servers that deploy the first intelligent agent to execute the interaction method provided in the embodiments of this disclosure.

[0046] In some embodiments, Figure 2 In the interaction method of the embodiment shown, operation S220 is executed by the first intelligent agent, while operations S210, S230 and S240 can be executed by other servers or computing devices deployed other than the first intelligent agent.

[0047] It should be noted that, for the purpose of explaining the interaction method based on agent collaboration provided in the embodiments of this disclosure, the embodiments of this disclosure use a first agent as the execution subject of the interaction method, and are not intended to limit the interaction method to only be executed by the first agent.

[0048] According to embodiments of this disclosure, the initial input information of the target object can be data of any modality, such as voice data, text data, or image data. Embodiments of this disclosure do not limit the specific data modality type of the initial input information.

[0049] According to embodiments of this disclosure, initial input information can be processed by calling a large language model to obtain an initial intent understanding result. Alternatively, the initial input information can be subjected to intent understanding through other methods such as keyword extraction to obtain keywords representing the initial intent understanding result. Embodiments of this disclosure do not limit the specific method for obtaining the initial intent understanding result.

[0050] According to embodiments of this disclosure, the initial intent understanding result can be the initial-level demand intent represented by the initial input information. For example, the initial intent understanding result can represent the topic intent of the target object's demand intent, such as a travel planning topic, a health consultation topic, etc. Furthermore, the initial intent understanding result may also include associated demand intents related to the topic intent.

[0051] For example, if the initial input information is "planning a travel itinerary for city A", the initial intent understanding results can include thematic intents related to the theme of travel planning, destination intents related to the destination "city A", and "planning a travel itinerary" representing the task requirement intent. The task requirement intent and the destination intent can be related requirement intents.

[0052] According to embodiments of this disclosure, interactive elements are used to respond to interactive behaviors. For example, a dialog box interactive element can respond to user text input behavior, determining the text input information as interactive information. By utilizing a first intelligent agent to perform interactive pattern requirement detection on the initial intent understanding result, the interactive mode or interactive element represented by the interactive pattern requirement attribute can be matched with the topic intent and associated requirement intent represented by the initial intent understanding result, so that the target object can naturally express its actual requirement information by performing interactive behaviors on the interactive element. Thus, by calling the interactive element that matches the interactive model requirement attribute to interact with the target object, and by responding to the target object's interactive behavior through the interactive element, interactive information that can represent the target object's deep requirements can be obtained, realizing the target object's requirement intent mining, and enabling the interactive information to more accurately represent the target object's multi-dimensional requirement attributes.

[0053] It should be noted that invoking interactive elements to interact with a target object can include displaying interactive elements in a display interface to provide interactive behavior objects for the target object, and obtaining interactive information by responding to the target object's text input operations, click operations, voice input operations, and other interactive behaviors through the interactive elements. Interactive elements can be components capable of responding to interactive behaviors performed by the target object, and can be displayed in a display interface, for example, based on element types such as icons, dialog boxes, and page links. The embodiments of this disclosure do not limit the specific display method of interactive elements.

[0054] According to embodiments of this disclosure, controlling a second intelligent agent to execute a target task based on interactive information can include using the second intelligent agent to determine a target task that meets the actual needs and intentions of a target object based on the interactive information, and then executing the target task to obtain an execution result that satisfies the needs of the target object. This execution result could be, for example, a travel guide document, including images, text, and a travel planning page linked to attraction ticket bookings. By using a first intelligent agent to invoke interactive elements that match the interactive pattern and needs attributes of the target object to interact with the target object and obtain interactive information, and by controlling the second intelligent agent to execute a target task that matches the needs and intentions represented by the interactive information, a multi-agent collaborative approach can be used to achieve needs and intention mining and task execution, thereby improving the matching degree between the execution result and the actual needs and intentions of the target object, and ultimately increasing the satisfaction of the response content.

[0055] In some embodiments, the target task includes at least one of the following: travel strategy planning task, copywriting generation task, graphic element generation task, and image editing task.

[0056] According to embodiments of this disclosure, the copywriting generation task can be, for example, using a large language model to perform semantic understanding of interactive information to generate copywriting of any type, such as promotional copywriting for a topic or tourism promotion.

[0057] According to embodiments of this disclosure, a graph element generation task can be used to generate image elements such as icons and cartoon characters displayed in videos or images. The graph element generation task may include a specified code generation task, such as text content described based on interactive information, such as "generate a red sphere and make the red sphere move randomly in the image." A second intelligent agent is controlled to understand the text content using a large language model and generate a specified code script, and the specified code script is executed to control the random movement of the red sphere in the image.

[0058] In some embodiments, image editing tasks may include editing operations such as cropping, compressing, and adjusting the tone of a specified image. By controlling a second intelligent agent to invoke a specified image editing tool to execute the image editing task, precise image editing can be achieved according to the intended needs represented by the interactive information, thereby improving the accuracy of the image editing result as the execution result.

[0059] In some embodiments, the travel strategy planning task may include multiple sub-tasks such as attraction location query, travel route planning, and travel trajectory generation. By controlling a second intelligent agent to execute the travel strategy planning task based on interactive information, the second intelligent agent can execute different sub-tasks by calling various functional tools. This ensures that the travel strategy information determined based on the execution results of multiple sub-tasks more accurately meets the user's actual needs and improves the user's travel experience.

[0060] It should be noted that the number of target tasks involved in this embodiment may be one or more. The second intelligent agent can be controlled to execute multiple target tasks to obtain multiple intermediate execution results, and an execution result for pushing to the target object can be generated based on the multiple intermediate execution results.

[0061] For example, an image editing task for image editing includes multiple target tasks, namely, a specified image query task, an object detection task, and an image cropping task. The image query task can be used to search for multiple images in a specified directory; the object detection task is used to identify images representing "kittens" from multiple images and to determine the image regions representing "kittens" within those images; and the image cropping task can crop the image regions representing "kittens" to obtain diverse image blocks representing "kittens."

[0062] According to embodiments of this disclosure, performing intent understanding on initial input information for a target object to obtain an initial intent understanding result may include: calling a large language model to perform structured semantic understanding on the initial input information to obtain multiple intent understanding information with structured attributes.

[0063] In some embodiments, multiple intent understanding information of a structured attribute can be information content corresponding to the intent attribute type.

[0064] For example, if the intent attribute type is "topic intent type", the information content corresponding to the topic intent type could be "tourism". Another example is if the intent attribute type is "destination" (related to the demand intent type), and the information content corresponding to "city A" (related to the demand intent type "destination").

[0065] In some embodiments, a first agent can be used to invoke a large language model to perform semantic understanding of the initial input information based on prompt words, thereby obtaining multiple intent understanding information with structured attributes. The prompt words are used to control the large language model to perform semantic understanding of the initial input information based on multiple intent attribute types represented by structured attributes, so that the multiple intent understanding information can correspond to multiple intent attribute types respectively. This allows the multiple intent understanding information with structured attributes to clearly represent the target object's topic intent and related need intent according to intent attribute types. This enables the first agent to detect the interaction scenario or interaction pattern of the target object's needs through the multiple intent understanding information with structured attributes, and accurately determine the interaction pattern need attributes.

[0066] According to embodiments of this disclosure, the intent understanding information includes an intent attribute type and intent fields corresponding to the intent attribute type. The intent fields may represent field data such as text fields or identifier fields corresponding to the intent attribute type.

[0067] In some embodiments, the large language model can semantically understand multiple intent understanding information with structured attributes and enable the intent understanding information to be represented based on the control of prompt words. However, the intent field in the initial input information corresponding to the preset demand attribute type indicated by the prompt words may have defects in demand intent representation, such as semantic ambiguity and missing text. Therefore, the first intelligent agent can determine the missing information content in the interaction mode demand attributes by processing the demand intent representation defects in the intent understanding information, enabling the target object to provide the preset demand attribute type through interaction. This improves the accuracy and completeness of the interaction information's representation of the target object's actual demand intent, thereby ensuring that the execution result of the target task meets the user's actual needs.

[0068] For example, if the initial input information is "travel planning," and the multiple intent understanding information only contains the topic intent "travel," the intent field corresponding to the associated demand intent type "destination" is missing information. Therefore, the first agent, under the control of prompt words, can perform interaction pattern demand detection based on multiple intent understanding information. This determines that the interaction pattern demand attributes represent multiple interactive elements for the travel planning scenario, including destination interaction option elements representing "City A" and "City B." By having the target object interact with these interaction option elements, the intent field corresponding to the target object's destination demand can be determined, thereby improving the execution result of the travel strategy planning task to meet the target object's actual travel needs.

[0069] In some embodiments, there is a hierarchical relationship between multiple intent understanding information. For example, a topic intent type and a topic intent field corresponding to the topic intent type have a first level, and an associated intent type and an intent field corresponding to the associated intent type may have a second level.

[0070] In some embodiments, the multiple intent fields corresponding to the associated intent type can also have different structural levels. For example, the intent fields corresponding to the travel duration intent type include "summer vacation" and "one week". The intent field "summer vacation" has a second level, and the intent field "one week" has a third level.

[0071] According to embodiments of this disclosure, by representing the structured attributes of multiple intent understanding information through a hierarchical relationship, the first agent can more accurately and deeply understand the clarity of the target object's current needs regarding the topic intent according to the hierarchical relationship. It can also call matching interactive elements to interact with the target object by outputting the interaction mode requirement attributes, enabling the target object to interact according to an interaction mode adapted to the clarity of its needs. This allows multiple interactive information to progressively represent the target object's actual needs according to the hierarchical relationship, thereby achieving the mining of the target object's intent needs. This ensures that subsequent execution results satisfy the potential intent needs expressed by the target object through the interactive information, and also prevents the second agent from executing tasks by processing information with unclear or missing intent needs, which could affect the accuracy of the response content. This improves the matching degree between the execution results and the target object's actual needs.

[0072] According to embodiments of this disclosure, the interaction method based on agent collaboration further includes: determining multiple types of interaction elements based on interaction mode requirement attributes, and the invocation order of the multiple types of interaction elements.

[0073] For example, the interaction mode requirement attribute represents a scenario mode for interacting with the intent field of the theme intent type "tourism," and indicates the absence of intent fields such as "destination," "tour duration," and "number of tourists." Multiple interactive elements corresponding to the interaction mode requirement attribute can include prompt icons to indicate intent attribute types such as "destination," "tour duration," and "number of tourists," and multiple prompt icons are displayed in the order of "destination," "tour duration," and "number of tourists" to guide the target audience to input interactive information through prompt icons according to the target audience's thinking style. This improves the accuracy and sufficiency of the interactive information in expressing the target audience's actual needs and intents, thereby improving the matching degree between the target task and execution result and the target audience's actual needs and intents, and enhancing the target audience's interactive experience and efficiency.

[0074] In some embodiments, the interaction mode requirement attribute also characterizes the degree of semantic accuracy requirement of the interaction information, which needs to meet a preset accuracy condition. For example, in interaction scenarios corresponding to thematic intents such as health consultation or precision instrument repair consultation, it is necessary to use a specific interaction mode to control the interaction information input by the target object to meet the preset accuracy condition, so as to avoid errors in the creation and execution of the target task due to the semantic accuracy of the interaction information not meeting the requirements. Therefore, matching interaction elements can be determined based on the interaction mode requirement attribute that characterizes the degree of semantic accuracy requirement of the interaction information to meet the preset accuracy condition, so that the interaction information meets the preset accuracy condition.

[0075] In some embodiments, invoking an interactive element that matches the interaction mode requirement attributes to interact with the target object may include: invoking an interactive option element that represents preset option data to interact with the target object.

[0076] According to embodiments of this disclosure, an interactive option element is used to determine target option data in preset option data in response to interactive behavior, and the interactive information may include the target option data.

[0077] For example, multiple interactive option elements can represent preset option data corresponding to specific drug names, such as "Drug A," "Drug B," and "Drug C." The target user can determine that the target option data is "Drug A" by interacting with the interactive option element representing "Drug A." This avoids deviations in the health consultation task performed by the second intelligent agent due to unclear or incorrect drug names input by the target user, ensuring that the accuracy of the health consultation results meets the actual needs of the target user.

[0078] Figure 3 The diagram illustrates an interaction scenario of an agent-based collaborative interaction method according to an embodiment of the present disclosure.

[0079] like Figure 3As shown, this application scenario includes a first interactive interface 300, which displays the initial input information of the target object: "Are there any ways to relieve frequent headaches?". The first intelligent agent performs intent understanding on the initial input information by calling a large language model. The resulting initial understanding includes multiple intent understanding information with structured attributes. The first intelligent agent performs interaction pattern requirement detection on the multiple intent understanding information. The resulting interaction pattern requirement attributes represent the semantic clarity of the interaction information related to the health consultation topic intent and need to meet preset accuracy conditions. Therefore, multiple interactive option elements can be determined as interactive elements matching the interaction pattern requirement attributes. These multiple interactive option elements can be multiple option icons displayed in the target option box 310. The text content in the option icons represents preset option data. For example, the target option box 310 may include option icons representing the preset option data "Drug A" and "Drug B" respectively.

[0080] The target object can interact with the option icons in the target option box 310. The option icons, responding to the interaction, determine the target option data as follows: "Duration of attack: 10-20 minutes," "Frequency of headache attacks: 2-3 times per week," "Taking the following medication: Medication B," and "Location of pain: Top of the head." Based on this, the second agent can perform the target task using the interaction information including the target option data, resulting in the following execution result: "Based on the information you entered, the initial assessment is that the headache is caused by overwork. It is recommended to rest early and exercise more each day, and to visit a relevant specialist hospital for examination and consultation."

[0081] According to embodiments of this disclosure, the interaction mode demand attribute can also characterize whether the emotional demand attribute of the initial intent understanding result meets preset emotional demand conditions. For example, for specified theme demand intentions such as tourism theme demand intentions or problem-solving theme demand intentions, preset emotional demand conditions can be met.

[0082] In some embodiments, the first intelligent agent can use its initial intent understanding results and historical interaction preferences of the target object to perform semantic understanding in order to determine whether the target object needs to interact with the target object in an interaction mode that meets the preset emotional needs conditions.

[0083] For example, the first intelligent agent determines that the target object frequently inputs demand information in interactive scenarios such as videos and games by understanding historical interaction preferences. Thus, the emotional demand attributes of the initial intention understanding result can be represented by the output interaction mode demand attributes to meet the preset emotional demand conditions.

[0084] In some embodiments, invoking an interactive element that matches the interaction mode requirement attributes to interact with the target object may also include: invoking a virtual object interactive element to interact with the target object.

[0085] According to embodiments of this disclosure, virtual object interaction elements are used to drive virtual objects to perform verbal communication with target objects. For example, an authorized cartoon character can be driven to explain a tourist destination, prompting the target object to interact more engagingly with the video elements delivered by the cartoon character. By fully understanding the target object's needs and intentions, the target object can be guided to input interactive information reflecting its actual needs and intentions. This allows the target object to satisfy its emotional needs through the immersive interaction provided by the virtual object interaction elements, and to input interactive information with high efficiency. This, in turn, improves the accuracy and comprehensiveness of the second agent's understanding of the interactive information, enhances the accuracy of the execution results, and improves the target object's interactive experience.

[0086] Figure 4 The diagram illustrates an interaction scenario of an agent-based collaborative interaction method according to another embodiment of the present disclosure.

[0087] like Figure 4 As shown, this application scenario includes a second interactive interface 400, which displays a virtual object 420 corresponding to the emotional needs attribute represented by the interaction mode's requirement attributes. The virtual object 420 can perform corresponding yoga movements during a spoken explanation of a yoga course, allowing the target audience to interact with the first guidance element 411 and the second guidance element 412 indicated by the virtual object 420 in an immersive interactive scenario to input the course item they want to experience or the course time they need to book. Furthermore, the target audience can also input voice or text information into the input box 431 used by the virtual object 420 to provide interactive information. It should be understood that the virtual object interactive elements can include the virtual object 420, the first guidance element 411, and the second guidance element 412. The input box 431 can be introduced as a dialogue interactive element into the current interaction mode to provide rich interactive input entry points.

[0088] The second intelligent agent can execute the target task based on the interaction information determined by the target object's interaction behavior with virtual object interaction elements, and generate a course experience list as the execution result for the target object.

[0089] According to embodiments of this disclosure, the interaction mode requirement attribute also characterizes the degree of semantic ambiguity of the initial intent understanding result, which satisfies a preset ambiguity condition.

[0090] For example, if the number of missing intent fields corresponding to multiple related demand intents in the initial intent understanding result regarding the "tourism" theme meets the preset number threshold, it can be determined that the initial intent understanding result meets the preset fuzziness condition.

[0091] In some embodiments, the first agent can be controlled to perform semantic understanding of the initial input information based on preset prompts that represent preset ambiguity conditions, so that the output initial intent understanding result includes "the initial input information is semantically ambiguous".

[0092] For example, if the initial input is "solve the problem," the preset prompts could include "need to understand the semantics of the problem, determine the subject and knowledge points of the problem." Then, the initial semantic understanding result could indicate "the semantics of the problem-solving requirement are ambiguous, requiring further inquiry."

[0093] In some embodiments, invoking an interactive element that matches the interaction mode requirement attributes to interact with the target object may also include: invoking a dialog interactive element to interact with the target object.

[0094] According to embodiments of this disclosure, dialogue interaction elements are used to guide the target object's needs and intentions. For example, dialogue interaction elements may include question information generated by a first agent based on the need and intention type indicated by the interaction mode's need attribute. By displaying the question information, the target object is prompted to perform interactive behaviors such as text input or voice input according to the text information, thereby obtaining information content corresponding to the need and intention type, thus improving the completeness of the interactive information in representing the actual need and intention.

[0095] For example, a dialogue interaction element could include a question information display box such as "What is your question?". If the target user inputs the interaction information "The question is 2+2×3=", a question information display box corresponding to the next dialogue interaction element could be displayed: "Would you like a detailed explanation?". Thus, dialogue interaction elements can be used to probe the target user's intent and guide them to input interactive information expressing their intent through interactive behavior.

[0096] Figure 5 The diagram illustrates an interaction scenario of an agent-based collaborative interaction method according to yet another embodiment of the present disclosure.

[0097] like Figure 5As shown, this application scenario includes a third interactive interface 500. The third interactive interface 500 can include the initial input information from the user, "Need to plan a trip." The initial intent understanding result can only include the main intent "travel," with semantic ambiguity meeting a preset ambiguity condition. Therefore, the third interactive interface 500 can display the first dialogue interaction element 511 to follow up on the target object's related needs intent, guiding the target object to input the interaction information "City B, three-day trip" for the first dialogue interaction element 511 through dialog box 501. The first intelligent agent can also continue to call the second dialogue interaction element 512 "How many people are traveling? What type of attractions do you like?" based on the current interaction information and the initial intent understanding result to continue guiding the target object's needs intent, enabling the target object to input subsequent interaction information through interactive operations. Thus, by controlling the second intelligent agent to execute the travel strategy planning task based on the user's input interaction information, the travel planning scheme is displayed as the execution result on the third interactive interface 500, realizing the push of the execution result to the target object.

[0098] According to embodiments of this disclosure, controlling a second intelligent agent to perform a target task based on interactive information may include: using a first intelligent agent to perform semantic understanding of the interactive information to obtain task requirement information; controlling the second intelligent agent to perform task orchestration based on the task requirements to obtain the target task; and controlling the second intelligent agent to perform the target task.

[0099] In some embodiments, a first intelligent agent can perform semantic understanding of the interaction information, the initial intent understanding result, and the initial input information to determine task requirement information. Task requirement information may include, for example, data of any modality, such as text or images, describing task elements like execution conditions and parameters of the target task. This transforms the requirement intent represented by the interaction information and initial input information into relevant information corresponding to the target task, thereby improving the efficiency of the second intelligent agent in creating and executing the target task and enhancing the accuracy of the execution results.

[0100] According to embodiments of this disclosure, the second intelligent agent may possess task orchestration capabilities. The second intelligent agent can determine one or more target tasks that satisfy the task requirement information through semantic understanding of the task requirement information. Thus, by leveraging the intent understanding and interaction pattern matching capabilities of the first intelligent agent to dynamically identify the target object's intent, and combining the task orchestration and task execution capabilities of the second intelligent agent, multi-agent collaboration can be achieved to determine execution results that satisfy the diverse intents of the target object, thereby improving interaction satisfaction.

[0101] In some embodiments, controlling the second intelligent agent to orchestrate tasks based on task requirements to obtain target tasks includes: controlling the planning intelligent agent to orchestrate tasks based on task requirement information to obtain multiple target tasks and dependencies between the multiple target tasks. Dependencies can represent the execution order of the multiple target tasks.

[0102] According to embodiments of this disclosure, the second intelligent agent includes a planning intelligent agent, which issues multiple target tasks and dependencies between the multiple target tasks to the execution intelligent agent, and the execution intelligent agent executes the multiple target tasks according to the dependencies to obtain the execution result.

[0103] For example, a planning agent can orchestrate tasks based on task requirements related to the intent of a presentation video, resulting in multiple target tasks such as problem-solving, text generation, subtitle generation, and script generation. The execution order of these target tasks is then defined as a dependency. By invoking the execution agent to execute these target tasks according to their dependencies, a presentation video can be generated as the execution result. The text generation task is then executed based on the problem-solving result, yielding the presentation text as an intermediate execution result. The script generation task is then executed based on the presentation text, producing voice-driven and action-driven data to drive the virtual object's narration, serving as an intermediate execution result. The subtitle generation task is then executed based on the voice-driven data, determining the subtitle file as the intermediate specified result. Finally, by synthesizing these multiple intermediate execution results, a presentation video can be generated.

[0104] According to embodiments of this disclosure, the control planning agent may perform task orchestration based on task requirement information, which may include: the control planning agent orchestrating tasks based on task requirement information and functional description information for candidate execution agents, determining multiple target tasks, and the execution agents among the candidate execution agents for performing the target tasks.

[0105] According to embodiments of this disclosure, dependencies can also characterize the invocation order of multiple execution agents. The functional description information for candidate execution agents can represent information describing the task execution capability boundaries of the candidate execution agents, such as the task types they can execute, task parameters, and execution result accuracy. By controlling the planning agent to process task requirement information and the functional description information for candidate execution agents for task orchestration, the functional description information can be used as a prompt to control the planning agent to fully understand the capabilities of the candidate execution agents and assign multiple target tasks that meet the task requirement information to execution agents whose execution capability boundaries meet the task requirement conditions of the target tasks. This improves the accuracy of task orchestration, avoids mismatches between the execution capability boundaries of the execution agents and the requirement conditions of the target tasks, which could lead to errors in intermediate execution results, thereby improving the accuracy of execution results and ultimately enhancing the interaction satisfaction of the target object.

[0106] According to embodiments of this disclosure, a planning agent orchestrates tasks based on functional description information and task requirement information obtained through dynamic multi-level intent understanding of the target object. The planning agent, through task decomposition and phase division, comprehensively evaluates the execution capability boundaries, task adaptability, and current resource status of each candidate execution agent, and selects the optimal combination of execution agents to collaboratively execute multiple target tasks. Furthermore, the planning agent outputs dependencies to clearly represent the execution order of target tasks and the interactive collaboration process between multiple execution agents, ensuring smooth information transmission between them, avoiding execution conflicts and redundant use of computing or storage resources, and effectively improving task completion efficiency and quality.

[0107] Figure 6 The illustration shows a schematic diagram of the principle of agent-based collaboration according to an embodiment of the present disclosure.

[0108] like Figure 6 As shown, the target object 601 interacts with the first intelligent agent 610 by operating a computer device. The first intelligent agent 610 calls a large language model to perform semantic understanding on the initial input information input by the target object 601, and obtains the initial intent understanding result. The initial intent understanding result includes multiple intent understanding information with structured attributes. For example, the multiple intent understanding information can be "topic intent type: tourism; destination: city A; time: during summer vacation".

[0109] The first intelligent agent 610 performs interaction pattern requirement detection on the initial intent understanding result, and finds that the interaction pattern requirement attribute meets the emotional requirement condition. It then invokes virtual object interaction elements to interact with the target object 601 to determine the interaction information. The first intelligent agent 610 generates task requirement information by processing the interaction information and the initial intent understanding result, and sends the task requirement information to the second intelligent agent 620.

[0110] The second intelligent agent terminal 620 includes multiple second intelligent agents, namely a planning intelligent agent 6211 and multiple candidate execution intelligent agents, namely the first candidate execution intelligent agent 6221, the second candidate execution intelligent agent 6222, ..., the nth candidate execution intelligent agent 622n. The planning intelligent agent 6211 orchestrates tasks based on task requirement information, obtaining two target tasks: a scenic spot ticket reservation task and a travel route planning task. The planning intelligent agent 6211 calls the first candidate execution intelligent agent 6221 and the second candidate execution intelligent agent 6222 as two execution intelligent agents to execute the scenic spot ticket reservation task and the travel route planning task respectively, thus obtaining two intermediate execution results: a ticket reservation QR code and a travel route map. The first intelligent agent 610 synthesizes the ticket reservation QR code and the travel route map to obtain the target page, which is then pushed to the target object 601 as the execution result.

[0111] According to embodiments of this disclosure, the interaction method based on agent collaboration may further include: responding to an execution feedback request for an execution result, controlling a second agent to update a target task based on feedback information carried in the execution feedback request, and executing the updated target task to obtain a feedback execution result.

[0112] In some embodiments, the execution feedback request may be determined by the interactive operation performed by the target object in response to the execution result. For example, the feedback information carried by the execution feedback request may include "modify the itinerary and replace attraction A with attraction B".

[0113] In some embodiments, updated task requirement information can be generated by utilizing a first intelligent agent to process feedback and interaction information. The first intelligent agent then controls a second intelligent agent to orchestrate tasks based on the updated task requirement information to obtain an updated target task. The second intelligent agent can then execute the updated target task to obtain a new execution result as feedback.

[0114] For example, the feedback implementation result could be a travel strategy planning document that replaces attraction A with attraction B.

[0115] In some embodiments, controlling the second agent to update the target task based on the feedback information carried in the execution feedback request may include: using the first agent to perform semantic understanding of the execution result and the target task based on the feedback information to obtain feedback task requirement information; and controlling the second agent to perform task orchestration based on the feedback task requirement information to obtain the updated target task.

[0116] In one embodiment, a first agent can be used to process execution results, target tasks, and interaction information based on feedback information as prompts, outputting updated task requirement information as feedback task requirement information. This allows the task requirement information to be updated based on the feedback prompts regarding the execution results and target tasks, enabling the feedback task requirement information to more accurately represent the current needs and intentions of the target object. Furthermore, the updated target task can be executed by controlling the execution agent to obtain a feedback execution result that satisfies the current needs of the target object.

[0117] Figure 7 The illustration shows a schematic diagram of the principle of agent-based collaboration according to another embodiment of the present disclosure.

[0118] like Figure 7 As shown, the first target object 701 interacts with the first intelligent agent 711 to input initial input information and interaction information. The first intelligent agent 711 determines the execution result by executing the method provided in this embodiment and pushes the execution result to the first target object 701. The first target object 701 controls the first intelligent agent 711 to generate a feedback execution result based on the adjustment information content of the feedback information by transmitting feedback information to the first intelligent agent 711, and pushes the feedback execution result to the first target object 701. The first intelligent agent 711 can also store the feedback information and the feedback execution result in a knowledge base to update the task execution strategy. For example, it can fine-tune the parameters of the large model used by the first intelligent agent 711 to obtain an updated first intelligent agent 712.

[0119] The updated first agent 712 can interact with other target objects 702 based on its optimized capabilities to push optimized execution results to them.

[0120] According to embodiments of this disclosure, interactive elements such as graphical interfaces and natural language interaction windows can be used to respond to the interactive behavior of the target object and determine feedback information. This allows a first and second intelligent agent to collaborate in processing the feedback information and executing updated target tasks, enabling timely adjustments or modifications to the execution results based on the target object's desired outcome. Furthermore, the feedback information can be stored in a knowledge base for the target object, allowing the first and second intelligent agents to learn from the feedback information and update model parameters and task execution strategies. This enables the agent to automatically backtrack and adjust its capabilities, thereby continuously optimizing the interaction process with the target object.

[0121] Figure 8 A block diagram of an agent-based interactive device according to an embodiment of the present disclosure is shown schematically.

[0122] like Figure 8 As shown, the interactive device 800 based on intelligent agent collaboration includes: an intent understanding module 810, a detection module 820, a calling module 830, and an execution module 840.

[0123] The intent understanding module 810 is used to understand the initial input information for the target object and obtain the initial intent understanding result.

[0124] The detection module 820 is used to detect interaction mode requirements based on the initial intent understanding results of the first intelligent agent, and obtain the interaction mode requirement attributes.

[0125] Module 830 is invoked to call interactive elements that match the interaction mode requirements attributes to interact with the target object and obtain interactive information that represents the target object's requirements and intentions. The interactive elements are used to respond to interactive behaviors.

[0126] The execution module 840 is used to control the second intelligent agent to execute the target task based on interactive information, obtain the execution result, and push the execution result to the target object.

[0127] According to embodiments of this disclosure, the interaction mode requirement attribute characterizes the degree of semantic accuracy requirement of the interaction information, indicating that it meets a preset accuracy condition. The calling module includes a first calling unit.

[0128] The first calling unit is used to call the interactive option element representing the preset option data to interact with the target object. The interactive option element is used to respond to the interactive behavior to determine the target option data in the preset option data. The interactive information includes the target option data.

[0129] According to embodiments of this disclosure, the interaction mode requirement attribute further characterizes whether the semantic ambiguity of the initial intent understanding result meets a preset ambiguity condition. The invocation module further includes a second invocation unit.

[0130] The second invocation unit is used to invoke dialogue interaction elements to interact with the target object. The dialogue interaction elements are used to guide the target object's needs and intentions.

[0131] According to embodiments of this disclosure, the interaction mode requirement attribute also characterizes the emotional requirement attribute of the initial intent understanding result as satisfying a preset emotional requirement condition; wherein, the calling module further includes a third calling unit.

[0132] The third calling unit is used to call virtual object interaction elements to interact with the target object. The virtual object interaction elements are used to interact with the target object by driving the virtual object to perform verbal communication.

[0133] According to embodiments of this disclosure, the interaction device based on agent collaboration further includes a first determining module.

[0134] The first determining module is used to determine multiple types of interactive elements and the order in which these interactive elements are invoked based on the interaction mode requirement attributes.

[0135] According to embodiments of this disclosure, the intent understanding module includes a first obtaining unit.

[0136] The first acquisition unit is used to call the large language model to perform structured semantic understanding on the initial input information and obtain multiple intent understanding information with structured attributes.

[0137] According to embodiments of this disclosure, multiple intent understanding information have a hierarchical relationship, and the intent understanding information includes intent attribute types and intent fields corresponding to the intent attribute types.

[0138] According to embodiments of this disclosure, the execution module includes: a task requirement information acquisition unit and a control unit.

[0139] The task requirement information acquisition unit is used to perform semantic understanding of the interaction information by the first intelligent agent to obtain task requirement information.

[0140] The control unit is used to control the second intelligent agent to arrange tasks based on task requirement information, obtain the target task, and control the second intelligent agent to execute the target task.

[0141] According to embodiments of this disclosure, the control unit includes a control subunit.

[0142] The control subunit is used to control the planning agent to arrange tasks based on task requirement information, thereby obtaining multiple target tasks and the dependencies between the multiple target tasks; wherein, the second agent includes the planning agent, which issues multiple target tasks and the dependencies between the multiple target tasks to the execution agent, and the execution agent executes multiple target tasks according to the dependencies to obtain the execution result.

[0143] According to an embodiment of this disclosure, the control subunit is further configured as follows: the control planning agent performs task orchestration on task requirement information and functional description information for candidate execution agents, determines multiple target tasks, and the execution agents among the candidate execution agents for executing the target tasks, wherein the dependency relationship represents the calling order of the multiple execution agents.

[0144] According to embodiments of this disclosure, the interaction device based on agent collaboration further includes an update module.

[0145] The update module is used to respond to the execution feedback request for the execution result, control the second agent to update the target task based on the feedback information carried in the execution feedback request, execute the updated target task, and obtain the feedback execution result.

[0146] According to embodiments of this disclosure, the update module includes: a semantic understanding unit and a second acquisition unit.

[0147] The semantic understanding unit is used to perform semantic understanding of the execution results and target tasks based on feedback information by the first intelligent agent, and obtain feedback task requirement information.

[0148] The second acquisition unit is used to control the second intelligent agent to perform task arrangement based on the feedback task requirement information, and obtain the updated target task.

[0149] According to embodiments of this disclosure, the target task includes at least one of the following: travel strategy planning task, copywriting generation task, graphic element generation task, and image editing task.

[0150] Figure 9 A schematic block diagram of an artificial intelligence agent according to an embodiment of the present disclosure is shown.

[0151] In embodiments of this disclosure, such as Figure 9 As shown, the AI ​​agent 900 may include an input module 910, a processing module 920, and an output module 930.

[0152] Input module 910 is used to receive initial input information and interaction information;

[0153] Processing module 920 is used to determine the initial intent understanding result based on the initial input information received by the input module, use a first intelligent agent to perform interaction mode requirement detection on the initial intent understanding result to obtain interaction mode requirement attributes; call the interaction element that matches the interaction mode requirement attributes to interact with the target object to obtain interaction information representing the target object's requirement intent; control a second intelligent agent to execute the target task based on the interaction information, and obtain the execution result as output information.

[0154] Output module 930 is used to output the output information obtained by the processing module.

[0155] According to embodiments of this disclosure, the input module 910 is responsible for receiving or sensing information such as queries, requests, instructions, signals, or data from the outside world (e.g., users or the external environment), and converting it into a format that the AI ​​agent 900 can understand and process. The input module 910 is the primary link for the AI ​​agent 900 to interact with the outside world, enabling the AI ​​agent 900 to efficiently and accurately obtain necessary "sensory" information from the outside world and respond to this information.

[0156] In the example, input module 910 can input the initial input information and interaction information described above.

[0157] In the example, processing module 920 is the core support for the AI ​​agent 900's ability to handle complex tasks. Processing module 920 can execute the agent-based collaborative interaction methods described above.

[0158] In the example, the performance of the processing module 920 is closely related to the large model on which the AI ​​agent 900 is based. To fully leverage the capabilities of the large model, the internal structure of the processing module 920 can be designed to be highly configurable and scalable to handle various types of tasks and requirements in real-world scenarios.

[0159] Understandably, while large language models possess excellent language understanding and generation capabilities, like humans, their ability to solve tasks is limited without the aid of any tools. However, once the AI ​​agent 900 is given the ability to invoke tools, it can perform tasks such as using a calculator to complete mathematical calculations, using Python to perform data analysis, and using a search engine to generate weather forecasts.

[0160] In the example, output module 930 can output the execution results described above.

[0161] The AI ​​agent 900 according to embodiments of this disclosure can simply and effectively improve the level of intelligence, as well as enhance flexibility and versatility.

[0162] According to embodiments of this disclosure, this disclosure also provides an electronic device, a readable storage medium, and a computer program product.

[0163] According to an embodiment of the present disclosure, an electronic device includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, the instructions being executed by the at least one processor to enable the at least one processor to perform the method described above.

[0164] According to embodiments of the present disclosure, a non-transitory computer-readable storage medium stores computer instructions, wherein the computer instructions are used to cause a computer to perform the method described above.

[0165] According to an embodiment of this disclosure, a computer program product includes a computer program that, when executed by a processor, implements the method described above.

[0166] Figure 10 A schematic block diagram of an example electronic device for implementing an agent-based collaborative interaction method of embodiments of the present disclosure is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device may also represent various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the present disclosure described and / or claimed herein.

[0167] like Figure 10 As shown, device 1000 includes a computing unit 1001, which can perform various appropriate actions and processes according to a computer program stored in read-only memory (ROM) 1002 or a computer program loaded into random access memory (RAM) 1003 from storage unit 1008. The RAM 1003 may also store various programs and data required for the operation of device 1000. The computing unit 1001, ROM 1002, and RAM 1003 are interconnected via bus 1004. Input / output (I / O) interface 1005 is also connected to bus 1004.

[0168] Multiple components in device 1000 are connected to I / O interface 1005, including: input unit 1006, such as keyboard, mouse, etc.; output unit 1007, such as various types of monitors, speakers, etc.; storage unit 1008, such as disk, optical disk, etc.; and communication unit 1009, such as network card, modem, wireless transceiver, etc. Communication unit 1009 allows device 1000 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.

[0169] The computing unit 1001 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 1001 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 1001 performs the various methods and processes described above, such as agent-based collaborative interaction methods. For example, in some embodiments, the agent-based collaborative interaction method can be implemented as a computer software program tangibly contained in a machine-readable medium, such as storage unit 1008. In some embodiments, part or all of the computer program can be loaded and / or installed on device 1000 via ROM 1002 and / or communication unit 1009. When the computer program is loaded into RAM 1003 and executed by the computing unit 1001, one or more steps of the agent-based collaborative interaction method described above can be performed. Alternatively, in other embodiments, the computing unit 1001 may be configured to perform an agent-based interaction method by any other suitable means (e.g., by means of firmware).

[0170] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SoCs), payload-programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.

[0171] The program code used to implement the methods of this disclosure may be written in any combination of one or more programming languages. This program code may be provided to a processor or controller of a general-purpose computer, special-purpose computer, or other programmable data processing apparatus, such that when executed by the processor or controller, the program code causes the functions / operations specified in the flowcharts and / or block diagrams to be implemented. The program code may be executed entirely on a machine, partially on a machine, as a standalone software package partially on a machine and partially on a remote machine, or entirely on a remote machine or server.

[0172] In the context of this disclosure, a machine-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.

[0173] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device for displaying information to the user (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor); and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the computer. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).

[0174] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as a data server), or computing systems that include middleware components (e.g., an application server), or computing systems that include frontend components (e.g., a user computer with a graphical user interface or web browser through which a user can interact with implementations of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., a communication network). Examples of communication networks include local area networks (LANs), wide area networks (WANs), and the Internet.

[0175] Computer systems can include clients and servers. Clients and servers are generally located far apart and typically interact via communication networks. Client-server relationships are created by computer programs running on the respective computers and having a client-server relationship with each other. Servers can be cloud servers, distributed system servers, or servers incorporating blockchain technology.

[0176] It should be understood that the various forms of processes shown above can be used to rearrange, add, or delete steps. For example, the steps described in this disclosure can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution disclosed in this disclosure can be achieved, and this is not limited herein.

[0177] The specific embodiments described above do not constitute a limitation on the scope of protection of this disclosure. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this disclosure should be included within the scope of protection of this disclosure.

Claims

1. An interaction method based on agent collaboration, comprising: The initial input information for the target object is subjected to intent understanding to obtain an initial intent understanding result. The initial intent understanding result includes multiple intent understanding information with structured attributes. The intent understanding information includes intent attribute types and intent fields corresponding to the intent attribute types. The interaction pattern requirement is detected by the first intelligent agent based on the initial intent understanding result, and the interaction pattern requirement attribute is obtained. The interactive element that matches the required attributes of the interaction mode is invoked to interact with the target object, thereby obtaining interactive information that represents the required intent of the target object. The interactive element is used to respond to the interactive behavior. Based on the interactive information, the second intelligent agent is controlled to execute the target task, obtain the execution result, and push the execution result to the target object; The interaction mode requirement attribute further characterizes the semantic ambiguity of the initial intent understanding result, indicating whether it meets a preset ambiguity condition; the step of invoking an interaction element matching the interaction mode requirement attribute to interact with the target object includes: The dialog interaction element is invoked to interact with the target object, and the dialog interaction element is used to guide the target object to its needs and intentions.

2. The method according to claim 1, wherein, The interaction mode requirement attribute indicates that the semantic accuracy requirement of the interaction information meets the preset accuracy conditions. The step of invoking an interactive element that matches the interaction mode requirement attribute to interact with the target object includes: An interactive option element representing preset option data is invoked to interact with the target object, wherein the interactive option element is used to determine the target option data in the preset option data in response to the interactive behavior, and the interactive information includes the target option data.

3. The method according to claim 1, wherein, The interaction mode requirement attribute also indicates that the emotional requirement attribute of the initial intention understanding result meets the preset emotional requirement condition. The step of invoking an interactive element that matches the interaction mode requirement attribute to interact with the target object includes: The virtual object interaction element is invoked to interact with the target object, wherein the virtual object interaction element is used to interact with the target object by driving the virtual object to perform verbal communication.

4. The method according to claim 1, wherein, The method further includes: Based on the interaction mode requirement attributes, multiple types of interactive elements are determined, as well as the order in which these interactive elements are invoked.

5. The method according to claim 1, wherein, The process of performing intent understanding on the initial input information for the target object to obtain the initial intent understanding result includes: The initial input information is subjected to structured semantic understanding by calling a large language model, resulting in multiple intent understanding information with structured attributes.

6. The method according to claim 5, wherein, The multiple intent-understanding information pieces have a hierarchical structural relationship.

7. The method according to claim 1, wherein, The step of controlling the second intelligent agent to execute the target task based on the interactive information includes: The first intelligent agent performs semantic understanding on the interaction information to obtain task requirement information; The second intelligent agent is controlled to perform task orchestration based on the task requirement information to obtain the target task, and then the second intelligent agent is controlled to execute the target task.

8. The method according to claim 7, wherein, The control of the second intelligent agent involves task orchestration based on the task requirement information to obtain the target task, including: The control planning agent orchestrates tasks based on the task requirement information to obtain multiple target tasks and the dependencies between the multiple target tasks; The second intelligent agent includes the planning intelligent agent, which issues multiple target tasks and dependencies between the multiple target tasks to the execution intelligent agent. The execution intelligent agent executes the multiple target tasks according to the dependencies to obtain the execution result.

9. The method according to claim 8, wherein, The control planning agent performs task orchestration based on the task requirement information, including: The planning agent is controlled to orchestrate tasks based on the task requirement information and the functional description information for candidate execution agents, thereby determining multiple target tasks and execution agents among the candidate execution agents for executing the target tasks, wherein the dependency relationship represents the calling order of the multiple execution agents.

10. The method according to claim 1, wherein, The method further includes: In response to an execution feedback request for the execution result, the second agent is controlled to update the target task based on the feedback information carried in the execution feedback request, and the updated target task is executed to obtain the feedback execution result.

11. The method according to claim 10, wherein, The step of controlling the second agent to update the target task based on the feedback information carried in the execution feedback request includes: The first intelligent agent performs semantic understanding of the execution result and the target task based on the feedback information to obtain feedback task requirement information; and The second intelligent agent is controlled to perform task orchestration based on the feedback task requirement information to obtain the updated target task.

12. The method according to claim 1, wherein, The target task includes at least one of the following: Tasks include travel strategy planning, copywriting generation, graphic element generation, and image editing.

13. An interactive device based on agent collaboration, comprising: The intent understanding module is used to understand the initial input information for the target object and obtain an initial intent understanding result. The initial intent understanding result includes multiple intent understanding information with structured attributes. The intent understanding information includes intent attribute types and intent fields corresponding to the intent attribute types. The detection module is used to detect interaction mode requirements based on the initial intent understanding results of the first intelligent agent, and obtain interaction mode requirement attributes. The calling module is used to call an interactive element that matches the interaction mode requirement attribute to interact with the target object and obtain interactive information that represents the target object's requirement intention. The interactive element is used to respond to the interactive behavior. The execution module is used to control the second intelligent agent to execute the target task based on the interaction information, obtain the execution result, and push the execution result to the target object; The interaction mode requirement attribute further characterizes the semantic ambiguity of the initial intent understanding result, indicating whether it meets a preset ambiguity condition; the invocation module includes: The second invocation unit is used to invoke a dialogue interaction element to interact with the target object, and the dialogue interaction element is used to guide the target object's needs and intentions.

14. The apparatus according to claim 13, wherein, The interaction mode requirement attribute indicates that the semantic accuracy requirement of the interaction information meets the preset accuracy conditions. The calling module includes: The first invocation unit is used to invoke an interactive option element representing preset option data to interact with the target object, wherein the interactive option element is used to determine the target option data in the preset option data in response to the interactive behavior, and the interactive information includes the target option data.

15. An artificial intelligence agent system configured to perform the method as described in any one of claims 1 to 12.

16. An electronic device comprising: At least one processor; as well as A memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor to enable the at least one processor to perform the method of any one of claims 1 to 12.

17. A non-transitory computer-readable storage medium storing computer instructions, wherein, The computer instructions are used to cause the computer to perform the method according to any one of claims 1 to 12.

18. A computer program product comprising a computer program that, when executed by a processor, implements the method according to any one of claims 1 to 12.

Citation Information

Patent Citations

  • Multi-agent-driven multi-mode cognitive method and device, electronic equipment and medium

    CN119961683A

  • Sample corpus generation method, training method and testing method for large model

    CN120196720A