Information processing method for agent, and related device

WO2025175706A9PCT designated stage Publication Date: 2026-08-13HUAWEI CLOUD COMPUTING TECHNOLOGIES CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-08-07
Publication Date
2026-08-13

Smart Images

  • Figure CN2024110316_13082026_PF_FP_ABST
    Figure CN2024110316_13082026_PF_FP_ABST
Patent Text Reader

Abstract

Provided in the present application are an information processing method for an agent, and a related device, which are used to simplify the process of development of embodied agents and improve development efficiency. The method comprises: displaying a conversation interface comprising an input field, and in response to a first input operation for the input field, acquiring a first instruction, wherein the language of the first instruction is natural language; and on the conversation interface, displaying first response information given in reply by a first agent, wherein the first response information is specific to the first instruction, and the first agent is determined on the basis of the first instruction.
Need to check novelty before this filing date? Find Prior Art

Description

An information processing method and related equipment for an intelligent agent

[0001] This application claims priority to Chinese Patent Application No. 202410204964.1, filed on February 23, 2024, entitled "An Embodied Intelligent Agent Development and Operation Management System Based on Multi-Agent Interaction," and Chinese Patent Application No. 202410660467.2, filed on May 22, 2024, entitled "An Information Processing Method and Related Equipment for an Intelligent Agent," both of which are incorporated herein by reference. Technical Field

[0002] This application relates to the field of cloud computing, and in particular to an information processing method and related equipment for an intelligent agent. Background Technology

[0003] With the development of artificial intelligence (AI), various industries have been affected by AI agents. An intelligent agent can be understood as an intelligent proxy capable of performing tasks on behalf of a user. Intelligent agents include embodied agents, which are intelligent agents that interact or provide feedback to physical entities, including virtual objects, digital entities, robots, etc., and can be represented graphically.

[0004] In traditional technical solutions, web services require users to navigate through multiple pages to complete different needs and tasks, resulting in time-consuming service operations. Furthermore, existing embodied agent development solutions require the cooperation of multiple modules. Developing based on these traditional service models leads to a complex and lengthy overall process with low development efficiency.

[0005] Summary of the Invention

[0006] This application provides an information processing method and related equipment for intelligent agents, which simplifies the development process of embodied intelligent agents and improves development efficiency.

[0007] Firstly, this application provides an information processing method for an intelligent agent, including:

[0008] The information processing method for intelligent agents provided in this application is applied to a cloud platform, which runs an intelligent agent system. This system includes multiple intelligent agents, including at least one of real-body intelligent agents, virtual-body intelligent agents, and AI intelligent agents. Real-body intelligent agents correspond to physical entities in a physical environment, while virtual-body intelligent agents correspond to virtual entities in a virtual environment. Both real-body and virtual-body intelligent agents can interact with or act upon physical entities, thus affecting the physical environment. AI intelligent agents are used to manage other intelligent agents in the intelligent agent system or to provide online services.

[0009] The cloud platform displays a session interface via a computer device logged into a user's registered account on the cloud platform, providing intelligent agent services. The session interface includes an input field. The cloud platform responds to operations on the input field, obtaining user-inputted commands. For example, the cloud platform responds to a first input operation on the input field, obtaining a first command, the language of which is natural language. The first input operation can be various, including touch operation, input via an input device, or other input methods such as voice input; specific methods are not limited here. Input devices include keyboards, mice, styluses, etc., and specific methods are not limited here. Furthermore, natural language in this application refers to language used for human communication, including both written and spoken forms; this application does not limit the language of the natural language. The cloud platform obtains the first command and displays a first response message from the first intelligent agent on the session interface. The first response message is a response to the first command. The language of the first response message is either natural language or a code language, and specific methods are not limited here. The first intelligent agent is determined based on the first instruction, and the intelligent agent system includes multiple intelligent agents, including the first intelligent agent. Further, the first instruction includes an executing subject and / or an event. The executing subject indicates the object executing the first instruction, or in other words, the executing subject is the subject of the first instruction. The event indicates the purpose of the first instruction, or the intent of the first instruction, representing what the first instruction aims to accomplish or achieve.

[0010] In this application, interaction with the first intelligent agent through natural language simplifies complex development or management processes. Furthermore, in schemes where the first intelligent agent includes an embodied intelligent agent or the first intelligent agent manages an embodied intelligent agent, the complexity of developing embodied intelligent agents is reduced, and development efficiency is improved. In some optional implementations of the first aspect, the first intelligent agent is a real embodied intelligent agent, that is, a physical entity in the physical environment corresponding to the first intelligent agent. When the executing subject of the first instruction is the first intelligent agent, and the events included in the first instruction match the functions of the first intelligent agent, it means that the first intelligent agent can complete the event indicated by the first instruction. Therefore, the first response information includes at least one of the following: the state of the first intelligent agent and the event state corresponding to the first instruction. The state of the first intelligent agent includes information reflecting its current state, such as whether the first intelligent agent has accepted the first instruction, whether the first intelligent agent is currently executing a task, and whether the first intelligent agent is malfunctioning / disconnected / has insufficient power. The specific content of the event state corresponding to the first instruction varies depending on the events included in the first instruction. In summary, it can include event-related information such as the current completion status of the event, the current stage of the event, the estimated time for completion, and whether any obstacles have been encountered during event processing. The language of the first response information is natural language.

[0011] In this application, when the first intelligent agent is a real embodied intelligent agent, the first instruction directs the first intelligent agent, and the first intelligent agent is capable of completing the events included in the first instruction, the first response information can have multiple possibilities, enriching the implementation methods of the technical solution of this application. Furthermore, based on the first response information, the state of the first intelligent agent and / or the events included in the first instruction can be understood, and it can be determined whether to adjust the intelligent agent executing the first instruction, thereby accelerating the completion of the events included in the first instruction and improving the efficiency of event completion.

[0012] In some optional implementations of the first aspect, the first intelligent agent is a real embodied intelligent agent, corresponding to a physical entity in the physical environment. If the executing entity of the first instruction is the first intelligent agent, and the event included in the first instruction does not match the function of the first intelligent agent, it means that the first intelligent agent cannot complete the event included in the first instruction. In this case, the first intelligent agent will call other intelligent agents to complete the event included in the first instruction. That is, the first response information includes call information for the embodied intelligent agent. The second intelligent agent is a real embodied intelligent agent, and the function of the second intelligent agent matches the event included in the first instruction; or, the second intelligent agent is the first AI intelligent agent, which is used to manage the intelligent agent system. The call information for the second intelligent agent refers to calling the second intelligent agent to implement the event included in the first instruction, or to respond to the first instruction. Furthermore, the language of the first response information can be natural language, that is, the language of the call information for the second intelligent agent can be natural language.

[0013] In this application, in the scheme where the first intelligent agent, as instructed by the first instruction, cannot complete the event included in the first instruction, the first response information replied by the first intelligent agent is actually a call information for the second intelligent agent, which then implements the event included in the first instruction. In other words, in this application, intelligent agents can also call each other through natural language, further improving the development efficiency of intelligent agents and enhancing their intelligence.

[0014] In some optional implementations of the first aspect, in a scheme where the first response information includes invocation information for the second agent, after displaying the first response message on the session interface, the cloud platform also displays a second response message from the second agent on the session interface, which is in response to the first instruction. That is, the second response message is actually the second agent's reply to the content of the first instruction. Furthermore, the second response information varies depending on the second agent. Optionally, in a scheme where the function of the second agent matches the event included in the first instruction, the second response information includes the state of the second agent and / or the event state corresponding to the first instruction. In a scheme where the second agent is a first AI agent, the second response information includes invocation information for the agent matching the event included in the first instruction, or an execution failure message. The execution failure message indicates that the agent system cannot execute the event included in the first instruction. In general, in a scheme where the second agent is a first AI agent, the first AI agent sends invocation information to other agents, instructing the agent to implement the event included in the first instruction, or the first AI agent replies with an execution failure message.

[0015] In this application, in a scenario where the first intelligent agent, as instructed by the first instruction, is unable to complete the event included in the first instruction, after the first intelligent agent replies with the first response information, the second intelligent agent can also reply with the second response information. Both the second intelligent agent and the second response information have multiple possibilities, enriching the implementation methods and application scenarios of the technical solution of this application, and enhancing the flexibility and practicality of the technical solution.

[0016] In some optional implementations of the first aspect, in schemes where the first instruction does not include an executing entity, the first intelligent agent is defaulted to the first AI intelligent agent, which can also be called the management intelligent agent, used to manage the intelligent agent system. Then, the first response information includes either a call message for the third intelligent agent or an execution failure message. The function of the third intelligent agent matches the event included in the first instruction. The call message for the third intelligent agent refers to calling the third intelligent agent to implement the event included in the first instruction. The execution failure message indicates that the intelligent agent system cannot execute the event included in the first instruction.

[0017] In this application, where the first instruction does not include an executing entity, the first AI agent manages and provides a default first response message. This means that even if the user's first instruction is incomplete, a response can still be received, simplifying user operations and further simplifying agent development. Furthermore, the first AI agent can invoke other agents to complete the events included in the first instruction or provide a message indicating that execution is not possible, allowing the user to understand the progress of the events included in the first instruction and enhancing the practicality of the technical solution presented in this application.

[0018] In some optional implementations of the first aspect, in the scheme where the first instruction includes an executing subject but not an event, the first intelligent agent is the intelligent agent instructed by the executing subject. The first response information includes a prompt message, which is used to remind the user to input an event, or to display the default function of the first intelligent agent, or to display the most recent operation of the first intelligent agent. Alternatively, the first response information is used to trigger the default operation of the first intelligent agent, such as triggering the default function of the first intelligent agent, or triggering the most recent operation of the first intelligent agent, etc. The specific content of the default operation can be set based on the needs of the actual application, and is not limited here.

[0019] In some optional implementations of the first aspect, the cloud platform responds to a second input operation in the input field by obtaining a second instruction. This second instruction directs the display of a data stream of a target event; in other words, the second instruction includes a data stream of the target event. The target event is an event that the intelligent agent system has completed or is in the process of executing. The language of the second instruction is natural language. The second input operation is similar to the aforementioned first input operation and has several possibilities, which will not be elaborated here. After obtaining the second instruction, the data stream of the target event is displayed.

[0020] In this application, the data stream of the target event can vividly and graphically display the state of the target event, allowing users to intuitively experience the state of the target event and improving the user experience.

[0021] In some alternative implementations of the first aspect, before displaying the data stream of the target event, the cloud platform displays third response information on the session interface. This third response information is in response to the second instruction; that is, it is a response to the second instruction. Therefore, the computer device displays the data stream of the target event in response to a touch operation on the third response information. This third response information can be a medium such as a link or file that indicates the data stream of the target event. The user performs a touch operation on this medium, causing the data stream of the target event to be displayed.

[0022] In this application, after the cloud platform obtains the second instruction, it can not only directly display the data stream of the target event, but also display the data stream of the target event by displaying the third response information in the session interface, which enriches the implementation method and application scenarios of the technical solution of this application and improves the flexibility of the technical solution.

[0023] In some optional implementations of the first aspect, the display location of the target event's data stream can vary, including displaying it in the session interface or on an interface outside the session interface; no specific limitation is made here. This allows for flexible configuration based on actual applications, further enhancing the flexibility of the technical solution presented in this application.

[0024] In some optional implementations of the first aspect, the data stream of the target event can have multiple possibilities, including a real-time data stream of the target event, a historical data stream of the target event, a first-person perspective data stream of the target event, or a third-person perspective data stream of the target event, without any specific limitation here. Among them, the first-person perspective is the perspective of the intelligent agent executing the target event, and the third-person perspective can also be called the God's-eye view or the observer's perspective, that is, the perspective of the intelligent agent system.

[0025] In this application, the data stream type from the target perspective can be of various types, which can be determined according to the content indicated by the second instruction or set by default, further enhancing the flexibility of the technical solution of this application.

[0026] In some alternative implementations of the first aspect, the cloud platform responds to a third input operation in the input field and obtains a third instruction. This third instruction instructs the setup of the target simulation environment; that is, the event included in the third instruction is the setup of the target simulation environment. Furthermore, the language of the third instruction is natural language. Based on the third instruction, the computer device displays the target simulation environment on the simulation interface.

[0027] In this application, the computer equipment can also display a simulation environment, through which the simulation or modeling of physical entities can be realized, and the embodied intelligent agent can be developed and tested.

[0028] In some optional implementations of the first aspect, the cloud platform responds to a fourth input operation in response to the input field, obtaining a fourth instruction. This fourth instruction instructs the generation of target-type data; that is, the event included in the fourth instruction is the generation of target-type data. Furthermore, the language of the fourth instruction is natural language. Target types can be various, including radar data, depth data, RGB data, trajectory data, etc., without specific limitations here. Based on the fourth instruction, the cloud platform displays fourth response information on the session interface, which includes data indicating the target type. Further, the fourth response information may include target-type data or a medium indicating the target-type data. In the latter scenario, the computer device responds to a touch operation in response to this medium, displaying the target-type data.

[0029] In this application, the data type generated by the simulation environment can also be indicated. There are multiple possibilities for the data type and the fourth response information, which enriches the implementation methods of the technical solution of this application.

[0030] In some optional implementations of the first aspect, the cloud platform responds to a fifth input operation in the input field, obtains a fifth instruction, the language of which is natural language, and the events included in the fifth instruction match the functionality of the virtual avatar agent. The cloud platform displays a fourth response message from the virtual avatar agent in the conversation interface, which includes the state of the virtual avatar agent and / or the event state corresponding to the fourth instruction.

[0031] In this application, virtual embodied intelligent agents can also be invoked to complete instructions, thereby changing the virtual environment. It can also serve as a simulation of changes to the real physical environment, providing a reference for applying corresponding physical entities in the physical environment, or in other words, providing simulation data, thus enhancing the practicality of the technical solution in this application.

[0032] In some alternative implementations of the first aspect, in a scheme where the executing entity included in the fifth instruction is not the virtual avatar intelligent agent shown in the previous implementation, before displaying the fourth response information, the cloud platform displays response information replied by the executing entity included in the fifth instruction. This response information includes invocation information for the virtual avatar intelligent agent, which instructs the virtual avatar intelligent agent to reply to the fifth instruction.

[0033] In some optional implementations of the first aspect, the cloud platform responds to a sixth input operation in the input field and obtains a sixth instruction. This sixth instruction instructs the generation of skill code for a fourth intelligent agent. In other words, the event included in the sixth instruction is the generation of skill code for the fourth intelligent agent. This skill code describes the functions possessed by the fourth intelligent agent and, in addition, its operational logic. Furthermore, the language of the fourth instruction is natural language. Based on the sixth instruction, the cloud platform displays the skill code of the fourth intelligent agent in the code interface.

[0034] In this application, the computer device can also generate and display the skill codes of the intelligent agent, providing a basis for code detection and improving the feasibility of the technical solution.

[0035] In some alternative implementations of the first aspect, the cloud platform can also test the skill code of the fourth agent and display the test results of the code, so that users can judge whether the code can be executed accurately, thereby improving the accuracy of agent development.

[0036] Secondly, this application provides an information processing apparatus for an intelligent agent, capable of implementing the method described in the first aspect or any possible implementation of the first aspect. The apparatus includes corresponding units or modules for executing the aforementioned method. The units or modules included in the apparatus can be implemented in software and / or hardware.

[0037] Thirdly, this application provides a computer device including a processor and a memory, wherein the processor stores instructions that, when executed on the processor, implement the method shown in the first aspect or any possible implementation of the first aspect.

[0038] Fourthly, this application provides a computer-readable storage medium storing instructions that, when executed on a processor, implement the method shown in the first aspect or any possible implementation of the first aspect.

[0039] Fifthly, this application provides a computer program product that, when executed on a processor, implements the method shown in the first aspect or any possible implementation of the first aspect.

[0040] The beneficial effects shown in any of the second to fifth aspects are similar to those in the first aspect or any possible implementation of the first aspect, and will not be repeated here. Attached Figure Description

[0041] Figure 1 is a schematic diagram of the system architecture provided in an embodiment of this application;

[0042] Figure 2 is a flowchart illustrating an information processing method for an intelligent agent provided in an embodiment of this application.

[0043] Figure 3 is a schematic diagram of a session interface provided in an embodiment of this application;

[0044] Figure 4 is a schematic diagram of an input field provided in an embodiment of this application;

[0045] Figure 5 is another schematic diagram of the input field provided in an embodiment of this application;

[0046] Figure 6 is another schematic diagram of the session interface provided in an embodiment of this application;

[0047] Figure 7 is another schematic diagram of the session interface provided in an embodiment of this application;

[0048] Figure 8 is a schematic diagram of an interface provided in an embodiment of this application;

[0049] Figure 9 is another schematic diagram of an interface provided in an embodiment of this application;

[0050] Figure 10 is another schematic diagram of the session interface provided in an embodiment of this application;

[0051] Figure 11 is a schematic diagram of a simulation interface provided in an embodiment of this application;

[0052] Figure 12 is a schematic diagram of the code interface provided in an embodiment of this application;

[0053] Figure 13 is another schematic diagram of an interface provided in an embodiment of this application;

[0054] Figure 14 is another schematic diagram of an interface provided in an embodiment of this application;

[0055] Figure 15 is a schematic diagram of the structure of a cloud platform provided in an embodiment of this application;

[0056] Figure 16 is a schematic diagram of a computing device provided in an embodiment of this application;

[0057] Figure 17 is a schematic diagram of a computing device cluster provided in an embodiment of this application;

[0058] Figure 18 is another schematic diagram of the computing device cluster provided in the embodiment of this application. Detailed Implementation

[0059] This application provides an information processing method and related equipment for intelligent agents, which simplifies the development process of embodied intelligent agents and improves development efficiency.

[0060] The embodiments of this application will now be described with reference to the accompanying drawings. Those skilled in the art will recognize that, with technological advancements and the emergence of new scenarios, the technical solutions provided in the embodiments of this application are equally applicable to similar technical problems.

[0061] The terms "first," "second," etc., used in the specification, claims, and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such terms are interchangeable where appropriate; this is merely a way of distinguishing objects with the same attributes in the embodiments of this application. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion, so that a process, method, system, product, or apparatus that comprises a series of units is not necessarily limited to those units, but may include other units not explicitly listed or inherent to those processes, methods, products, or apparatuses. Additionally, "at least one" means one or more, and "more than one" means two or more. "And / or" describes the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A alone, A and B simultaneously, or B alone, where A and B can be singular or plural. The character " / " generally indicates that the preceding and following related objects are in an "or" relationship. "At least one of the following" or similar expressions refer to any combination of these items, including any combination of single or multiple items. For example, at least one of a, b, or c can be expressed as: a, b, c, ab, ac, bc, or abc, where a, b, and c can be single or multiple.

[0062] First, please refer to Figure 1, which is a schematic diagram of the system architecture provided in the embodiment of this application.

[0063] As shown in Figure 1, a tenant logs into the cloud platform 30 via client 10 through the Internet 20 using the account and password registered on the cloud platform 30. The cloud platform 30 manages the infrastructure, which includes multiple data centers located in different regions, with each region having at least one cloud data center. For example, region 1 in Figure 1 includes cloud data center 1 and cloud data center 2, and region 2 includes cloud data center 3 and cloud data center 4. Each cloud data center has multiple servers running business instances (including at least one of virtual machines, containers, and dedicated servers).

[0064] A tenant, also known as a user, is a top-level object used to manage cloud services and / or cloud resources. Tenants register a tenant account and set a tenant password on the cloud platform 30 through a local client (e.g., a browser). The local client then remotely logs into the cloud platform 30 using the tenant account and password. The cloud platform 30 provides a configuration interface or API for tenants to configure and use cloud services.

[0065] In this embodiment, the cloud platform 30 runs an intelligent agent system, which includes multiple intelligent agents, including at least one of real intelligent agents, virtual intelligent agents, and AI intelligent agents.

[0066] A real embodied intelligent agent corresponds to a physical entity in a physical environment. It executes corresponding operations based on instructions, interacts with the physical entity, or changes its state. A virtual embodied intelligent agent corresponds to a virtual embodied intelligent agent in a virtual environment. It executes corresponding operations based on instructions, can change the virtual environment, and thus affect physical entities. In general, both real and virtual embodied intelligent agents refer to intelligent agents that interact with or can act upon physical entities. On client 10, embodied intelligent agents can be represented graphically; that is, embodied intelligent agents are actually entities represented graphically. For example, embodied intelligent agents can also be virtual objects, digital entities, or robots, etc., without further limitation here.

[0067] AI agents are used to manage other agents within an agent system or to provide online services. They are also used to assist in the development or operation of embodied agents.

[0068] In this embodiment, the user develops or manages the embodied intelligent agents included in the intelligent agent system running on the cloud platform 30 through the operation terminal 10. The specific implementation process is described in detail below.

[0069] In some optional implementations, the physical environment corresponding to the system architecture provided in this application embodiment can be scenarios such as logistics, port freight, factory area, home, office, etc., and the simulation environment can be a simulation of the aforementioned physical environment, which is not limited here.

[0070] It should be noted that in practical applications, client 10 can also be other types of devices, such as mobile phones, tablets, wearable devices, vehicles, drones, smart home devices, etc. Computer devices can be widely used in various scenarios, such as device-to-device (D2D), vehicle-to-everything (V2X) communication, machine-type communication (MTC), Internet of Things (IoT), virtual reality, augmented reality, industrial control, autonomous driving, telemedicine, smart grids, smart furniture, smart offices, smart wearables, smart transportation, smart cities, etc.

[0071] Please refer to Figure 2 below. Figure 2 is a flowchart illustrating the interaction method of the embodied intelligent agent provided in the embodiments of this application, including:

[0072] 201. The cloud platform displays a session interface, which includes an input field.

[0073] The cloud platform displays a conversational interface via computer devices. This interface showcases interactions between the agent and the user, or between agents themselves. These interactions can be expressed through natural language, either text or voice. In simple terms, the conversational interface can be understood as a chat interface.

[0074] The conversation interface includes an input field, which can be understood as a window for user interaction with the intelligent agent. The user inputs data into the input field to issue commands. In practical applications, the conversation interface and input field can take various forms; Figure 3 is used as an example for illustration below. Please refer to Figure 3, which is a schematic diagram of the conversation interface provided in an embodiment of this application.

[0075] As shown in Figure 3, the conversation interface 300 includes an input field 310. In the embodiment shown in Figure 3, the user enters the content "deliver cargo 1 from port A to port B" in the input field 310. When the user clicks the send control 311, the aforementioned content will be displayed in the conversation interface as shown in the conversation box 330.

[0076] In addition, as shown in Figure 3, the conversation interface 300 can also display the conversation subject avatar 320 to distinguish the sender of each conversation box 330.

[0077] It is understood that the input field 310 shown in Figure 3 is merely an illustration, and in practical applications, the input field 310 may have other forms. For example, please refer to Figure 4, which is a schematic diagram of the input field provided in an embodiment of this application.

[0078] As shown in Figure 4, in addition to the send control 311, the input field 310 may also include other functional controls 312. For example, the functional controls 312 shown in Figure 4, from left to right, are an attachment control, an email control, and a voice control. The attachment control is used to add or modify attachments, the email control is used for email transmission, and the voice control is used for voice input. In practical applications, the input field 310 may include more or fewer functional controls; this is not limited here.

[0079] 202. The cloud platform responds to the first input operation in the input field by obtaining the first instruction, and the language of the first instruction is natural language.

[0080] When a user performs a first input operation in the input field, the cloud platform responds by obtaining a first instruction. In this embodiment, the language of the first instruction is natural language; that is, the language used in the user's input is natural language. Natural language includes text and speech, and refers to the language used for human communication. In this embodiment, the language of the natural language is not limited; for example, it can be Chinese, English, Korean, etc., determined based on the needs of the actual application, and is not limited here.

[0081] For example, in the implementation shown in Figure 3, the user enters the content "deliver goods 1 from port A to port B" in the input field 310 and clicks the send control 311, so that the cloud platform obtains the first instruction "deliver goods 1 from port A to port B".

[0082] Optionally, users can perform the first input operation through an input device, such as a keyboard, mouse, stylus, etc., without specific limitations here. Optionally, if the computer device logged into a cloud platform account has a touchscreen screen, users can also perform the first input operation directly on the input field displayed on the computer device using a stylus or finger.

[0083] 203. The cloud platform displays the first response information of the first intelligent agent in the conversation interface. The first response information is for the first instruction, and the first intelligent agent is determined based on the first instruction.

[0084] After the cloud platform receives the first instruction, it displays the first response information from the first intelligent agent in the conversation interface. In other words, the sender of the first response information is the first intelligent agent.

[0085] In the embodiments of this application, the first instruction includes an executing subject and / or an event. The executing subject is used to indicate the object that completes the event. In layman's terms, the executing subject is the subject of the first instruction, indicating who completes the events included in the first instruction. The events of the first instruction are used to indicate the purpose or intention of the first instruction. In layman's terms, the event represents what the first instruction is to accomplish.

[0086] In this application example, the first intelligent agent is determined based on the first instruction. Since the content of the first instruction has multiple possibilities, the first intelligent agent also has multiple possibilities, which will be explained below.

[0087] (a) The first instruction includes the executing entity and the event.

[0088] In this scheme, it can be understood that the user specifies the execution subject and the event that the execution subject needs to complete. The execution subject specified by the user is the first intelligent agent. There are several possibilities for the user to specify the first intelligent agent, such as by using special symbols or by using the first intelligent agent as the subject of the input content; the specific method is not limited here.

[0089] For example, please refer to Figure 5, which is a schematic diagram of the input field provided in an embodiment of this application.

[0090] Optionally, as shown in Figure 5(a), the first agent can be designated as "Agent A" using the special symbol @. In practical applications, after the user enters the special symbol, the computer device can also display a list of agents, showing multiple agents included in the list, allowing the user to select the first agent from among them.

[0091] In addition to @ as shown in Figure 5, other special symbols can also be used, such as #, *, etc. They can be set according to the needs of actual applications, and there are no specific restrictions here.

[0092] Optionally, as shown in Figure 5(b), the user inputs: Agent A delivers cargo 1 from port A to port B. The computer device performs semantic recognition and analysis on this content, determines that the subject of the content is Agent A, and thus determines that the first agent is Agent A.

[0093] In practical applications, the first intelligent agent may be unable to complete the event specified by the user. In such cases, the content indicated by the first response information may vary, and will be explained below.

[0094] 1. The first intelligent agent is a real embodied intelligent agent, and the functions of the first intelligent agent match the events.

[0095] The first intelligent agent is a real, embodied intelligent agent, meaning it corresponds to a physical entity in the physical environment—a tangible object such as a robot, robotic arm, camera, or sensor. The function-event matching of the first intelligent agent means that the first intelligent agent can complete the event. Therefore, in this scheme, the first response information includes the state of the first intelligent agent and / or the event state corresponding to the first instruction.

[0096] The state of the first intelligent agent includes information that reflects its current state, such as whether the first intelligent agent has accepted the first instruction, whether the first intelligent agent is currently executing a task, and whether the first intelligent agent is malfunctioning / disconnected / has insufficient power.

[0097] The event status corresponding to the first instruction varies in specific content depending on the events included in the first instruction. In general, it can include event-related information such as the current completion status of the event, the current stage of the event, the estimated time to complete the event, and whether any obstacles have been encountered in the event processing.

[0098] For example, suppose the first instruction is "Agent B delivers one box of dolls from port C to port D," meaning the first agent is agent B, and the event is delivering one box of dolls from port C to port D. The first response information can have different contents depending on the situation:

[0099] Optionally, if agent B is currently executing the previous task, its first response could be "Executing the previous task, please wait." Optionally, if agent B's battery is low, its first response could be "Low battery, please select another agent to perform the task." Optionally, if agent B malfunctions, its first response could be "Under repair, expected to wait 20 minutes." The first response messages in the aforementioned examples all reflect the state of the first embodied agent.

[0100] Optionally, if the road from port C to port D is currently blocked, the first response from agent B could be "The road is currently blocked; please replan your route." Alternatively, agent B's first response could also be "Okay, on my way to pick up the goods," or "Goods have been picked up and are being transported." The first response information in the aforementioned examples all reflect the event state corresponding to the first instruction.

[0101] Optionally, if agent B is idle and the road from port C to port D is unobstructed, then agent B, based on the first instruction, will deliver one box of dolls from port C to port D, thus transferring the box of dolls between the ports and altering the physical entity in the physical environment. Alternatively, it can be understood as agent B interacting with the physical entity.

[0102] For example, suppose the first instruction is "Agent A takes a picture of port A", that is, the first agent is agent A, and the event is taking a picture of port A. Then the first response information can be a picture of port A, which means that the event status corresponding to the first instruction is that the event is completed.

[0103] Optionally, if the first agent is unable to immediately execute the event included in the first instruction, the first response information may also include a prompt message to remind the user to switch agents, or to provide the user with alternative embodied agents capable of executing the event included in the first instruction.

[0104] It should also be noted that the language of the first response information can be natural language, code language, or a custom language, etc., and there is no specific limitation here. For example, in the scheme of first agent failure, the first response information can be a custom code that indicates that the first agent has failed and cannot complete the event included in the first instruction.

[0105] In this embodiment of the application, when the first intelligent agent is a real embodied intelligent agent, the first instruction directs the first intelligent agent, and the first intelligent agent can complete the event included in the first instruction, the first response information has multiple possibilities, enriching the implementation methods of the technical solution of this application. Furthermore, based on the first response information, the state of the first intelligent agent and / or the event included in the first instruction can be understood, and it can be determined whether to adjust the intelligent agent executing the first instruction, thereby accelerating the completion of the event included in the first instruction and improving the efficiency of event completion.

[0106] 2. The first intelligent agent is a real embodied intelligent agent, and the functions of the first intelligent agent do not match the events.

[0107] If the function of the first intelligent agent does not match the event, it means that the first intelligent agent cannot complete the event. In this case, the first intelligent agent calls the second intelligent agent to respond to the first instruction. That is, the first response information includes the call information for the second intelligent agent. The second intelligent agent is the first AI intelligent agent, which is used to manage the intelligent agent system. Alternatively, the second intelligent agent is a real embodied intelligent agent, and the function of the second intelligent agent matches the event included in the first instruction. In the embodiments of this application, the first AI intelligent agent is also referred to as the management intelligent agent.

[0108] The following description is based on a schematic diagram. Please refer to Figure 6, which is a schematic diagram of the session interface provided in an embodiment of this application.

[0109] In the embodiment shown in Figure 6, the first instruction is "Intelligent agent C, deliver 1 box of dolls from port C to port D" as an example. That is, the first intelligent agent is intelligent agent C, and the first instruction includes the event of delivering 1 box of dolls from port C to port D.

[0110] If agent C is unable to complete the event included in the first instruction, in the embodiment shown in Figure 6(a), agent C's first response is to invoke the first AI agent to execute the event included in the first instruction. In the embodiment shown in Figure 6(b), agent C's first response is to invoke agent B to execute the event included in the first instruction. The function of agent B matches the event included in the first instruction.

[0111] It is understood that the embodiment shown in Figure 6 is merely an example of the first response information. In practical applications, the first response information replied by agent C may not include the event, but only the executing entity (i.e., only agent B). In this scheme, agent B can identify the event included in the first instruction based on the session interface.

[0112] In addition, the language of the first response information can be natural language, which means that in the embodiments of this application, different intelligent agents can also call each other through natural language.

[0113] In this application, in the scheme where the first intelligent agent, as instructed by the first instruction, is unable to complete the event included in the first instruction, the first response information replied by the first intelligent agent is actually a call information for the second intelligent agent, which then implements the event included in the first instruction. In other words, in this application, intelligent agents can also call each other through natural language, further simplifying the development efficiency of intelligent agents and enhancing their intelligence.

[0114] In some optional implementations, after the first response information from the first embodied intelligent agent is displayed in the session interface, the computer device may also display a second response information from the second intelligent agent in response to the content of the first instruction. That is, the second response information is actually the second intelligent agent's reply to the content of the first instruction.

[0115] Optionally, if the second agent is a management agent, then the second response information includes invocation information for the agent matching the event included in the first instruction, or an execution failure message. The execution failure message indicates that the agent system cannot execute the event included in the first instruction.

[0116] For example, in the embodiment shown in Figure 6(a), after agent C replies with the first response information, the management agent replies with the second response information. The second response information indicates the event that calls agent B to execute the first instruction. Although only the content of the second response information "@agentB" is shown in Figure 6(a), agent B can identify the content of the conversation interface and determine that the event to be executed is the event of the first instruction, namely "deliver 1 box of dolls from port C to port C".

[0117] Optionally, if the second agent is an embodied agent whose function matches the event included in the first instruction, then the second response information includes the state of the second agent and / or the event state corresponding to the first instruction. Its specific content is similar to the first response information in the aforementioned technical solution "1. The first agent is a real embodied agent, and the function of the first agent matches the event," as shown above, and will not be repeated here.

[0118] For example, in the embodiment shown in Figure 6(b), agent C replies with a first response message instructing agent B to be invoked. After displaying the first response message, the session interface displays a second response message from agent B. The content of the second response message is "OK, received," indicating that the second agent (i.e., agent B) is currently in a normal working state and can execute the events included in the first instruction.

[0119] It is understandable that although the first response information is displayed before the second response information is displayed, the second response information is a reply from the second agent invoked by the first response information, and the content of the second response information is actually a response to the first instruction.

[0120] In this application, in a scenario where the first intelligent agent, as instructed by the first instruction, is unable to complete the event included in the first instruction, after the first intelligent agent replies with the first response information, the second intelligent agent can also reply with the second response information. Both the second intelligent agent and the second response information have multiple possibilities, enriching the implementation methods and application scenarios of the technical solution of this application, and enhancing the flexibility and practicality of the technical solution.

[0121] 3. The first intelligent agent is a virtual embodied intelligent agent, and the functions of the first intelligent agent match the events.

[0122] In the embodiments of this application, the executing entity included in the first instruction can also be a virtual embodied intelligent agent, which corresponds to a virtual entity in a virtual environment. The virtual environment can also be understood as a simulation environment of the physical environment, and the virtual entity is the virtualization or simulation of the physical entity, that is, the virtual embodied intelligent agent is the virtualization or simulation of the real embodied intelligent agent.

[0123] In this technical solution, the content of the first response information is similar to that in the aforementioned technical solution "1. The first intelligent agent is a real embodied intelligent agent, and the function of the first intelligent agent matches the event," as detailed above. The difference lies in that the state of the first intelligent agent included in the first response information is the state of a virtual embodied intelligent agent; the event state corresponding to the first instruction is also the event state in a virtual environment.

[0124] In practical applications, users can determine the operational status of the intelligent agent system in a virtual environment based on the first response information, thereby deciding whether to adjust the intelligent agent system or apply it to a real physical scenario. In other words, evaluating the system through a virtual environment helps reduce development costs and improve development efficiency.

[0125] 4. The first intelligent agent is a virtual embodied intelligent agent, and the functions of the first intelligent agent do not match the events.

[0126] In this technical solution, the content of the first response information is similar to that in the aforementioned technical solution where "2. the first intelligent agent is a real embodied intelligent agent, and the function of the first intelligent agent does not match the event," as detailed above, and will not be repeated here. The difference lies in that, in this technical solution, the first intelligent agent can also invoke other virtual embodied intelligent agents.

[0127] (ii) The first instruction does not include the executing entity, but includes the event.

[0128] In a scheme where the first instruction only includes events, a first AI agent is assumed to be the first agent, and the first response information includes a call message for a third agent and / or an execution failure message. The function of the third agent matches the events included in the first instruction.

[0129] The following description uses specific examples. Please refer to Figure 7, which is a schematic diagram of the session interface provided in the embodiment of this application.

[0130] It should be noted that in the embodiment shown in Figure 7, the first instruction is "deliver one box of food from port C to port D" as an example, that is, the event included in the first instruction is delivering one box of dolls from port C to port D. The first instruction does not include an executing subject, so the first AI agent is assumed to be the first embodied AI agent.

[0131] Optionally, in the embodiment shown in Figure 7(a), the first response information replied by the first AI agent is a call message to the third agent (i.e., agent A). In this scheme, after displaying the first response information, the response information replied by the third agent can also be displayed.

[0132] The response information from the third intelligent agent includes the state of the third intelligent agent and / or the event state corresponding to the first instruction. Its specific content is similar to the first response information in the aforementioned technical solution "1. The first intelligent agent is a real embodied intelligent agent, and the function of the first intelligent agent matches the event," as shown above, and will not be repeated here. For example, in the embodiment shown in Figure 7(a), the response information from the third intelligent agent is "Charging, requires 3 minutes to fully charge," indicating that the third intelligent agent (i.e., agent A) is currently in a charging state.

[0133] Optionally, in the embodiment shown in Figure 7(b), the first response information of the first AI agent is an execution failure message, which means that "this system does not transport food", which means that none of the agents included in the current agent system can execute the first instruction.

[0134] In this application, in schemes where the first instruction does not include an executing entity, the first AI agent defaults to providing a first response message. That is, even if the user's first instruction is incomplete, a response can still be received, simplifying user operations and further simplifying agent development. Furthermore, the first AI agent can invoke other agents to complete the events included in the first instruction, or provide a message indicating that execution is not possible, allowing the user to understand the progress of the events included in the first instruction and enhancing the practicality of the technical solution in this application.

[0135] (iii) The first instruction includes the executing entity but does not include the event.

[0136] In a scheme where the first instruction includes an executing entity but not an event, the first intelligent agent is the intelligent agent instructed by the executing entity. The first response information includes a prompt message, which is used to remind the user to input an event, or to display the default function of the first intelligent agent, or to display the most recent operation of the first intelligent agent.

[0137] As described above, in this embodiment, interaction with the intelligent agent through natural language simplifies the complex development or management process. Furthermore, in schemes where the first intelligent agent includes an embodied intelligent agent or the first intelligent agent manages an embodied intelligent agent, the complexity of developing the embodied intelligent agent is reduced, and development efficiency is improved.

[0138] The foregoing description introduced the interaction between the intelligent agent and the user, as well as the interaction between intelligent agents, provided in the embodiments of this application. The foregoing embodiments are not limited to the events included in the first instruction, but are determined based on the needs of actual applications.

[0139] For example, in the embodiments shown in Figures 3 to 7 above, the first instruction is for a port freight scenario, and the included event is the transportation of goods at the port. In practical applications, the intelligent agent system can also be applied to other scenarios, such as office scenarios and home scenarios, and the first instruction will change accordingly.

[0140] In the embodiments of this application, the intelligent agent system also has other functions, such as data stream display, building or modifying the simulation environment, code generation and detection, etc., which will be described below.

[0141] In some alternative implementations, the intelligent agent system can also display data streams; that is, the cloud platform can also display data streams that reflect the state or trajectory of physical entities in a physical environment, or the state or trajectory of virtual entities in a simulation environment. There are several possible ways for the cloud platform to display the target data stream, which are described below:

[0142] The cloud platform responds to the second input operation in the input field, obtains the second instruction, and the second instruction indicates the display of the data stream of the target event. The target event is an event that the intelligent agent system has completed or is in the process of executing. The language of the second instruction is natural language.

[0143] Optionally, the cloud platform can directly display the data stream of the target event based on the second instruction. The implementation of the second input operation is similar to that of the first input operation described above, and will not be repeated here.

[0144] Optionally, after receiving the second instruction, the cloud platform can also display a third response message in the session interface. This third response message is specific to the second instruction. In other words, the third response message indicates the data stream of the target event. The cloud platform then responds to the touch operation corresponding to the third response message, displaying the data stream of the target event.

[0145] In this application, after the cloud platform obtains the second instruction, it can not only directly display the data stream of the target event, but also display the data stream of the target event by displaying the third response information in the session interface, which enriches the implementation method and application scenarios of the technical solution of this application and improves the flexibility of the technical solution.

[0146] In the embodiments of this application, the data stream of the target event can have multiple possibilities, including a real-time data stream of the target event, a historical data stream of the target event, a first-person perspective data stream of the target event, or a third-person perspective data stream of the target event. The first-person perspective is the perspective of the intelligent agent executing the target event, also known as the perspective of the party involved. The third-person perspective is the perspective of the intelligent agent system, or it can also be understood as a global perspective. In layman's terms, the third-person perspective can also be called a God's-eye view or an observer's perspective; for example, in a port freight scenario, the third-person perspective is the perspective observed by a camera above the port.

[0147] Optionally, since the target event includes events that have completed or are in progress with the intelligent agent system, the data stream type of the target event can be set by default. For example, in a scheme where the target event is an event where the intelligent agent system has completed execution, the data stream of the target event is set to a historical data stream by default; in a scheme where the target event is an event where the intelligent agent system is in progress, the data stream of the target event is set to a real-time data stream by default.

[0148] In the embodiments of this application, there are multiple possibilities for the display location of the target event data stream. It can be displayed in the session interface or in an interface outside the session interface.

[0149] In this application embodiment, the display position of the target event data stream and the target event data stream have multiple possibilities, which can be determined according to the actual application, further improving the flexibility of the technical solution of this application.

[0150] The process of a computer device displaying a target event's data stream is further explained below with reference to the schematic diagrams. Please refer to Figures 8 to 10, where Figures 8 and 9 are schematic diagrams of the interface provided in the embodiments of this application, and Figure 10 is a schematic diagram of the session interface provided in the embodiments of this application.

[0151] It should be noted that in the embodiments shown in Figures 8 to 10, the content of the first instruction is "I want to see the video stream of agent B delivering a box of dolls from port C to port D," that is, the target event is delivering a box of dolls from port C to port D, and the data stream is a first-person perspective video stream, i.e., from agent B's own perspective.

[0152] In the embodiment shown in Figure 8, after the session interface 300 displays the first instruction, the data stream from the perspective of agent B is directly displayed on the data stream display interface 400. Specifically, the display area of ​​the data stream can be shown as the shaded area in Figure 8.

[0153] In the embodiment shown in Figure 9, after the session interface 300 displays the first instruction, it displays the third response information from agent B. The third response information is a URL. In response to a click operation on this URL, the computer device displays the data stream display interface 400, thereby displaying the data stream of the target event. Alternatively, the user clicks the link, and the data stream display interface 400 is displayed.

[0154] It should be noted that the embodiment shown in Figure 9 uses a URL as an example of third response information. In practical applications, third response information can also include other types of media, such as links, files, etc., as long as it is a data stream used to indicate the target event. Furthermore, the data stream display interface 400 can also be displayed through other software. Other software refers to software different from the software providing the session interface 300, such as standalone video playback software.

[0155] In some optional implementations, the third response information may also include historical data streams for different time periods. Users can select the time period they need, either by using natural language or by clicking the corresponding link or selection box, etc. The specifics are not limited here.

[0156] Optionally, after the user clicks on the medium included in the third response information, a prompt sub-interface can be displayed to ask the user whether they confirm the display of the target event's data stream. Upon user confirmation, the data stream display interface 400 is triggered. There are several possible ways for the user to confirm, such as clicking the confirmation control included in the prompt sub-interface, or remaining inactive within a preset time period. The choice can be made based on the actual application needs and is not limited here.

[0157] It should also be noted that in the embodiments shown in Figures 8 and 9, a session interface 300 and a data stream display interface 400 are displayed. In practical applications, the data stream display interface 400 may also cover the session interface 300, or the data stream display interface 400 may overlap with the session interface 300, or the data stream interface 400 may obscure the session interface 300, etc. The display positions of these two interfaces can be selected based on the actual application, and are not limited here.

[0158] Optionally, as shown in Figures 8 and 9, the data stream display interface 400 may also include function controls 410. From left to right, the function controls 410 are a volume control, a rotation control, a rewind control, a pause control, a forward control, and a screenshot control, used to control video volume, rotate the video display angle, play back the video, pause playback, fast forward the video, and take a screenshot, respectively. In practical applications, more or fewer function controls may be included, such as a playback speed control, etc., which is not limited here.

[0159] In the embodiment shown in Figure 10, after the session interface 300 displays the first instruction, the cloud platform displays a video stream of the target event responded to by agent B on the session interface 300. Specifically, the display area of ​​the video stream can be as shown in the shaded area of ​​Figure 10. That is, the data stream of the target event can be embedded in the session interface for display.

[0160] It should also be noted that the shaded areas shown in Figures 8 to 10 are just an example of the video stream display area. In practical applications, this area can also be scaled, rotated, etc. based on user operations, which is not limited here.

[0161] In the embodiments shown in Figures 8 to 10 above, the first instruction specifies the perspective of the data stream of the target event. In practical applications, the perspective may not be specified. In this scheme, the data stream of the target event displayed by the computer device can have multiple possibilities:

[0162] Optionally, the default setting can be to display a first-person view or a third-person view video stream of the target event. A toggle control can also be displayed to switch the view of the target event's data stream.

[0163] Optionally, the system can default to displaying both first-person and third-person view data streams of the target event. Then, based on the user's selection, the data stream from one of the views can be scaled.

[0164] In the embodiments shown in Figures 8 to 10 above, the first instruction specifies which specific event the target event is. In practical applications, a perspective may also be specified, but not the specific event. In this scheme, the data stream of the target event displayed by the cloud platform has several possibilities: Optionally, it can be set to display the data stream of the event currently being executed by the embodied agent by default. Optionally, it can be set to display the historical data stream of the most recent event executed by the current embodied agent by default. Optionally, it can be set to display the data stream of the default time period by default. No specific limitations are imposed here.

[0165] It should also be noted that the embodiments shown in Figures 8 to 10 are based on video streams. In practical applications, the data stream can also be the data stream of other sensors, such as infrared sensors, laser sensors, radar, inertial measurement units (IMUs), etc., and no specific limitation is made here.

[0166] In some optional implementations, the intelligent agent system provided in this application also constructs or modifies a simulation environment. The simulation environment is used to develop or test embodied intelligent agents. In summary, the cloud platform responds to a third input operation on the input field, obtains a third instruction, the language of which is natural language, and the third instruction instructs the construction of the target simulation environment. That is, the event included in the third instruction is the construction of the target simulation environment. Based on the third instruction, the cloud platform displays the target simulation environment on the simulation interface.

[0167] Optionally, if the third instruction does not include an executing agent, the computer device may display response information from the first AI agent in the session interface, which invokes the simulated agent to execute the third instruction.

[0168] Optionally, if the third instruction includes a second AI agent as the executing entity, then the second AI agent executes the third instruction. The second AI agent is used to construct the simulation environment and can also be called a simulation agent.

[0169] Optionally, if the execution entity included in the third instruction is unable to display the target simulation environment, then the execution entity may invoke the second AI agent to execute the third instruction.

[0170] The three possible implementation methods described above are similar to the various possibilities of the cloud platform calling the first intelligent agent to reply to the first instruction through the first response information, and will not be repeated here.

[0171] The simulation interface is used for the visualization of 3D or 2D virtual environments, which are also known as simulation environments. Furthermore, there are various ways the cloud platform can display the simulation interface, similar to the data flow display interface 400 in the embodiments shown in Figures 8 to 10 above, as described above, and will not be repeated here.

[0172] For example, a user inputs a simulation environment requirement, such as "Build a simulation environment of a dock, including 3 ports and 6 piles of cargo, with sunny weather and morning time." This instruction does not specify an executor; the first AI agent calls a second AI agent to execute the instruction. The second AI agent can be a single agent responsible for building the aforementioned simulation environment; or it can be multiple agents, each with different functions, such as changing lighting, the number of objects, textures, and positions. In a scheme with a single second agent, there is no need to build other agents, reducing computational load. In a scheme with multiple second agents, the responsibilities of each agent are clearly defined, facilitating maintenance.

[0173] The different functions mentioned above include, but are not limited to: changing the object attributes such as type, quantity, texture, position, and shape of objects in the simulation environment; changing environmental factors such as lighting and humidity in the environment; and generating events and test cases to be executed by embodied intelligent agents.

[0174] For example, please refer to Figure 11, which is a schematic diagram of the simulation interface provided in the embodiment of this application.

[0175] As shown in Figure 11, the simulation interface displays a 3D dock, including 3 ports and 6 piles of cargo. In the embodiment shown in Figure 11, the 3 ports are adjacent, but due to obstructed view, port C, which is adjacent to port B, is not shown in Figure 11.

[0176] In some alternative implementations, the simulation interface may also include function controls 510. In the embodiment shown in Figure 11, the functions of the function controls 510 from left to right are save, cut, refresh, undo, and select tool, respectively. In practical applications, the simulation interface may include more or fewer function controls, which is not limited here.

[0177] In a simulation environment, a virtual embodied intelligent agent corresponds to a simulated entity within that environment. For example, in the embodiment shown in Figure 11, assuming the simulation environment also includes simulated robot 1 and simulated robot 2, both used for transporting goods, then a user can use a session interface to invoke the virtual embodied intelligent agent to simulate goods transport, or to simulate mutual invocation between embodied intelligent agents.

[0178] In other words, in this embodiment, the cloud platform responds to the fifth input operation in the input field, obtains the fifth instruction, the language of the fifth instruction is natural language, and the events included in the fifth instruction match the functions of the virtual avatar intelligent agent. The cloud platform displays the fourth response information replied by the virtual avatar intelligent agent in the conversation interface. This fourth response information includes the state of the virtual avatar intelligent agent and / or the event state corresponding to the fourth instruction.

[0179] For example, suppose the simulation environment shown in Figure 11 includes a virtual embodied agent A, whose function is to transport goods between port A and port B. The fifth instruction input by the user includes the event of transporting one box of goods from port A to port B.

[0180] Optionally, if the executing entity included in the fifth instruction is a virtual avatar intelligent agent A, then the cloud platform displays the aforementioned fourth response information in the session interface. For example, virtual avatar intelligent agent A replies "OK, received" in the session interface, and corresponding to the simulation environment, virtual avatar A is in a freight transportation state.

[0181] Optionally, if the function of the executing entity included in the fifth instruction does not match the aforementioned event, then the fourth response information can be a call message for the virtual avatar intelligence A, causing the virtual avatar intelligence A to execute the aforementioned event. Additionally, the session interface can also display the response information replied by the virtual avatar intelligence A.

[0182] Optionally, if the fifth instruction does not include an executing entity, the cloud platform may also display the response information of the first AI agent before displaying the fourth response information. This response information is for the call information of the virtual embodied AI agent A.

[0183] In this application, virtual embodied intelligent agents can also be invoked to complete instructions, thereby changing the virtual environment. It can also serve as a simulation of changes to the real physical environment, providing a reference for applying corresponding physical entities in the physical environment, or in other words, providing simulation data, thus enhancing the practicality of the technical solution in this application.

[0184] In this embodiment of the application, the computer device can also determine the data type generated by the simulation environment.

[0185] Optionally, the computer device responds to a fourth input operation on the input field by obtaining a fourth instruction. The fourth instruction is in natural language and indicates the generation of data of the target type. Based on the fourth instruction, a fourth response message is displayed on the session interface, indicating the data of the target type. Optionally, the fourth response information may include the data of the target type, or it may include a medium indicating the data of the target type. In the latter scenario, the cloud platform responds to a touch operation on the medium by displaying the data of the target type.

[0186] Optionally, the cloud platform can also configure the data types generated by the simulation environment by default. In other words, in scenarios where the user does not specify the type of data to be generated, the cloud platform displays the generated data of the default type.

[0187] In this embodiment, the target type can be of various kinds, including radar data, depth data, RGB data, trajectory data, etc., and no specific limitation is made here.

[0188] In this application, the cloud platform can also display a simulation environment, through which the simulation or modeling of physical entities can be realized, and the embodied intelligent agent can be developed and tested. Furthermore, this application can also indicate the data type generated by the simulation environment; the data type and the fourth response information have multiple possibilities, enriching the implementation methods of the technical solution of this application.

[0189] In some optional implementations, the intelligent agent system provided in this application embodiment can also perform code generation and detection. In summary, the cloud platform responds to a sixth input operation in the input field, obtains a sixth instruction, the language of which is natural language, and the sixth instruction instructs the generation of skill code for a fourth intelligent agent. Based on the sixth instruction, the cloud platform displays the skill code of the fourth intelligent agent in the code interface. That is, the event included in the sixth instruction is the generation of skill code for the fourth intelligent agent. This skill code is used to describe the functions possessed by the fourth intelligent agent, and can also describe the working logic of the fourth intelligent agent.

[0190] Optionally, if the sixth instruction does not include an executing agent, the computer device can display a response from the first AI agent in the session interface, which invokes a third AI agent to execute the sixth instruction. The third AI agent is used to generate the agent's code.

[0191] Optionally, if the execution subject of the sixth instruction is a third AI agent, then the third AI agent executes the sixth instruction.

[0192] Optionally, if the capabilities of the executing agent included in the sixth instruction are insufficient to generate the skill code of the fourth agent, then the executing agent may invoke the third AI agent to execute the sixth instruction.

[0193] The three possible implementation methods described above are similar to the various possibilities of the first response information mentioned earlier, and will not be repeated here.

[0194] In addition, there are several possible ways for computer devices to display code interfaces. Similar to the data flow display interface 400 in the embodiments shown in Figures 8 to 10 above, see above, and will not be repeated here.

[0195] In some optional implementations, the code interface may display functional controls in addition to code, for implementing different functions. For example, please refer to Figure 12, which is a schematic diagram of the code interface provided in an embodiment of this application.

[0196] As shown in Figure 12, the code interface includes a code language control, which is used to set the language type of the code displayed on the code interface. For example, when the user clicks this control, the code interface can display a code language selection bar, which includes multiple code languages ​​for the user to select.

[0197] As shown in Figure 12, the code interface can also include code formatting controls to set the code format, including the font, font size, color, etc.

[0198] As shown in Figure 12, the code interface can also include panel style controls to set the style of the code interface, including the background color of the code interface, whether to use night mode, etc.

[0199] In some optional implementations, the intelligent agent system provided in this application embodiment can also detect code. The cloud platform tests the skill code of the fourth intelligent agent and displays the test results of the fourth intelligent agent's skill code.

[0200] Optionally, as shown in Figure 12, the code interface may include test controls. When the user clicks on the test controls, the cloud platform tests the skill code of the fourth intelligent agent displayed on the code interface.

[0201] Optionally, skill code detection can also be triggered within the conversation interface. For example, the user enters "detect the skill code of the fourth agent" in the input field of the conversation interface. The first AI agent calls the code detection agent to detect the skill code of the fourth agent. Then, the code detection agent can reply with the test result of this detection in the conversation interface. Optionally, the code detection agent and the third AI agent used to generate the code mentioned above can be different AI agents or the same AI agent; both can serve as code agents.

[0202] Optionally, the code generated by the cloud platform can be stored in a location by a first AI agent or a code agent, with a specified code name. When users need to test the code in a virtual environment, they can specify the code name and a detection event. The cloud platform can also use a code detection agent to test the code corresponding to that code name. Code agents include code detection agents and code generation agents.

[0203] In this application, the cloud platform can also generate and display the agent's skill code, providing a foundation for code detection and improving the feasibility of the technical solution. Furthermore, by detecting the code, users can determine whether the code can be executed accurately, thereby improving the accuracy of agent development.

[0204] In the embodiments shown in Figures 8 to 12 above, the data flow display interface, code interface, and simulation interface are all displayed based on the user's input operations in the input field. In practical applications, they can also be displayed in other ways, which will be explained below.

[0205] In some alternative implementations, the cloud platform may also display a menu bar, which includes display controls for different interfaces. When a user clicks on a control, the cloud platform displays the interface corresponding to that control.

[0206] In some optional implementations, the cloud platform can also display a search bar. When a user enters an instruction such as "display XX interface" in the search bar, the cloud platform displays the corresponding interface.

[0207] For example, please refer to Figure 13, which is a schematic diagram of the interface provided in an embodiment of this application.

[0208] As shown in Figure 13, the menu bar includes four display controls 610, which are, from left to right, the data flow interface display control, the code interface display control, the simulation interface display control, and the session interface display control.

[0209] Optionally, if the user clicks on the code interface to display controls, the cloud platform can display the code interface shown in Figure 12. If the user clicks on the simulation interface to display controls, the cloud platform can display the simulation interface shown in Figure 11.

[0210] In some alternative implementations, when a user clicks a display control, the cloud platform can also display a next-level menu. For example, in the embodiment shown in Figure 13, when a user clicks a display control on the data stream interface, the cloud platform displays a perspective selection menu, allowing the user to select to display a first-view and / or third-view data stream.

[0211] For example, in the embodiment shown in Figure 13, the cloud platform displays a video stream from a third-person perspective, and the state of each intelligent agent can also be marked in the video stream.

[0212] In some alternative implementations, the cloud platform can also display an introductory interface for the agents, which showcases the capabilities of each agent.

[0213] For example, please refer to Figure 14, which is a schematic diagram of the interface provided in an embodiment of this application.

[0214] As shown in Figure 14, intelligent agents include running intelligent agents, development intelligent agents, and system users. Running intelligent agents assist in the operation of intelligent agents and can specifically be management intelligent agents, intelligent agents A used for transporting goods, intelligent agents B used for clearing obstacles, etc., used in actual operation. Development intelligent agents assist in the development of intelligent agents and can specifically be simulation intelligent agents, code inspection intelligent agents, intelligent agents that modify the simulation environment, etc.

[0215] As shown in Figure 14, the computer device can display the agent's functions behind the agent's avatar. In practical applications, the agent's functions can also be displayed in other ways.

[0216] Optionally, the agent's functionality can be displayed when the input device moves to a specific area, which includes an area for displaying the agent's avatar.

[0217] Alternatively, the agent's functions can be displayed through specific actions, such as right-clicking or double-clicking the agent's avatar.

[0218] Optionally, the agent's functionality can also be displayed in the conversation interface. For example, a user can @ an agent to trigger a function description. Or, a user can @ an agent and ask it to describe its functions.

[0219] It should be noted that the interfaces shown in the foregoing embodiments are merely examples and do not limit the actual interfaces displayed in the interaction methods of the embodied intelligent agents provided in the embodiments of this application.

[0220] Please refer to Figure 15 below, which is a schematic diagram of the structure of the cloud platform provided in an embodiment of this application. The cloud platform runs an intelligent agent system, which includes multiple intelligent agents.

[0221] In some alternative implementations, the cloud platform 1500 includes:

[0222] Display unit 1501 is used to display a session interface, which includes an input field.

[0223] The transceiver unit 1502 is used to respond to a first input operation on the input field and obtain a first instruction, wherein the language of the first instruction is natural language.

[0224] The display unit 1501 is also used to display the first response information of the first agent's reply on the conversation interface. The first response information is for the first instruction, and the first agent is determined based on the first instruction.

[0225] In some optional implementations, the first intelligent agent is a real embodied intelligent agent. If the executing subject of the first instruction is the first intelligent agent, and the event included in the first instruction matches the function of the first intelligent agent, then the first response information includes at least one of the following: the state of the first intelligent agent, and the event state corresponding to the first instruction.

[0226] In some optional implementations, the first intelligent agent is a real embodied intelligent agent. If the executing entity included in the first instruction is the first intelligent agent, and the event included in the first instruction does not match the function of the first intelligent agent, then the first response information includes: invocation information for the second intelligent agent, wherein the second intelligent agent is a real embodied intelligent agent, and the function of the second intelligent agent matches the event. Alternatively, the second intelligent agent is a first AI intelligent agent, which is used to manage the intelligent agent system.

[0227] In some alternative implementations, the display unit 1501 is further configured to display second response information of the second agent's reply on the conversation interface, the second response information being directed to the first instruction.

[0228] In some optional implementations, if the first instruction does not include an executing entity, then the first intelligent agent is a first AI intelligent agent, and the first response information includes: invocation information for a third intelligent agent, wherein the function of the third intelligent agent matches the event included in the first instruction.

[0229] In some alternative implementations, where the first instruction includes an executing entity but not an event, the first agent is the agent instructed by the executing entity. The first response information includes a prompt message for reminding the user to input an event, or for displaying the default function of the first agent, or for displaying the most recent operation of the first agent.

[0230] In some optional implementations, the transceiver unit 1502 is further configured to respond to a second input operation on the input field, acquire a second instruction, the second instruction instructing the display of a data stream of a target event, the target event being an event that the intelligent agent system has completed or is in progress, and the language of the second instruction is natural language. The display unit 1501 is further configured to display the data stream of the target event.

[0231] In some alternative implementations, the display unit 1501 is further configured to display third response information on the session interface, the third response information being in response to the second instruction. A touch operation in response to the third response information displays a data stream of the target event.

[0232] In some alternative implementations, the display unit 1501 is specifically used to display the data stream of the target event on the session interface, or to display the data stream of the target event on an interface other than the session interface.

[0233] In some optional implementations, the data stream of the target event includes: a real-time data stream of the target event, a historical data stream of the target event, a first-person perspective data stream of the target event, or a third-person perspective data stream of the target event. The first-person perspective is the perspective of the agent executing the target event, and the third-person perspective is the perspective of the agent system.

[0234] In some optional implementations, the transceiver unit 1502 is further configured to respond to a third input operation on the input field, acquire a third instruction, the language of which is natural language, and the third instruction instructs the setup of the target simulation environment. The display unit 1501 is further configured to display the target simulation environment on the simulation interface based on the third instruction.

[0235] In some optional implementations, the transceiver unit 1502 is further configured to respond to a fourth input operation on the input field, acquire a fourth instruction, the language of which is natural language, and the fourth instruction indicates the generation of data of the target type. The display unit 1501 is further configured to display a fourth response message on the session interface based on the fourth instruction, the fourth response message indicating data of the target type.

[0236] In some optional implementations, the transceiver unit 1502 is further configured to respond to a fifth input operation on the input field, acquire a fifth instruction, wherein the language of the fifth instruction is natural language, and the events included in the fifth instruction match the functions of the virtual avatar intelligent agent. The display unit 1501 is further configured to display fourth response information replied by the virtual avatar intelligent agent in the conversation interface, wherein the fourth response information includes the state of the virtual avatar intelligent agent, and / or the event state corresponding to the fourth instruction.

[0237] In some optional implementations, the transceiver unit 1502 is further configured to respond to a fifth input operation on the input field, acquire a fifth instruction, the language of which is natural language, and instruct the generation of skill code for a fourth intelligent agent. The display unit 1501 is further configured to display the skill code of the fourth intelligent agent on the code interface based on the fifth instruction.

[0238] In some optional implementations, the processing unit 1503 is further configured to test the skill code of the fourth agent. The display unit 1501 is further configured to display the test results of the skill code of the fourth agent.

[0239] The display unit 1501, transceiver unit 1502, and processing unit 1503 can all be implemented in software or in hardware. For example, the implementation of processing unit 1503 will be described below. Similarly, the implementation of acquisition unit 801 can be referenced to the implementation of processing unit 1503.

[0240] As an example of a software functional unit, processing unit 1503 may include code running on a computing instance. The computing instance may include at least one of a physical host (computing device), a virtual machine, or a container. Further, the aforementioned computing instance may be one or more. For example, processing unit 1503 may include code running on multiple hosts / virtual machines / containers. It should be noted that the multiple hosts / virtual machines / containers used to run the code may be distributed in the same region or in different regions. Further, the multiple hosts / virtual machines / containers used to run the code may be distributed in the same availability zone (AZ) or in different AZs, each AZ including one or more geographically proximate data centers. Typically, a region may include multiple AZs.

[0241] Similarly, multiple hosts / virtual machines / containers used to run this code can be distributed within the same Virtual Private Cloud (VPC) or across multiple VPCs. Typically, a VPC is set up within a region. Communication between two VPCs within the same region, as well as between VPCs in different regions, requires a communication gateway to be set up within each VPC to enable interconnection between VPCs.

[0242] As an example of a hardware functional unit, the processing unit 1503 may include at least one computing device, such as a server. Alternatively, the processing unit 1503 may also be a device implemented using an application-specific integrated circuit (ASIC) or a programmable logic device (PLD). The PLD may be implemented using a complex programmable logical device (CPLD), a field-programmable gate array (FPGA), generic array logic (GAL), or any combination thereof.

[0243] The processing unit 1503 includes multiple computing devices that can be distributed in the same region or in different regions. Similarly, the processing unit 1503 can be distributed in the same Availability Zone (AZ) or in different AZs. Likewise, the processing unit 1503 can be distributed in the same Virtual Private Cloud (VPC) or in multiple VPCs. These multiple computing devices can be any combination of computing devices such as servers, ASICs, PLDs, CPLDs, FPGAs, and GALs.

[0244] It should be noted that the display unit 1501, transceiver unit 1502, and processing unit 1503 respectively implement different steps in the information processing method of the intelligent agent to realize all the functions of the cloud platform 1500. The cloud platform 1500 is used to implement the operations performed by the computer device in the embodiments shown in Figures 1 to 14 above, so as to realize the interaction method of the intelligent agent provided in the embodiments of this application, which will not be described again here.

[0245] Please refer to Figure 16, which is a schematic diagram of a computing device provided in an embodiment of this application. The computing device 1600 includes a processor 1601, a communication interface 1602, a bus 1603, and a memory 1604. The processor 1601, communication interface 1602, and memory 1604 communicate via the bus 1603. In practical applications, communication can also be achieved through other means such as wireless transmission; specific methods are not limited here.

[0246] The computing device 1600 may be a server or a terminal device. It should be understood that this application does not limit the number of processors and memory in the computing device 1600.

[0247] Processor 1601 may include any one or more processors such as a central processing unit (CPU), a graphics processing unit (GPU), a microprocessor (MP), or a digital signal processor (DSP).

[0248] The communication interface 1602 uses transceiver modules such as, but not limited to, network interface cards and transceivers to enable communication between the computing device 1600 and other devices or communication networks.

[0249] Bus 1603 can be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, etc. Buses can be categorized as address buses, data buses, control buses, etc. For ease of illustration, only one line is used in Figure 16, but this does not imply that there is only one bus or one type of bus. Bus 1603 can include pathways for transmitting information between various components of computing device 1600 (e.g., memory 1604, processor 1601, communication interface 1602).

[0250] Memory 1604 may include volatile memory, such as random access memory (RAM). Memory 1604 may also include non-volatile memory, such as read-only memory (ROM), flash memory, hard disk drive (HDD), or solid state drive (SSD).

[0251] The memory 1604 stores executable program code, and the processor 1601 executes the executable program code to implement the functions of the aforementioned display unit 1501, transceiver unit 1502, and processing unit 1503, thereby realizing the data processing method of the intelligent agent. That is, the memory 1604 stores instructions for executing the data processing method of the intelligent agent.

[0252] This application also provides a computing device cluster, which includes at least one computing device. The computing device can be a server, such as a central server, an edge server, or a local server in a local data center. In some optional embodiments, the computing device can also be a terminal device such as a desktop computer, a laptop computer, or a smartphone.

[0253] Please refer to Figures 17 and 18, which are schematic diagrams of the structure of the computing device cluster provided in the embodiments of this application.

[0254] As shown in Figure 17, the computing device cluster includes at least one computing device 1600. The memory 1604 of one or more computing devices 1600 in the computing device cluster may store the same instructions for executing the data processing method for the intelligent agent provided in the embodiments of this application.

[0255] In some possible implementations, the memory 1604 of one or more computing devices 1600 in the computing device cluster may also store partial instructions for executing the data processing method of the intelligent agent. In other words, a combination of one or more computing devices 1600 can jointly execute the instructions for executing the data processing method of the intelligent agent.

[0256] It should be noted that the memory 1604 in different computing devices 1600 within the computing device cluster can store different instructions, which are used to execute certain functions of the cloud platform. That is, the instructions stored in the memory 1604 of different computing devices 1600 can implement the functions of one or more units among the display unit 1501, transceiver unit 1502, and processing unit 1503.

[0257] In some possible implementations, one or more computing devices in a computing device cluster can be connected via a network. This network can be a wide area network (WAN) or a local area network (LAN), etc. Figure 18 illustrates one possible implementation. As shown in Figure 18, two computing devices 1600A and 1600B are connected via a network. Specifically, they are connected to the network through communication interfaces in each computing device. In this type of possible implementation, the memory 1604 in computing device 1600A stores instructions for performing the functions of the display unit 1501. Simultaneously, the memory 1604 in computing device 1600B stores instructions for performing the functions of the transceiver unit 1502 and the processing unit 1503.

[0258] The connection method between the computing device clusters shown in Figure 18 can be based on the fact that in the data processing method of the intelligent agent provided in this application, the display operation and other operations are executed separately. That is, the function of the display unit 1501 is assigned to the computing device 1600A, and the functions of the transceiver unit 1502 and the processing unit 1503 are assigned to the computing device 1600B.

[0259] It should be understood that the functions of computing device 1600A shown in Figure 18 can also be performed by multiple computing devices 1600. Similarly, the functions of computing device 1600B can also be performed by multiple computing devices 1600.

[0260] This application also provides another computing device cluster. The connection relationship between the computing devices in this computing device cluster can be similarly referred to the connection method of the computing device cluster described in Figures 17 and 18, and will not be repeated here.

[0261] This application also provides a computer program product containing instructions. The computer program product may be a software or program product containing instructions, capable of running on a computing device or stored on any usable medium. When the computer program product is run on at least one computer device, it causes the at least one computer device to perform the data processing method of the intelligent agent described above.

[0262] This application also provides a computer-readable storage medium. The computer-readable storage medium can be any available medium that a computing device can store, or a data storage device such as a data center containing one or more available media. The available medium can be a magnetic medium (e.g., floppy disk, hard disk, magnetic tape), an optical medium (e.g., DVD), or a semiconductor medium (e.g., solid-state drive). The computer-readable storage medium includes instructions that instruct the computing device to execute the data processing method of the aforementioned intelligent agent.

[0263] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working processes of the systems, devices, and units described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here.

[0264] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit them. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the protection scope of the technical solutions of the embodiments of this application.

Claims

1. An information processing method for an intelligent agent, characterized in that, The method is applied to a cloud platform, which runs an intelligent agent system, the intelligent agent system comprising multiple intelligent agents, and the method includes: Display a session interface, which includes an input field; In response to a first input operation on the input field, a first instruction is obtained, wherein the language of the first instruction is natural language; The first response information of the first agent is displayed in the conversation interface. The first response information is in response to the first instruction, and the first agent is determined based on the first instruction.

2. The method according to claim 1, characterized in that, The first intelligent agent is a real embodied intelligent agent; If the executing entity included in the first instruction is the first intelligent agent, and the event included in the first instruction matches the function of the first intelligent agent, then the first response information includes: the state of the first intelligent agent, and / or the event state corresponding to the first instruction.

3. The method according to claim 1, characterized in that, The first intelligent agent is a real embodied intelligent agent; If the executing entity included in the first instruction is the first intelligent agent, and the event included in the first instruction does not match the function of the first intelligent agent, then the first response information includes: Regarding the invocation information of the second intelligent agent, the second intelligent agent is a real embodied intelligent agent, and the function of the second intelligent agent matches the event; or, the second intelligent agent is a first artificial intelligence (AI) intelligent agent, which is used to manage the intelligent agent system.

4. The method according to claim 3, characterized in that, After the first response information of the first agent is displayed on the conversation interface, the method further includes: The second response information of the second agent is displayed in the conversation interface, and the second response information is in response to the first instruction.

5. The method according to claim 1, characterized in that, If the first instruction does not include an executing entity, then the first intelligent agent is a first AI intelligent agent, and the first response information includes: Regarding the invocation information of the third intelligent agent, the function of the third intelligent agent matches the events included in the first instruction.

6. The method according to any one of claims 1 to 5, characterized in that, The method further includes: In response to a second input operation on the input field, a second instruction is obtained, the second instruction instructing the display of a data stream of a target event, the target event being an event that the intelligent agent system has completed or is in the process of executing, and the language of the second instruction is natural language; Displays the data stream of the target event.

7. The method according to claim 6, characterized in that, Prior to displaying the data stream of the target event, the method further includes: The third response information is displayed on the session interface, and the third response information is in response to the second instruction; The data stream displaying the target event includes: In response to a touch operation on the third response information, a data stream of the target event is displayed.

8. The method according to claim 6 or 7, characterized in that, The data stream for displaying the target event includes: The data stream of the target event is displayed in the session interface, or the data stream of the target event is displayed in an interface other than the session interface.

9. The method according to any one of claims 6 to 8, characterized in that, The data stream of the target event includes: the real-time data stream of the target event, the historical data stream of the target event, the first-person perspective data stream of the target event, or the third-person perspective data stream of the target event; Wherein, the first perspective is the perspective of the agent executing the target event, and the third perspective is the perspective of the agent system.

10. The method according to any one of claims 1 to 9, characterized in that, The method further includes: In response to a third input operation on the input field, a third instruction is obtained, wherein the language of the third instruction is natural language and the third instruction instructs the construction of a target simulation environment; Based on the third instruction, the target simulation environment is displayed on the simulation interface.

11. The method according to claim 10, characterized in that, The method further includes: In response to a fourth input operation on the input field, a fourth instruction is obtained, wherein the language of the fourth instruction is natural language and the fourth instruction instructs the generation of data of the target type; Based on the fourth instruction, a fourth response message is displayed on the session interface, the fourth response message indicating the data of the target type.

12. The method according to claim 10 or 11, characterized in that, The method further includes: In response to a fifth input operation on the input field, a fifth instruction is obtained, wherein the language of the fifth instruction is natural language, and the events included in the fifth instruction match the functions of the virtual embodied intelligent agent. The fourth response information is displayed in the conversation interface, which includes the state of the virtual avatar and / or the event state corresponding to the fourth instruction.

13. The method according to any one of claims 1 to 12, characterized in that, The method further includes: In response to a sixth input operation on the input field, a sixth instruction is obtained, wherein the language of the sixth instruction is natural language and the sixth instruction indicates the generation of skill code for a fourth agent; Based on the sixth instruction, the skill code of the fourth agent is displayed in the code interface.

14. The method according to claim 13, characterized in that, The method further includes: Test the skill code of the fourth embodied intelligent agent; The test results of the skill codes of the fourth embodied intelligent agent are displayed.

15. An information processing device for an intelligent agent, characterized in that, Includes a module for implementing the method according to any one of claims 1 to 14.

16. A computer equipment cluster, characterized in that, It includes at least one computing device, each computing device including a processor and memory; The processor of the at least one computing device is configured to execute instructions stored in the memory of the at least one computing device to cause the cluster of computing devices to perform the method as described in any one of claims 1 to 14.

17. A computer-readable storage medium, characterized in that, The computer-readable storage medium includes computer program instructions that, when executed by a cluster of computing devices, cause the cluster of computing devices to perform the method as described in any one of claims 1 to 14.

18. A computer program product containing instructions, characterized in that, When the instructions are executed by a cluster of computer devices, the cluster of computer devices causes the cluster of computer devices to perform the method as described in any one of claims 1 to 14.