Information processing method for agent, and related device
Through the cloud platform's intelligent body system, natural language interacts with the intelligent body to simplify the development process of the embodied intelligent body, solve the complex and lengthy problems of traditional web services, and improve the development efficiency and intelligence of the intelligent body.
Patent Information
- Application Number
- PCT/CN2024/110316
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-05-22
- Filing Date
- 2024-08-07
- Publication Date
- 2025-08-28
AI Technical Summary
Traditional web services require multiple page jumps to complete different needs, resulting in complex and lengthy services and low efficiency in development of embodied intelligent bodies.
Through the cloud platform's intelligent body system, natural language interacts with the intelligent body to simplify the development process of the embodied intelligent body. The intelligent body is called and managed through natural language to achieve automatic matching and scheduling of events.
It reduces the complexity of development of embodied intelligent bodies, improves development efficiency and intelligence of the intelligent bodies, simplifies user operations, and enriches implementation methods and application scenarios.
Smart Images

Figure CN2024110316_28082025_PF_FP_ABST
Abstract
Description
An intelligent information processing method and related equipment
[0001] This application claims priority to Chinese patent application No. 202410204964.1, filed with the State Intellectual Property Office on February 23, 2024, entitled “A Development and Operation Management System for Embodied Intelligent Agents Based on Multi-Agent Interaction,” and to Chinese patent application No. 202410660467.2, filed with the State Intellectual Property Office on May 22, 2024, entitled “An Information Processing Method for Intelligent Agents and Related Devices.” The entire contents of each application are incorporated herein by reference. Technical Field
[0002] The present application relates to the field of cloud computing, and in particular to an information processing method for an intelligent entity and related equipment. Background Art
[0003] With the development of artificial intelligence (AI), AI agents are impacting all industries. An agent can be understood as an intelligent agent that can complete tasks on behalf of users. This includes embodied agents, which interact with or provide feedback to physical entities. These agents include virtual objects, digital entities, and robots, and can be represented graphically.
[0004] Traditional web services require users to navigate through multiple pages to complete different needs and tasks, resulting in inefficient service delivery. Furthermore, existing embodied intelligent development solutions require the coordination of multiple modules. Development based on these traditional service models results in a complex and lengthy process, resulting in low development efficiency.
[0005] Summary of the Invention
[0006] This application provides an intelligent agent information processing method and related equipment for simplifying the process of embodied intelligent agent development and improving development efficiency.
[0007] In a first aspect, the present application provides an information processing method of an intelligent agent, comprising:
[0008] The information processing method for an intelligent agent provided in this application is applied to a cloud platform, which runs an intelligent agent system, and the intelligent agent system includes multiple intelligent agents. These multiple intelligent agents include at least one of a real embodied intelligent agent, a virtual embodied intelligent agent, and an AI intelligent agent. The real embodied intelligent agent corresponds to a physical entity in a physical environment, and the virtual embodied intelligent agent corresponds to a virtual entity in a virtual environment. Both the real embodied intelligent agent and the virtual embodied intelligent agent can interact with or act on physical entities, and will have an impact on the physical environment. The AI intelligent agent is used to manage other intelligent agents in the intelligent agent system, or to provide online services.
[0009] The cloud platform displays a conversation interface through a computer device that is logged into a user's account registered on the cloud platform and provides agent services. The conversation interface includes an input field. The cloud platform responds to operations on the input field and obtains instructions input by the user. Exemplarily, the cloud platform responds to a first input operation on the input field and obtains a first instruction, where the language of the first instruction is a natural language. The first input operation can be a touch operation or input via an input device. In addition, other input methods such as voice input are also possible, and the specifics are not limited here. Input devices include keyboards, mice, styluses, and other devices, and the specifics are not limited here. Furthermore, the term "natural language" in this application refers to the language used for human communication, including text and voice, and this application does not limit the language of natural language. The cloud platform obtains the first instruction and displays a first response message from the first agent on the conversation interface. The first response message is specific to the first instruction, that is, the first response message is a response to the first instruction. The language of the first response message is a natural language or a code language, and the specifics are not limited here. The first agent is determined based on the first instruction, and the agent system includes multiple agents, including the first agent. Specifically, the first instruction includes an execution subject and / or an event. The execution subject indicates the object that executes the first instruction, or in other words, the execution subject is the subject of the first instruction. The event indicates the purpose of the first instruction, or the intention of the first instruction, and represents what the first instruction is intended to accomplish or achieve.
[0010] In this application, natural language interaction with a first agent simplifies complex development or management processes. Furthermore, in scenarios where the first agent includes an embodied agent or the first agent manages an embodied agent, the complexity of embodied agent development is reduced and development efficiency is improved. In some optional implementations of the first aspect, the first agent is a real embodied agent, meaning that the first agent corresponds to a physical entity in a physical environment. If the executor of the first instruction is the first agent, and the event included in the first instruction matches the function of the first agent, then the first agent can complete the event indicated by the first instruction. The first response information then includes at least one of the status of the first agent and the status of the event corresponding to the first instruction. The status of the first agent includes information reflecting the current status of the first agent, such as whether the first agent accepted the first instruction, whether the first agent currently has a task to execute, and whether the first agent is faulty, offline, or low on battery. The specific content of the event status corresponding to the first instruction varies depending on the event included in the first instruction. In general, it can include event-related information such as the completion degree of the event at the current moment, the stage of the event at the current moment, the estimated time for completion of the event, whether obstacles were encountered during event processing, etc. The language of the first response information is natural language.
[0011] In this application, if the first agent is a real embodied agent, the first instruction instructs the first agent, and the first agent can complete the event included in the first instruction, the first response information may have multiple possibilities, enriching the implementation methods of the technical solution of this application. In addition, based on the first response information, the status of the first agent and / or the event included in the first instruction can be understood, and it can be determined whether to adjust the agent executing the first instruction, thereby accelerating the completion of the event included in the first instruction and improving the efficiency of the event completion.
[0012] In some optional implementations of the first aspect, the first agent is a real embodied agent, corresponding to a physical entity in a physical environment. When the execution subject included in the first instruction is the first agent, and the event included in the first instruction does not match the function of the first agent, it means that the first agent cannot complete the event included in the first instruction, then the first agent will call other agents to complete the event included in the first instruction. That is to say, the first response information includes call information for the embodied agent, the second agent is a real embodied agent, and the function of the second agent matches the event included in the first instruction, or the second agent is a first AI agent, and the first AI agent is used to manage the agent system. Among them, the call information for the second agent refers to calling the second agent to implement the event included in the first instruction, or to reply to the first instruction. In addition, the language of the first response information can be a natural language, that is, the language of the call information for the second agent can be a natural language.
[0013] In this application, if a first agent, instructed by a first instruction, fails to complete the event specified in the first instruction, the first response message sent by the first agent is actually a call message to a second agent, which then completes the event specified in the first instruction. In other words, in this application, agents can call each other via natural language, further improving agent development efficiency and enhancing the agent's intelligence.
[0014] In some optional implementations of the first aspect, in scenarios where the first response message includes call information for a second agent, after the cloud platform displays the first response message on the conversation interface, it also displays a second response message from the second agent in response to the first instruction. That is, the second response message is actually the second agent's reply to the content of the first instruction. Furthermore, the second response message varies depending on the second agent. Optionally, in scenarios where the function of the second agent matches the event included in the first instruction, the second response message includes the status of the second agent and / or the status of the event corresponding to the first instruction. In scenarios where the second agent is a first AI agent, the second response message includes call information for an agent that matches the event included in the first instruction, or an "unable to execute" message. The "unable to execute" message indicates that the agent system cannot execute the event included in the first instruction. In general, in scenarios where the second agent is a first AI agent, the first AI agent sends a call message to another agent, instructing the agent to implement the event included in the first instruction, or the first AI agent replies with an "unable to execute" message.
[0015] In this application, in a scenario where a first agent, instructed by a first instruction, fails to complete the event specified in the first instruction, a second agent can also reply with a second response message after the first agent replies with a first response message. This allows for multiple possibilities for both the second agent and the second response message, enriching the implementation methods and application scenarios of this technical solution and enhancing its flexibility and practicality.
[0016] In some optional implementations of the first aspect, in the solution where the first instruction does not include an execution subject, the first agent is defaulted to the first AI agent, which may also be referred to as a management agent, for managing the agent system. Then the first response information includes call information for a third agent, or an inability to execute information. The function of the third agent matches the event included in the first instruction. The call information for the third agent refers to calling the third agent to implement the event included in the first instruction. The inability to execute information indicates that the agent system cannot execute the event included in the first instruction.
[0017] In this application, in scenarios where the first instruction does not include an execution subject, the managing first AI agent defaults to replying with a first response message. This means that even if the user enters an incomplete first instruction, a reply can still be received, simplifying user operations and further simplifying agent development. Furthermore, the first AI agent can call on other agents to complete the event included in the first instruction, or reply with a message indicating that the execution could not be performed, allowing the user to understand the progress of the event included in the first instruction, thereby enhancing the practicality of the technical solution of this application.
[0018] In some optional implementations of the first aspect, in scenarios where the first instruction includes an execution subject but not an event, the first agent is the agent indicated by the execution subject. The first response information includes prompt information, which is used to remind the user to input the event, or to display the default function of the first agent, or to display the first agent's most recent operation. Alternatively, the first response information is used to trigger the first agent's default operation, such as triggering the first agent's default function or triggering the first agent's most recent operation. The specific content of the default operation can be set based on actual application needs and is not specifically limited here.
[0019] In some optional implementations of the first aspect, the cloud platform receives a second instruction in response to a second input operation on the input field. The second instruction instructs the display of a data stream for a target event, or in other words, the event included in the second instruction is the data stream for displaying the target event. The target event is an event that has been executed or is currently being executed by the agent system. The second instruction is expressed in natural language. The second input operation is similar to the first input operation and has multiple possibilities, which are not further described here. After receiving the second instruction, the data stream for the target event is displayed.
[0020] In this application, the data stream of the target event can vividly display the status of the target event, allowing users to intuitively feel the status of the target event, thereby improving the user experience.
[0021] In some optional implementations of the first aspect, before displaying the data stream of the target event, the cloud platform displays third response information on the session interface. The third response information is specific to the second instruction, that is, the third response information is a response to the second instruction. Therefore, the computer device may display the data stream of the target event in response to a touch operation on the third response information. The third response information may be a medium such as a link or file indicating the data stream of the target event. The user touches the medium to display the data stream of the target event.
[0022] In this application, after the cloud platform obtains the second instruction, it can not only directly display the data stream of the target event, but also display the data stream of the target event by displaying the third response information on the conversation interface, enriching the implementation method and application scenarios of the technical solution of this application and improving the flexibility of the technical solution.
[0023] In some optional implementations of the first aspect, the target event data stream may be displayed in a variety of locations, including within the conversation interface or elsewhere, without limitation. This configuration can be flexibly tailored to the specific application, further enhancing the flexibility of the technical solution of this application.
[0024] In some optional implementations of the first aspect, the target event data stream may include a variety of possibilities, including a real-time data stream of the target event, a historical data stream of the target event, a first-person perspective data stream of the target event, or a third-person perspective data stream of the target event, without limitation herein. The first-person perspective is the perspective of the agent executing the target event, while the third-person perspective, also known as a "God's eye view" or "bystander's eye view," is the perspective of the agent system.
[0025] In the present application, there are multiple possible types of data streams in the target perspective, which can be determined or set by default based on the content indicated by the second instruction, further enhancing the flexibility of the technical solution of the present application.
[0026] In some optional implementations of the first aspect, the cloud platform obtains a third instruction in response to a third input operation on the input field. The third instruction instructs to build a target simulation environment. That is, the event included in the third instruction is to build the target simulation environment. The third instruction is written in a natural language. Based on the third instruction, the computer device displays the target simulation environment on the simulation interface.
[0027] In this application, the computer device can also display a simulation environment, through which the simulation or simulation of physical entities can be realized, and the embodied intelligent body can be developed and tested.
[0028] In some optional implementations of the first aspect, the cloud platform responds to a fourth input operation on the input bar and obtains a fourth instruction. The fourth instruction indicates the generation of target type data, that is, the event included in the fourth instruction is the generation of target type data. In addition, the language of the fourth instruction is a natural language. There are many possible target types, including radar (lidar) data, depth data, red, yellow and blue color mode (RGB) data, trajectory data, etc., which are not specifically limited here. Based on the fourth instruction, the cloud platform displays a fourth response message on the conversation interface, and the fourth response message includes data indicating the target type. Further, the fourth response message may include target type data, or include a medium for data indicating the target type. In the latter solution, the computer device responds to a touch operation on the medium and displays the target type data.
[0029] In the present application, the data type generated by the simulation environment can also be indicated. There are multiple possibilities for the data type and the fourth response information, which enriches the implementation method of the technical solution of the present application.
[0030] In some optional implementations of the first aspect, the cloud platform receives a fifth instruction in response to a fifth input operation on the input field. The fifth instruction is in a natural language, and the event included in the fifth instruction matches the function of the virtual embodied intelligent agent. The cloud platform displays fourth response information provided by the virtual embodied intelligent agent on the conversation interface. The fourth response information includes the status of the virtual embodied intelligent agent and / or the status of the event corresponding to the fourth instruction.
[0031] In this application, a virtual embodied agent can be called upon to complete instructions, thereby changing the virtual environment. This can also serve as a simulation of changes in the real physical environment, providing a reference for applying corresponding physical entities in the physical environment, or providing simulation data, thereby improving the practicality of the technical solution of this application.
[0032] In some optional implementations of the first aspect, in a scenario where the execution subject included in the fifth instruction is not the virtual embodied intelligent body shown in the previous implementation, before displaying the fourth response information, the cloud platform displays the response information replied by the execution subject included in the fifth instruction, and the response information includes call information for the virtual embodied intelligent body, which is used to instruct the virtual embodied intelligent body to reply to the fifth instruction.
[0033] In some optional implementations of the first aspect, the cloud platform responds to a sixth input operation on the input bar and obtains a sixth instruction. The sixth instruction instructs the generation of a skill code for the fourth agent. That is, the event included in the sixth instruction is the generation of a skill code for the fourth agent. The skill code is used to describe the functions possessed by the fourth agent and can also describe the working logic of the fourth agent. In addition, the language of the fourth instruction is a natural language. Based on the sixth instruction, the cloud platform displays the skill code of the fourth agent on the code interface.
[0034] In this application, the computer device can also generate and display the skill code of the intelligent agent, providing a basis for code detection and improving the feasibility of the technical solution.
[0035] In some optional implementations of the first aspect, the cloud platform can also test the skill code of the fourth intelligent agent and display the test results of the code, so that users can judge whether the code can be executed accurately, thereby improving the accuracy of intelligent agent development.
[0036] In a second aspect, the present application provides an information processing device for an intelligent agent, which can implement the method described in the first aspect, or any possible implementation of the first aspect. The device includes corresponding units or modules for executing the above-mentioned method. The units or modules included in the device can be implemented in software and / or hardware.
[0037] In a third aspect, the present application provides a computer device comprising a processor and a memory, wherein the processor stores instructions. When the instructions stored in the memory are executed on the processor, the method shown in the aforementioned first aspect or any possible implementation of the first aspect is implemented.
[0038] In a fourth aspect, the present application provides a computer-readable storage medium, which stores instructions. When the instructions are executed on a processor, the method shown in the first aspect or any possible implementation of the first aspect is implemented.
[0039] In a fifth aspect, the present application provides a computer program product, which, when executed on a processor, implements the method shown in the aforementioned first aspect or any possible implementation of the first aspect.
[0040] The beneficial effects shown in any of the second to fifth aspects are similar to those of the first aspect or any possible implementation method of the first aspect, and will not be repeated here. BRIEF DESCRIPTION OF THE DRAWINGS
[0041] FIG1 is a schematic diagram of a system architecture provided in an embodiment of the present application;
[0042] FIG2 is a flow chart of an information processing method for an intelligent agent provided in an embodiment of the present application;
[0043] FIG3 is a schematic diagram of a conversation interface provided in an embodiment of the present application;
[0044] FIG4 is a schematic diagram of an input field provided in an embodiment of the present application;
[0045] FIG5 is another schematic diagram of an input field provided in an embodiment of the present application;
[0046] FIG6 is another schematic diagram of a conversation interface provided in an embodiment of the present application;
[0047] FIG7 is another schematic diagram of a conversation interface provided in an embodiment of the present application;
[0048] FIG8 is a schematic diagram of an interface provided in an embodiment of the present application;
[0049] FIG9 is another schematic diagram of an interface provided in an embodiment of the present application;
[0050] FIG10 is another schematic diagram of a conversation interface provided in an embodiment of the present application;
[0051] FIG11 is a schematic diagram of a simulation interface provided in an embodiment of the present application;
[0052] FIG12 is a schematic diagram of a code interface provided in an embodiment of the present application;
[0053] FIG13 is another schematic diagram of an interface provided in an embodiment of the present application;
[0054] FIG14 is another schematic diagram of an interface provided in an embodiment of the present application;
[0055] FIG15 is a schematic diagram of the structure of a cloud platform provided in an embodiment of the present application;
[0056] FIG16 is a schematic diagram of the structure of a computing device provided in an embodiment of the present application;
[0057] FIG17 is a schematic diagram of a structure of a computing device cluster provided in an embodiment of the present application;
[0058] FIG18 is another structural diagram of a computing device cluster provided in an embodiment of the present application. DETAILED DESCRIPTION
[0059] The embodiments of the present application provide an intelligent agent information processing method and related equipment for simplifying the process of embodied intelligent agent development and improving development efficiency.
[0060] The embodiments of the present application are described below in conjunction with the accompanying drawings. Those skilled in the art will appreciate that, with the development of technology and the emergence of new scenarios, the technical solutions provided in the embodiments of the present application are also applicable to similar technical problems.
[0061] The terms "first", "second", etc. in the specification and claims of this application and the above-mentioned drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or sequential order. It should be understood that the terms used in this way are interchangeable when appropriate, and this is merely a way of distinguishing objects of the same attributes when describing the embodiments of the present application. In addition, the terms "including" and "having" and any of their variations are intended to cover non-exclusive inclusions, so that the process, method, system, product or device comprising a series of units need not be limited to those units, but may include other units that are not clearly listed or inherent to these processes, methods, products or devices. In addition, "at least one" refers to one or more, and "a plurality" refers to two or more. "and / or" describes the association relationship of associated objects and indicates that three relationships can exist. For example, A and / or B can represent: the situation where A exists alone, A and B exist simultaneously, and B exists alone, where A and B can be singular or plural. The character " / " generally indicates that the associated objects before and after are in an "or" relationship. "At least one of the following" or similar expressions refers to any combination of these items, including any combination of single or plural items. For example, at least one of a, b, or c can mean: a, b, c, ab, ac, bc, or abc, where a, b, and c can be single or plural.
[0062] First, please refer to FIG1 , which is a schematic diagram of the system architecture provided in an embodiment of the present application.
[0063] As shown in Figure 1, a tenant logs into cloud platform 30 via client 10 over the internet 20, using their registered account and password. Cloud platform 30 manages infrastructure, which includes multiple data centers located in different regions, with each region having at least one cloud data center. For example, Region 1 shown in Figure 1 includes Cloud Data Center 1 and Cloud Data Center 2, while Region 2 includes Cloud Data Center 3 and Cloud Data Center 4. Each cloud data center is equipped with multiple servers, each of which runs business instances (including at least one of virtual machines, containers, and dedicated hosts).
[0064] Tenants, also known as users, are top-level objects used to manage cloud services and / or cloud resources. Tenants register a tenant account and set a tenant password on the cloud platform 30 using a local client (e.g., a browser). The local client then remotely logs in to the cloud platform 30 using the tenant account and password. The cloud platform 30 provides a configuration interface or API for tenants to configure and use cloud services.
[0065] In an embodiment of the present application, the cloud platform 30 runs an intelligent agent system, which includes multiple intelligent agents. These multiple embodied intelligent agents include at least one of a real embodied intelligent agent, a virtual embodied intelligent agent, and an AI intelligent agent.
[0066] A real embodied agent corresponds to a physical entity in a physical environment, performs corresponding operations based on instructions, interacts with the physical entity, or changes the state of the physical entity. A virtual embodied agent corresponds to a virtual embodied agent in a virtual environment, performs corresponding operations based on instructions, can change the virtual environment, and thus act on the physical entity. In general, whether it is a real embodied agent or a virtual embodied agent, it refers to an agent that interacts with a physical entity or can act on a physical entity. On the client 10, the embodied agent can be represented graphically, that is, the embodied agent is actually an entity represented graphically. Exemplarily, the embodied agent can also be a virtual object, a digital entity, a robot, or other entity, which is not limited here.
[0067] AI agents are used to manage other agents in the agent system or provide online services. They are used to assist in the development or operation of embodied agents.
[0068] In the embodiment of the present application, the user develops or manages the embodied intelligent agent included in the intelligent agent system running on the cloud platform 30 through the operation terminal 10. The specific implementation process is described in detail below.
[0069] In some optional implementations, the physical environment corresponding to the system architecture provided in the embodiments of the present application may be logistics, port freight, factory areas, homes, offices, and other scenarios, and the simulation environment may be a simulation of the aforementioned physical environment, which is not specifically limited here.
[0070] It should be noted that, in actual applications, the client 10 may also be other types of devices, such as mobile phones, tablet computers, wearable devices, vehicles, drones, smart home devices, etc. Computer devices can be widely used in various scenarios, such as device-to-device (D2D), vehicle-to-everything (V2X) communication, machine-type communication (MTC), Internet of Things (IoT), virtual reality, augmented reality, industrial control, autonomous driving, telemedicine, smart grid, smart furniture, smart office, smart wearables, smart transportation, smart cities, etc.
[0071] Next, please refer to FIG2 , which is a flow chart of an interactive method for an embodied intelligent agent provided in an embodiment of the present application, including:
[0072] 201. The cloud platform displays a conversation interface, which includes an input field.
[0073] The cloud platform displays a conversational interface through a computer device. This interface is used to display interactions between agents and users, or between agents themselves. These interactions can be expressed through natural language, such as text or voice. In layman's terms, the conversational interface can be understood as a chat interface.
[0074] The conversation interface includes an input field, which can be understood as a window for user interaction with an agent. Users input into the input field to issue commands. In practical applications, the conversation interface and input field can take many forms, as illustrated below using Figure 3 as an example. Figure 3 is a schematic diagram of a conversation interface provided in an embodiment of the present application.
[0075] As shown in FIG3 , conversation interface 300 includes an input field 310 . In the embodiment shown in FIG3 , a user enters the content "Send cargo 1 from port A to port B" in input field 310 . The user clicks the send control 311 , and the content is displayed in the conversation interface as shown in conversation box 330 .
[0076] In addition, as shown in FIG. 3 , the conversation interface 300 may further display conversation subject avatars 320 for distinguishing the senders of each conversation frame 330 .
[0077] It is understood that the input field 310 shown in FIG3 is merely a schematic diagram. In practical applications, the input field 310 may also have other forms. For example, please refer to FIG4 , which is a schematic diagram of the input field provided in an embodiment of the present application.
[0078] As shown in Figure 4 , in addition to the send control 311, the input field 310 may also include other functional controls 312. For example, the functional controls 312 shown in Figure 4 are, from left to right, an attachment control, an email control, and a language control. The attachment control is used to add or modify attachments, the email control is used for email transmission, and the voice control is used for voice input. In actual applications, the input field 310 may also include a greater or lesser number of functional controls, and the specific number is not limited here.
[0079] 202. The cloud platform responds to a first input operation on the input field and obtains a first instruction, where the language of the first instruction is a natural language.
[0080] The user performs a first input operation in the input field, and the cloud platform responds to the operation and obtains a first instruction. In the embodiment of the present application, the language of the first instruction is a natural language, that is, the language used in the content input by the user is a natural language. Natural language includes text and speech, and refers to the language used by humans to communicate. In the embodiment of the present application, the language of the natural language is not limited. For example, it can be Chinese, English, Korean, etc., which is determined based on the needs of the actual application and is not limited here.
[0081] Exemplarily, in the implementation shown in FIG3 , the user enters the content “Send cargo 1 from port A to port B” in the input bar 310 and clicks the send control 311 , so that the cloud platform obtains the first instruction “Send cargo 1 from port A to port B”.
[0082] Optionally, the user can perform the first input operation through an input device, which includes a keyboard, a mouse, a stylus, and other devices, the specifics of which are not limited here. Optionally, if the screen of the computer device with the cloud platform account logged in is a touch screen, the user can also perform the first input operation directly in the input field displayed on the computer device using a stylus or a finger.
[0083] 203. The cloud platform displays the first response information replied by the first agent on the conversation interface. The first response information is for the first instruction, and the first agent is determined based on the first instruction.
[0084] After the cloud platform obtains the first instruction, it displays the first response information replied by the first agent on the conversation interface, that is, the sender of the first response information is the first agent.
[0085] In the embodiments of the present application, the first instruction includes an execution subject and / or an event. The execution subject indicates the object that completes the event. In layman's terms, the execution subject is the subject of the first instruction, indicating who will complete the event included in the first instruction. The event of the first instruction indicates the purpose or intention of the first instruction. In layman's terms, the event represents what the first instruction is intended to accomplish.
[0086] In the example of this application, the first agent is determined based on the first instruction. Since the content of the first instruction has multiple possibilities, the first agent also has multiple possibilities, which are explained below.
[0087] (1) The first instruction includes an execution subject and an event.
[0088] In this solution, it can be understood that the user specifies the execution subject and the event to be completed by the execution subject. The execution subject specified by the user is the first agent. The user can specify the first agent in various ways, such as by using special symbols or by making the first agent the subject of the input content. The specific methods are not limited here.
[0089] For example, please refer to FIG5 , which is a schematic diagram of an input field provided in an embodiment of the present application.
[0090] Optionally, as shown in FIG5(a), the first agent is designated as "Agent A" by the special symbol @. In actual applications, after the user enters the special symbol, the computer device may also display an agent list, showing multiple agents included in the agent list, so that the user can select the first agent from among them.
[0091] In addition, in addition to the @ shown in FIG. 5 , the special symbol may also be other symbols, such as #, *, etc., which can be set based on the needs of actual applications and are not specifically limited here.
[0092] Optionally, as shown in Figure 5 (b), the user inputs content: Agent A delivers cargo 1 from Port A to Port B. The computer device performs semantic recognition and analysis on the content, and determines that the subject of the content is Agent A, thereby determining that the first agent is Agent A.
[0093] In actual applications, the first agent may not be able to complete the event specified by the user. In this case, there are many possibilities for the content indicated by the first response information, which are explained below.
[0094] 1. The first agent is a real embodied agent, and the function of the first agent matches the event.
[0095] The first agent is a real, embodied agent, meaning it corresponds to a physical entity in the physical environment, a real, tangible object such as a robot, robotic arm, camera, or sensor. The first agent's functionality matches the event, meaning the first agent can complete the event. In this scenario, the first response information includes the state of the first agent and / or the state of the event corresponding to the first instruction.
[0096] The status of the first agent includes information that can reflect the current status of the first agent, such as whether the first agent accepts the first instruction, whether the first agent currently has a task to execute, whether the first agent is faulty / offline / low on power, etc.
[0097] The event status corresponding to the first instruction varies depending on the event included in the first instruction. Generally speaking, it can include event-related information such as the completion degree of the event at the current moment, the current stage of the event, the estimated completion time of the event, and whether any obstacles were encountered during event processing.
[0098] For example, let's assume the first instruction is "Agent B delivers a box of dolls from Port C to Port D." That is, the first agent is Agent B, and the event is delivering a box of dolls from Port C to Port D. For different situations, the first response message can have different contents:
[0099] Optionally, if agent B is currently executing the previous task, the first response message of agent B may be "Executing the previous task, please wait." Optionally, if agent B is currently low on power, the first response message of agent B may be "Low power, please select another agent to execute the task." Optionally, if agent B fails, the first response message of agent B may be "Under maintenance, estimated waiting time is 20 minutes." The first response messages of the above examples all reflect the status of the first embodied agent.
[0100] Alternatively, if the road from Port C to Port D is currently blocked, Agent B's first response message could be "Road blocked, please re-route." Alternatively, Agent B's first response message could be "OK, on my way to pick up the goods," "Goods already picked up, shipping now," etc. The first response messages in the aforementioned examples all reflect the event status corresponding to the first instruction.
[0101] Alternatively, if Agent B is idle and the route from Port C to Port D is clear, Agent B, based on the first instruction, delivers the box of dolls from Port C to Port D, thus transferring the box of dolls between ports and changing the physical entities in the physical environment. Alternatively, Agent B can be understood as interacting with the physical entities.
[0102] For example, assume the first instruction is "Agent A takes a photo of Port A," meaning the first agent is Agent A and the event is taking a photo of Port A. Then the first response message could be "Photo of Port A," meaning the event status for the first instruction is Completed.
[0103] Optionally, if the first agent cannot immediately execute the event included in the first instruction, the first response information may also include prompt information, which is used to remind the user to change the agent, or provide the user with other optional embodied agents that can execute the event included in the first instruction.
[0104] It should also be noted that the language of the first response information can be a natural language, a code language, or a custom language, and the specific language is not limited here. For example, in the scenario where the first agent fails, the first response information can be a custom code indicating that the first agent has failed and cannot complete the event included in the first instruction.
[0105] In the embodiments of the present application, if the first agent is a real embodied agent, the first instruction instructs the first agent, and the first agent can complete the event included in the first instruction, the first response information may have multiple possibilities, enriching the implementation methods of the technical solution of the present application. In addition, based on the first response information, the status of the first agent and / or the event included in the first instruction can be understood, and it can be determined whether to adjust the agent executing the first instruction, thereby accelerating the completion of the event included in the first instruction and improving the efficiency of the event completion.
[0106] 2. The first agent is a real embodied agent, and the function of the first agent does not match the event.
[0107] The function of the first agent does not match the event, indicating that the first agent cannot complete the event. Then, the first agent calls the second agent to reply to the first instruction. In other words, the first response information includes call information for the second agent. The second agent is the first AI agent, and the first AI agent is used to manage the agent system. Alternatively, the second agent is a real embodied agent, and the function of the second agent matches the event included in the first instruction. In the embodiment of the present application, the first AI agent is also called the management agent.
[0108] The following is an explanation with reference to a schematic diagram. Please refer to FIG6 , which is a schematic diagram of a conversation interface provided in an embodiment of the present application.
[0109] In the embodiment shown in Figure 6, the first instruction is "Agent C, send 1 box of dolls from Port C to Port D" as an example, that is, the first agent is Agent C, and the event included in the first instruction is sending 1 box of dolls from Port C to Port D as an example.
[0110] Agent C is unable to complete the event included in the first instruction. In the embodiment shown in Figure 6(a), Agent C's first response message is to call the first AI agent to execute the event included in the first instruction. In the embodiment shown in Figure 6(b), Agent C's first response message is to call Agent B to execute the event included in the first instruction. Agent B's function matches the event included in the first instruction.
[0111] It is understandable that the embodiment shown in FIG6 is merely an example of the first response information. In actual applications, the first response information returned by agent C may also not include the event, but only include the execution subject (that is, only agent B). In this solution, agent B can identify the event included in the first instruction based on the conversation interface.
[0112] In addition, the language of the first response information can be a natural language, which means that in the embodiment of the present application, different intelligent agents can also call each other through natural language.
[0113] In this application, if a first agent, instructed by a first instruction, fails to complete the event specified in the first instruction, the first response message sent by the first agent is actually a call message to a second agent, which then completes the event specified in the first instruction. In other words, in this application, agents can call each other via natural language, further simplifying agent development efficiency and improving their intelligence.
[0114] In some optional embodiments, after the conversation interface displays the first response message from the first embodied agent, the computer device may also display a second response message from the second agent on the conversation interface, where the second response message is specific to the first instruction. In other words, the second response message is actually the second agent's reply to the first instruction.
[0115] Optionally, if the second agent is a management agent, the second response information includes call information for an agent matching the event included in the first instruction, or an execution failure information indicating that the agent system cannot execute the event included in the first instruction.
[0116] For example, in the embodiment shown in Figure 6(a), after Agent C replies with a first response message, the management agent replies with a second response message. The second response message indicates an event that calls Agent B to execute the first instruction. Although Figure 6(a) only displays the content of the second response message as "@Agent B," Agent B recognizes the content of the conversation interface and can determine that the event to be executed is the event of the first instruction, namely, "Deliver 1 box of dolls from Port C to Port C."
[0117] Optionally, if the second agent is an embodied agent whose functions match the event included in the first instruction, the second response information includes the state of the second agent and / or the state of the event corresponding to the first instruction. The specific content of this information is similar to the first response information in the technical solution described above under "1. The first agent is a real embodied agent, and the functions of the first agent match the event." Please refer to the previous section and will not be further elaborated here.
[0118] For example, in the embodiment shown in FIG6(b), agent C sends a first response message instructing to call agent B. After displaying the first response message, the conversation interface displays a second response message from agent B. The second response message reads "OK, received," indicating that the second agent (i.e., agent B) is currently in normal working order and can execute the events included in the first instruction.
[0119] It is understandable that although the first response message is displayed before the second response message, the second response message is a reply from the second agent called by the first response message, and the content included in the second response message is actually a reply to the first instruction.
[0120] In this application, in a scenario where a first agent, instructed by a first instruction, fails to complete the event specified in the first instruction, a second agent can also reply with a second response message after the first agent replies with a first response message. This allows for multiple possibilities for both the second agent and the second response message, enriching the implementation methods and application scenarios of this technical solution and enhancing its flexibility and practicality.
[0121] 3. The first agent is a virtual embodied agent, and the function of the first agent matches the event.
[0122] In an embodiment of the present application, the execution subject included in the first instruction may also be a virtual embodied agent, which corresponds to a virtual entity in a virtual environment. The virtual environment can also be understood as a simulation of the physical environment, and the virtual entity is a virtualization or simulation of the physical entity, that is, the virtual embodied agent is a virtualization or simulation of a real embodied agent.
[0123] In this technical solution, the content of the first response is similar to the first response in the aforementioned technical solution "1. The first agent is a real embodied agent, and the first agent's function matches the event," as detailed above. The difference is that the state of the first agent included in the first response is the state of the virtual embodied agent, and the event state corresponding to the first instruction is also the event state in the virtual environment.
[0124] In practical applications, users can use this first response information to determine the operating status of the agent system in a virtual environment, thereby deciding whether to adjust the agent system or apply it to a real physical scenario. In other words, conducting evaluations in a virtual environment can help reduce development costs and improve efficiency.
[0125] 4. The first agent is a virtual embodied agent, and the function of the first agent does not match the event.
[0126] In this technical solution, the content of the first response message is similar to the first response message in the technical solution described above under "2. The first agent is a real embodied agent, and the first agent's functions do not match the event." For details, see the previous section and will not be repeated here. The difference is that in this technical solution, the first agent can also call other virtual embodied agents.
[0127] (2) The first instruction does not include the execution subject, but includes the event.
[0128] In a scenario where the first instruction includes only an event, the first AI agent is assumed to be the first agent, and the first response information includes call information for a third agent and / or an inability to execute information, wherein the function of the third agent matches the event included in the first instruction.
[0129] The following is an explanation with reference to a specific example. Please refer to FIG7 , which is a schematic diagram of a conversation interface provided in an embodiment of the present application.
[0130] It should be noted that in the embodiment shown in Figure 7, the first instruction is "deliver a box of food from Port C to Port D," and the event included in the first instruction is delivering a box of dolls from Port C to Port D. Since the first instruction does not include an execution subject, the first AI agent is assumed to be the first embodied agent.
[0131] Alternatively, in the embodiment shown in FIG7(a), the first response message returned by the first AI agent is a call message to a third agent (i.e., agent A). In this solution, after displaying the first response message, the response message returned by the third agent may also be displayed.
[0132] The response information returned by the third agent includes the state of the third agent and / or the event state corresponding to the first instruction. The specific content is similar to the first response information in the technical solution of "1. The first agent is a real embodied agent, and the function of the first agent matches the event". Please refer to the above text and will not be repeated here. For example, in the embodiment shown in Figure 7 (a), the response information returned by the third agent is "Charging, it will take 3 minutes to fully charge", indicating that the third agent (i.e., agent A) is currently in the charging state.
[0133] Optionally, in the embodiment shown in Figure 7 (b), the first response message replied by the first AI agent is an unexecution message, and the first response message indicates that "this system does not transport food", which means that all the agents included in the current agent system cannot execute the first instruction.
[0134] In this application, in scenarios where the first instruction does not include an execution subject, the first AI agent defaults to replying with a first response message. This means that even if the user enters an incomplete first instruction, a reply can still be received, simplifying user operations and further simplifying agent development. Furthermore, the first AI agent can call on other agents to complete the event included in the first instruction, or reply with a message indicating that the execution could not be performed, allowing the user to understand the progress of the event included in the first instruction, thereby enhancing the practicality of the technical solution of this application.
[0135] (3) The first instruction includes the execution subject but not the event.
[0136] In a scenario where the first instruction includes an execution subject but does not include an event, the first agent is the agent indicated by the execution subject. The first response information includes prompt information, which is used to remind the user to input the event, or to display the default function of the first agent, or to display the most recent operation of the first agent.
[0137] Based on the foregoing description, it can be seen that in the embodiments of the present application, natural language interaction with an agent simplifies the complex development or management process. Furthermore, in a solution where the first agent includes an embodied agent or the first agent manages an embodied agent, the complexity of embodied agent development is reduced and development efficiency is improved.
[0138] In the above description, the interaction between the agent and the user, as well as the interaction between agents, provided by the embodiments of the present application are introduced. The above embodiments do not limit the events included in the first instruction, which are determined based on the needs of actual applications.
[0139] For example, in the embodiments shown in Figures 3 to 7 above, the first instruction is for a port freight scenario, and the included event is cargo transportation at the port. In actual applications, the intelligent agent system can also be applied to other scenarios, such as office scenarios, home scenarios, etc., and the first instruction can be changed accordingly.
[0140] In the embodiment of the present application, the intelligent agent system also has other functions, such as data flow display, construction or modification of simulation environment, code generation and detection, etc., which are explained below.
[0141] In some optional implementations, the agent system can also display data streams. That is, the cloud platform can also display data streams that reflect the state or trajectory of physical entities in the physical environment, or the state or trajectory of virtual entities in the simulated environment. The cloud platform can display target data streams in a variety of ways, which are described below:
[0142] The cloud platform responds to the second input operation on the input bar and obtains a second instruction. The second instruction indicates the data flow of displaying the target event. The target event is an event that has been executed or is being executed by the intelligent system. The language of the second instruction is natural language.
[0143] Optionally, the cloud platform can directly display the data stream of the target event based on the second instruction. The second input operation is similar to the implementation of the first input operation, and will not be described in detail here.
[0144] Optionally, after receiving the second instruction, the cloud platform may also display a third response message on the conversation interface. The third response message is specific to the second instruction. In other words, the third response message indicates the data stream of the target event. The cloud platform then responds to the touch operation specific to the third response message by displaying the data stream of the target event.
[0145] In this application, after the cloud platform obtains the second instruction, it can not only directly display the data stream of the target event, but also display the data stream of the target event by displaying the third response information on the conversation interface, enriching the implementation method and application scenarios of the technical solution of this application and improving the flexibility of the technical solution.
[0146] In the embodiments of the present application, there are multiple possibilities for the data stream of the target event, including the real-time data stream of the target event, the historical data stream of the target event, the first-perspective data stream of the target event, or the third-perspective data stream of the target event. Among them, the first perspective is the perspective of the intelligent agent that executes the target event, which can also be called the perspective of the party involved. The third perspective is the perspective of the intelligent agent system, or it can also be understood as a global perspective. In layman's terms, the third perspective can also be called the God's perspective or the bystander's perspective. For example, in a port freight scenario, the third perspective is the perspective observed by the camera above the port.
[0147] Optionally, since the target event includes an event that has been completed or is currently being executed by the agent system, the data stream type of the target event can be set by default. For example, in a scenario where the target event is an event that has been completed by the agent system, the data stream of the target event is set by default to a historical data stream. In a scenario where the target event is an event that is currently being executed by the agent system, the data stream of the target event is set by default to a real-time data stream.
[0148] In the embodiment of the present application, there are multiple possible display locations for the data stream of the target event. The data stream of the target event can be displayed on the conversation interface, or on an interface other than the conversation interface.
[0149] In the embodiments of the present application, there are multiple possibilities for the display position of the data stream of the target event and the data stream of the target event, which can be determined according to actual applications, further improving the flexibility of the technical solution of the present application.
[0150] The following diagram further illustrates the process of displaying the data stream of a target event on a computer device. Please refer to Figures 8 to 10, where Figures 8 and 9 are schematic diagrams of interfaces provided by embodiments of the present application, and Figure 10 is a schematic diagram of a conversation interface provided by embodiments of the present application.
[0151] It should be noted that in the embodiments shown in Figures 8 to 10 , the first instruction is "I want to see the video stream of Agent B delivering a box of dolls from Port C to Port D." This means the target event is delivering a box of dolls from Port C to Port D, and the data stream is a first-person perspective video stream. That is, it's Agent B's own perspective.
[0152] In the embodiment shown in Figure 8 , after the conversation interface 300 displays the first instruction, the data stream from the perspective of agent B is directly displayed on the data stream display interface 400. Specifically, the display area of the data stream may be as shown in the shaded area shown in Figure 8 .
[0153] In the embodiment shown in FIG9 , after displaying the first instruction, conversation interface 300 displays a third response message from agent B. The third response message is a URL. In response to a click on the URL, the computer device displays data stream display interface 400 , thereby displaying the data stream of the target event. In other words, when a user clicks the link, data stream display interface 400 is displayed.
[0154] It should be noted that the embodiment shown in FIG9 uses the third response information including a URL as an example. In actual applications, the third response information may also include other types of media, such as links, files, etc., as long as they are used to indicate the data stream of the target event. Furthermore, the data stream display interface 400 may also be displayed using other software. Other software refers to software different from the software that provides the conversation interface 300, such as separate video playback software.
[0155] In some optional implementations, the third response information may also include historical data streams of different time periods. The user may select the time period he or she needs by using natural language or clicking a corresponding link or selection box, etc., which is not specifically limited here.
[0156] Optionally, after the user clicks the medium included in the third response information, a prompt sub-interface may be displayed to prompt the user to confirm whether to display the data stream of the target event. After the user confirms, the data stream display interface 400 is triggered. There are various possible ways for the user to confirm, such as clicking a confirmation control included in the prompt sub-interface, or refraining from any operation within a preset time period. The selection can be based on actual application needs and is not limited herein.
[0157] It should also be noted that in the embodiments shown in Figures 8 and 9 , both the conversation interface 300 and the data stream display interface 400 are displayed. In actual applications, the data stream display interface 400 may also overlay the conversation interface 300, or the data stream display interface 400 and the conversation interface 300 may overlap, or the data stream display interface 400 may obscure the conversation interface 300. The display positions of these two interfaces can be selected based on actual applications and are not specifically limited here.
[0158] Optionally, as shown in Figures 8 and 9, the data stream display interface 400 may further include function controls 410. From left to right, the function controls 410 include a volume control, a rotation control, a back control, a pause control, a forward control, and a screenshot control, respectively, for controlling the video volume, rotating the video display angle, replaying the video, pausing playback, fast-forwarding the video, and taking screenshots. In actual applications, a greater or fewer number of function controls may be included, such as a speed control, but the specifics are not limited here.
[0159] In the embodiment shown in Figure 10 , after the conversation interface 300 displays the first instruction, the cloud platform displays the video stream of the target event responded to by agent B on conversation interface 300. Specifically, the display area of this video stream can be shown as the shaded area in Figure 10 . In other words, the data stream of the target event can be embedded in the conversation interface for display.
[0160] It should also be noted that the shaded area shown in Figures 8 to 10 is only an example of the video stream display area. In actual applications, the area can also be scaled, rotated, etc. based on user operations, which is not limited here.
[0161] In the embodiments shown in FIG8 to FIG10 , the first instruction specifies the viewing angle of the target event's data stream. In practical applications, the viewing angle may not be specified. In this solution, the data stream of the target event displayed by the computer device may have multiple possibilities:
[0162] Optionally, the first-person perspective or the third-person perspective video stream of the target event can be displayed by default. A switch control can also be displayed to switch the perspective of the data stream of the target event.
[0163] Optionally, the data stream of the first perspective and the data stream of the third perspective of the target event can be displayed by default, and then the data stream of one of the perspectives can be zoomed based on the user's selection.
[0164] In the embodiments shown in Figures 8 to 10 above, the first instruction specifies which specific event the target event is. In actual applications, the perspective may also be specified, but the specific event is not specified. In this solution, there are multiple possibilities for the data stream of the target event displayed by the cloud platform: Optionally, the data stream of the event being executed by the current embodied intelligent agent can be set by default. Optionally, the historical data stream of the most recent event executed by the current embodied intelligent agent can be set by default. Optionally, the data stream of the default time period can be set by default. No specific limitations are made here.
[0165] It should also be noted that the embodiments shown in Figures 8 to 10 take video streams as an example. In actual applications, the data streams can also be data streams of other sensors, such as infrared sensors, laser sensors, radars, inertial measurement units (IMUs), etc., which are not specifically limited here.
[0166] In some optional implementations, the intelligent agent system provided in the embodiments of the present application further constructs or changes a simulation environment. The simulation environment is used to develop or test embodied intelligent agents. In summary, the cloud platform responds to a third input operation on the input field and obtains a third instruction. The third instruction is written in natural language and instructs the establishment of a target simulation environment. In other words, the third instruction includes an event for establishing the target simulation environment. Based on the third instruction, the cloud platform displays the target simulation environment on the simulation interface.
[0167] Optionally, if the third instruction does not include an execution subject, the computer device may display a response message from the first AI agent on the conversation interface, where the response message calls the simulation agent to execute the third instruction.
[0168] Optionally, if the execution subject included in the third instruction is a second AI agent, then the second AI agent executes the third instruction. The second AI agent is used to build a simulation environment and can also be called a simulation agent.
[0169] Optionally, if the capabilities of the execution subject included in the third instruction are unable to display the target simulation environment, then the execution subject may call the second AI agent to execute the third instruction.
[0170] The above three possible implementation methods are similar to the various possibilities of the cloud platform calling the first intelligent agent to respond to the first instruction through the first response information, and will not be repeated here.
[0171] The simulation interface is used to visualize a 3D or 2D virtual environment, which is also known as a simulation environment. In addition, the cloud platform can display the simulation interface in a variety of ways, similar to the data flow display interface 400 in the embodiments shown in Figures 8 to 10 above, as shown above, and will not be repeated here.
[0172] For example, the user inputs the requirements for the simulation environment, such as "construct a simulation environment for the dock, including 3 ports and 6 piles of cargo, and the weather is sunny and the time is in the morning." This instruction does not specify an execution subject, and the first AI agent calls the second AI agent to execute the instruction. The second AI agent can be a single agent, responsible for realizing the construction of the aforementioned simulation environment; or it can be multiple agents, each with different functions, such as changing the light, number of items, texture, position, etc. In the solution where there is only one second agent, there is no need to build other agents, which reduces the amount of calculation. In the solution where there are multiple second agents, the powers of each agent are clear, which is convenient for maintenance.
[0173] The different functions mentioned above include, but are not limited to: changing the type, quantity, texture, position, shape and other object properties in the simulation environment; changing environmental factors such as lighting and humidity in the environment; generating events and test cases executed by embodied intelligent agents, etc.
[0174] For example, please refer to FIG11 , which is a schematic diagram of a simulation interface provided in an embodiment of the present application.
[0175] As shown in Figure 11, the simulation interface displays a 3D terminal, including three ports and six cargo piles. In the embodiment shown in Figure 11, the three ports are adjacent, but Port C, which is adjacent to Port B, is not shown in Figure 11 due to obstructed vision.
[0176] In some optional embodiments, the simulation interface may further include function controls 510. In the embodiment shown in FIG11 , the functions of the function controls 510 from left to right are respectively Save, Cut, Refresh, Undo, and Select Tool. In actual applications, the simulation interface may further include a greater or lesser number of function controls, which is not specifically limited herein.
[0177] In a simulation environment, a virtual embodied agent corresponds to a simulated entity within the simulation environment. For example, in the embodiment shown in FIG11 , the simulation environment also includes simulated robots 1 and 2, both used to carry cargo. A conversational interface can then be used to allow users to invoke virtual embodied agents to simulate cargo handling, or to allow simulated embodied agents to invoke each other.
[0178] That is, in this embodiment of the present application, the cloud platform responds to a fifth input operation on the input field by obtaining a fifth instruction, the fifth instruction being in natural language and including an event that matches the function of the virtual embodied agent. The cloud platform then displays a fourth response message from the virtual embodied agent on the conversation interface, the fourth response message including the state of the virtual embodied agent and / or the state of the event corresponding to the fourth instruction.
[0179] For example, assume that the simulation environment shown in FIG11 includes a virtual embodied agent A, whose function is to transport cargo between Port A and Port B. The fifth instruction input by the user includes an event of transporting a box of cargo from Port A to Port B.
[0180] Optionally, if the fifth instruction includes an execution subject that is virtual embodied agent A, the cloud platform displays the aforementioned fourth response information on the conversation interface. For example, virtual embodied agent A responds "OK, received" on the conversation interface, and in the simulation environment, virtual embodied agent A is in a freight transport state.
[0181] Alternatively, if the function of the execution subject included in the fifth instruction does not match the aforementioned event, the fourth response information may be a call information directed to virtual embodied agent A, causing virtual embodied agent A to execute the aforementioned event. Additionally, the conversation interface may also display the response information returned by virtual embodied agent A.
[0182] Optionally, if the fifth instruction does not include an execution subject, the cloud platform may also display the response information replied by the first AI agent before displaying the fourth response information, and the response information is for the call information of the virtual embodied agent A.
[0183] In this application, a virtual embodied agent can be called upon to complete instructions, thereby changing the virtual environment. This can also serve as a simulation of changes in the real physical environment, providing a reference for applying corresponding physical entities in the physical environment, or providing simulation data, thereby improving the practicality of the technical solution of this application.
[0184] In an embodiment of the present application, the computer device may also determine the type of data generated by the simulation environment.
[0185] Optionally, the computer device responds to a fourth input operation on the input field by obtaining a fourth instruction, the fourth instruction being in a natural language and instructing the generation of data of the target type. Based on the fourth instruction, a fourth response message is displayed on the conversation interface, the fourth response message indicating the data of the target type. Optionally, the fourth response message may include the data of the target type, or include a medium indicating the data of the target type. In the latter scenario, the cloud platform responds to the touch operation on the medium by displaying the data of the target type.
[0186] Optionally, the cloud platform can also configure the data type generated by the simulation environment by default. That is, in a scenario where the user does not specify the type of generated data, the cloud platform displays the generated data of the default type.
[0187] In the embodiment of the present application, there are many possible target types, including radar data, depth data, RGB data, trajectory data, etc., which are not specifically limited here.
[0188] In this application, the cloud platform can also display a simulation environment, through which physical entities can be simulated or simulated, and embodied intelligent agents can be developed and tested. Furthermore, in this application, the data type generated by the simulation environment can be indicated. There are multiple possibilities for data types and fourth response information, enriching the implementation methods of the technical solutions of this application.
[0189] In some optional implementations, the intelligent agent system provided in the embodiments of the present application can also perform code generation and detection. In summary, the cloud platform responds to the sixth input operation for the input bar and obtains the sixth instruction. The language of the sixth instruction is natural language, and the sixth instruction instructs to generate the skill code of the fourth agent. Based on the sixth instruction, the cloud platform displays the skill code of the fourth agent on the code interface. In other words, the event included in the sixth instruction is to generate the skill code of the fourth agent. The skill code is used to describe the functions possessed by the fourth agent. In addition, it can also describe the working logic of the fourth agent.
[0190] Alternatively, if the sixth instruction does not include an execution subject, the computer device may display a response message from the first AI agent on the conversation interface, wherein the response message calls a third AI agent to execute the sixth instruction. The third AI agent is used to generate agent code.
[0191] Optionally, if the execution subject included in the sixth instruction is a third AI agent, then the third AI agent executes the sixth instruction.
[0192] Optionally, if the capabilities of the execution subject included in the sixth instruction cannot generate the skill code of the fourth agent, then the execution subject can call the third AI agent to execute the sixth instruction.
[0193] The above three possible implementations are similar to the various possibilities of the first response information mentioned above and will not be repeated here.
[0194] In addition, there are many possible ways for a computer device to display a code interface, which are similar to the data flow display interface 400 in the embodiments shown in Figures 8 to 10 above. Please refer to the above and will not be repeated here.
[0195] In some optional implementations, the code interface may display function controls in addition to displaying the code, for implementing different functions. For example, please refer to FIG12 , which is a schematic diagram of the code interface provided in an embodiment of the present application.
[0196] As shown in Figure 12, the code interface includes a code language control, which is used to set the language type of the code displayed on the code interface. Exemplarily, when the user clicks the control, the code interface can display a code language selection bar, which includes multiple code languages, which are determined by the user.
[0197] As shown in FIG12 , the code interface may further include a code format control for setting the format of the code, including the font, size, color, etc. of the code.
[0198] As shown in FIG12 , the code interface may further include a panel style control for setting the style of the code interface, including the background color of the code interface, whether to use night mode, and the like.
[0199] In some optional implementations, the agent system provided in the embodiment of the present application can also detect codes. The cloud platform tests the skill code of the fourth agent and displays the test result of the skill code of the fourth agent.
[0200] Optionally, as shown in FIG12 , the code interface may include a test control. When the user clicks the test control, the cloud platform tests the skill code of the fourth agent displayed on the code interface.
[0201] Optionally, the skill code detection can also be triggered in the conversation interface. For example, the user enters "Detect the skill code of the fourth agent" in the input field of the conversation interface, and the first AI agent calls the code detection agent to detect the skill code of the fourth agent. The code detection agent can then reply with the test results of this detection in the conversation interface. Optionally, the code detection agent and the third AI agent used to generate the code shown above can be different AI agents, or they can be the same AI agent, and both can become code agents.
[0202] Optionally, the code generated by the cloud platform can be stored by the first AI agent or the code agent in a specific location and assigned a code name. When the user needs to test the code in the virtual environment, by specifying the code name and detection event, the cloud platform can also use the code detection agent to test the code corresponding to the code name. Code agents include code detection agents and code generation agents.
[0203] In this application, the cloud platform can also generate and display the agent's skill code, providing a basis for code detection and improving the feasibility of the technical solution. In addition, by detecting the code, users can determine whether the code can be executed accurately, thereby improving the accuracy of agent development.
[0204] In the embodiments shown in Figures 8 to 12 above, the data flow display interface, code interface, and simulation interface are all displayed based on the user's input operations in the input field. In actual applications, they can also be displayed in other ways, which are explained below.
[0205] In some optional implementations, the cloud platform may further display a menu bar, which includes display controls for different interfaces. When a user clicks on the control, the cloud platform displays the interface corresponding to the control.
[0206] In some optional implementations, the cloud platform may further display a search bar. When a user enters "display XX interface" or other indication events in the search bar as instructions for displaying a certain interface, the cloud platform displays the corresponding interface.
[0207] For example, please refer to FIG13 , which is a schematic diagram of the interface provided in an embodiment of the present application.
[0208] As shown in FIG13 , the menu bar includes four display controls 610 , which are, from left to right, a data flow interface display control, a code interface display control, a simulation interface display control, and a session interface display control.
[0209] Optionally, the user clicks on the code interface display control, and the cloud platform can display the code interface shown in Figure 12. The user clicks on the simulation interface display control, and the cloud platform can display the simulation interface shown in Figure 11.
[0210] In some optional embodiments, when a user clicks a display control, the cloud platform may also display a next-level menu. For example, in the embodiment shown in FIG13 , when a user clicks a data stream interface display control, the cloud platform displays a perspective selection menu, allowing the user to select to display the data stream from the first perspective and / or the third perspective.
[0211] For example, in the embodiment shown in FIG13 , the cloud platform displays a third-person video stream, in which the status of each intelligent agent can also be marked.
[0212] In some optional implementations, the cloud platform may also display an agent introduction interface, which is used to demonstrate the functions of each agent.
[0213] For example, please refer to FIG14 , which is a schematic diagram of the interface provided in an embodiment of the present application.
[0214] As shown in Figure 14, agents include operational agents, development agents, and system users. Operational agents assist in the operation of agents and can include agents used in actual operation, such as management agents, agent A for cargo handling, and agent B for clearing roadblocks. Development agents assist in the development of agents and can include simulation agents, code checking agents, and agents that modify the simulation environment.
[0215] As shown in Figure 14, the computer device can display the functions of the agent behind the agent's avatar. In actual applications, the functions of the agent can also be displayed in other ways.
[0216] Optionally, the functions of the intelligent agent may be displayed when the input device moves to a specific area, and the specific area includes an avatar display area of the intelligent agent.
[0217] Optionally, the functions of the agent can be displayed through specific operations, such as right-clicking or double-clicking the agent's avatar.
[0218] Optionally, you can also display the agent's features in the conversation interface. For example, a user can @ an agent and trigger it to reply with a feature introduction. Alternatively, a user can @ an agent and have it introduce its features.
[0219] It should be noted that the interface shown in the aforementioned embodiments is merely an example and does not limit the interface actually displayed in the interaction method of the embodied intelligent body provided in the embodiments of the present application.
[0220] Next, please refer to Figure 15, which is a schematic diagram of the structure of the cloud platform provided in an embodiment of the present application. The cloud platform runs an intelligent agent system, which includes multiple intelligent agents.
[0221] In some optional implementations, the cloud platform 1500 includes:
[0222] The display unit 1501 is configured to display a conversation interface, which includes an input field.
[0223] The transceiver unit 1502 is configured to obtain a first instruction in response to a first input operation on the input field, where the language of the first instruction is a natural language.
[0224] The display unit 1501 is further used to display the first response information replied by the first agent on the conversation interface. The first response information is for the first instruction, and the first agent is determined based on the first instruction.
[0225] In some optional embodiments, the first agent is a real embodied agent. If the execution subject included in the first instruction is the first agent, and the event included in the first instruction matches the function of the first agent, the first response information includes at least one of the following: the state of the first agent and the state of the event corresponding to the first instruction.
[0226] In some optional embodiments, the first agent is a real embodied agent. If the execution subject included in the first instruction is the first agent, and the event included in the first instruction does not match the function of the first agent, then the first response information includes: call information for a second agent, the second agent is a real embodied agent, and the function of the second agent matches the event. Alternatively, the second agent is a first AI agent, and the first AI agent is used to manage the agent system.
[0227] In some optional embodiments, the display unit 1501 is further configured to display a second response message from the second agent on the conversation interface, where the second response message is in response to the first instruction.
[0228] In some optional embodiments, if the first instruction does not include an execution subject, the first agent is a first AI agent, and the first response information includes: call information for a third agent, and the function of the third agent matches the event included in the first instruction.
[0229] In some optional embodiments, in a scenario where the first instruction includes an execution subject but does not include an event, the first agent is the agent indicated by the execution subject. The first response information includes prompt information, which is used to remind the user to input the event, or to display the default function of the first agent, or to display the most recent operation of the first agent.
[0230] In some optional embodiments, the transceiver unit 1502 is further configured to respond to a second input operation on the input field and obtain a second instruction, wherein the second instruction instructs displaying the data stream of a target event, where the target event is an event that has been completed or is being executed by the agent system, and the second instruction is in natural language. The display unit 1501 is further configured to display the data stream of the target event.
[0231] In some optional implementations, the display unit 1501 is further configured to display a third response message on the conversation interface, the third response message being directed to the second instruction, and to display a data stream of the target event in response to a touch operation directed to the third response message.
[0232] In some optional implementations, the display unit 1501 is specifically configured to display the data stream of the target event on the conversation interface, or to display the data stream of the target event on an interface other than the conversation interface.
[0233] In some optional embodiments, the data stream of the target event includes: a real-time data stream of the target event, a historical data stream of the target event, a first-perspective data stream of the target event, or a third-perspective data stream of the target event. The first perspective is the perspective of the agent executing the target event, and the third perspective is the perspective of the agent system.
[0234] In some optional embodiments, the transceiver unit 1502 is further configured to respond to a third input operation on the input field and obtain a third instruction, wherein the third instruction is in a natural language and instructs to build a target simulation environment. The display unit 1501 is further configured to display the target simulation environment on the simulation interface based on the third instruction.
[0235] In some optional embodiments, the transceiver unit 1502 is further configured to, in response to a fourth input operation on the input field, obtain a fourth instruction, where the fourth instruction is in a natural language and instructs the generation of data of a target type. The display unit 1501 is further configured to display a fourth response message on the conversation interface based on the fourth instruction, where the fourth response message indicates the data of the target type.
[0236] In some optional embodiments, the transceiver unit 1502 is further configured to respond to a fifth input operation on the input field and obtain a fifth instruction, wherein the fifth instruction is in a natural language and the event included in the fifth instruction matches the function of the virtual embodied agent. The display unit 1501 is further configured to display a fourth response message from the virtual embodied agent on the conversation interface, wherein the fourth response message includes the state of the virtual embodied agent and / or the state of the event corresponding to the fourth instruction.
[0237] In some optional embodiments, the transceiver unit 1502 is further configured to, in response to a fifth input operation on the input field, obtain a fifth instruction, where the fifth instruction is in a natural language and instructs the generation of a skill code for the fourth agent. The display unit 1501 is further configured to display the skill code for the fourth agent on the code interface based on the fifth instruction.
[0238] In some optional implementations, the processing unit 1503 is further configured to test the skill code of the fourth agent. The display unit 1501 is further configured to display the test result of the skill code of the fourth agent.
[0239] The display unit 1501, the transceiver unit 1502, and the processing unit 1503 can all be implemented by software or hardware. For example, the implementation of the processing unit 1503 will be described below using the processing unit 1503 as an example. Similarly, the implementation of the acquisition unit 801 can refer to the implementation of the processing unit 1503.
[0240] As an example of a software functional unit, the processing unit 1503 may include code running on a computing instance. The computing instance may include at least one of a physical host (computing device), a virtual machine, and a container. Furthermore, the computing instance may be one or more. For example, the processing unit 1503 may include code running on multiple hosts / virtual machines / containers. It should be noted that the multiple hosts / virtual machines / containers used to run the code may be distributed in the same region or in different regions. Furthermore, the multiple hosts / virtual machines / containers used to run the code may be distributed in the same availability zone (AZ) or in different AZs, each AZ including one data center or multiple geographically close data centers. Typically, a region may include multiple AZs.
[0241] Similarly, multiple hosts / virtual machines / containers running the code can be distributed within the same virtual private cloud (VPC) or across multiple VPCs. Typically, a VPC is set up within a region. Cross-region communication between two VPCs within the same region, or between VPCs in different regions, requires a communication gateway within each VPC to interconnect the VPCs.
[0242] As an example of a hardware functional unit, processing unit 1503 may include at least one computing device, such as a server. Alternatively, processing unit 1503 may be implemented using an application-specific integrated circuit (ASIC) or a programmable logic device (PLD). The PLD may be implemented using a complex programmable logical device (CPLD), a field-programmable gate array (FPGA), a generic array logic (GAL), or any combination thereof.
[0243] The multiple computing devices included in processing unit 1503 can be distributed in the same region or in different regions. The multiple computing devices included in processing unit 1503 can be distributed in the same AZ or in different AZs. Similarly, the multiple computing devices included in processing unit 1503 can be distributed in the same VPC or in multiple VPCs. The multiple computing devices can be any combination of servers, ASICs, PLDs, CPLDs, FPGAs, GALs, and other computing devices.
[0244] It should be noted that the full functionality of cloud platform 1500 is achieved by implementing the various steps of the agent information processing method via display unit 1501, transceiver unit 1502, and processing unit 1503. Cloud platform 1500 is used to implement the operations performed by the computer device in the embodiments shown in Figures 1 to 14 above to implement the agent interaction method provided in the embodiments of this application, and will not be further described here.
[0245] Please refer to Figure 16, which is a schematic diagram of the structure of a computing device provided in an embodiment of the present application. The computing device 1600 includes a processor 1601, a communication interface 1602, a bus 1603, and a memory 1604. The processor 1601, the communication interface 1602, and the memory 1604 communicate with each other via the bus 1603. In practical applications, communication can also be achieved through other means such as wireless transmission, which is not limited here.
[0246] The computing device 1600 may be a server or a terminal device. It should be understood that this application does not limit the number of processors and memories in the computing device 1600 .
[0247] The processor 1601 may include any one or more processors such as a central processing unit (CPU), a graphics processing unit (GPU), a microprocessor (MP), or a digital signal processor (DSP).
[0248] The communication interface 1602 uses a transceiver module such as, but not limited to, a network interface card or a transceiver to implement communication between the computing device 1600 and other devices or a communication network.
[0249] Bus 1603 may be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, among others. Buses may be classified as address buses, data buses, control buses, and the like. For ease of illustration, FIG16 shows only one line, but this does not imply a single bus or type of bus. Bus 1603 may include a path for transmitting information between various components of computing device 1600 (e.g., memory 1604, processor 1601, and communication interface 1602).
[0250] The memory 1604 may include volatile memory, such as random access memory (RAM). The memory 1604 may also include non-volatile memory, such as read-only memory (ROM), flash memory, hard disk drive (HDD), or solid state drive (SSD).
[0251] Memory 1604 stores executable program code. Processor 1601 executes this executable program code to implement the functions of display unit 1501, transceiver unit 1502, and processing unit 1503, thereby implementing the agent's data processing method. In other words, memory 1604 stores instructions for executing the agent's data processing method.
[0252] The present application also provides a computing device cluster comprising at least one computing device. The computing device may be a server, such as a central server, an edge server, or a local server in a local data center. In some optional implementations, the computing device may also be a terminal device such as a desktop computer, a laptop computer, or a smartphone.
[0253] Please refer to Figures 17 and 18, which are both structural diagrams of the computing device cluster provided in embodiments of the present application.
[0254] As shown in Figure 17, the computing device cluster includes at least one computing device 1600. The memory 1604 in one or more computing devices 1600 in the computing device cluster may store the same instructions for executing the data processing method of the intelligent agent provided in the embodiment of the present application.
[0255] In some possible implementations, the memory 1604 of one or more computing devices 1600 in the computing device cluster may also store partial instructions for executing the agent's data processing method. In other words, the combination of one or more computing devices 1600 can jointly execute instructions for executing the agent's data processing method.
[0256] It should be noted that the memory 1604 in different computing devices 1600 in the computing device cluster can store different instructions, each used to execute a portion of the functions of the cloud platform. In other words, the instructions stored in the memory 1604 in different computing devices 1600 can implement the functions of one or more of the display unit 1501, the transceiver unit 1502, and the processing unit 1503.
[0257] In some possible implementations, one or more computing devices in a computing device cluster may be connected via a network. The network may be a wide area network (WAN) or a local area network (LAN), among others. FIG. 18 illustrates a possible implementation. As shown in FIG. 18 , two computing devices 1600A and 1600B are connected via a network. Specifically, the connection to the network is achieved via a communication interface in each computing device. In this type of possible implementation, the memory 1604 in the computing device 1600A stores instructions for executing the functions of the display unit 1501. Simultaneously, the memory 1604 in the computing device 1600B stores instructions for executing the functions of the transceiver unit 1502 and the processing unit 1503.
[0258] The connection method between the computing device clusters shown in Figure 18 can be based on the fact that in the data processing method of the intelligent body provided in this application, display operations and operations other than display operations are performed separately, that is, the function of the display unit 1501 is considered to be executed by the computing device 1600A, and the functions of the transceiver unit 1502 and the processing unit 1503 are considered to be executed by the computing device 1600B.
[0259] It should be understood that the functionality of the computing device 1600A shown in FIG18 may also be implemented by multiple computing devices 1600. Similarly, the functionality of the computing device 1600B may also be implemented by multiple computing devices 1600.
[0260] The present application also provides another computing device cluster, wherein the connection relationship between the computing devices in the computing device cluster can be similar to the connection relationship between the computing device clusters described in FIG17 and FIG18, and will not be repeated here.
[0261] The present application also provides a computer program product comprising instructions. The computer program product may be software or a program product comprising instructions that can be run on a computing device or stored on any available medium. When the computer program product is run on at least one computing device, the at least one computing device executes the aforementioned agent data processing method.
[0262] The present application also provides a computer-readable storage medium. The computer-readable storage medium can be any available medium that can be stored by a computing device or a data storage device such as a data center that contains one or more available media. The available medium can be a magnetic medium (e.g., a floppy disk, a hard disk, a magnetic tape), an optical medium (e.g., a DVD), or a semiconductor medium (e.g., a solid-state drive). The computer-readable storage medium includes instructions that instruct the computing device to execute the data processing method of the above-mentioned intelligent agent.
[0263] Those skilled in the art will clearly understand that, for the convenience and brevity of description, the specific working processes of the systems, devices and units described above can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.
[0264] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them. Although the present application has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the protection scope of the technical solutions of the embodiments of the present application.
Claims
1. An information processing method of an intelligent agent, characterized in that: The method is applied to a cloud platform, the cloud platform runs an intelligent agent system, the intelligent agent system includes multiple intelligent agents, and the method includes: Displaying a conversation interface, wherein the conversation interface includes an input field; In response to a first input operation on the input field, obtaining a first instruction, where the language of the first instruction is a natural language; The first response information replied by the first agent is displayed on the conversation interface, the first response information is for the first instruction, and the first agent is determined based on the first instruction.
2. The method according to claim 1, characterized in that The first agent is a real embodied agent; If the execution subject included in the first instruction is the first intelligent agent, and the event included in the first instruction matches the function of the first intelligent agent, then the first response information includes: the status of the first intelligent agent, and / or the event status corresponding to the first instruction.
3. The method according to claim 1, characterized in that The first agent is a real embodied agent; If the execution subject included in the first instruction is the first agent, and the event included in the first instruction does not match the function of the first agent, the first response information includes: Regarding the calling information of the second agent, the second agent is a real embodied agent, and the function of the second agent matches the event, or the second agent is a first artificial intelligence AI agent, and the first AI agent is used to manage the agent system.
4. The method according to claim 3, characterized in that After the conversation interface displays the first response information replied by the first agent, the method further includes: The second response information replied by the second agent is displayed on the conversation interface, and the second response information is for the first instruction.
5. The method according to claim 1, wherein If the first instruction does not include an execution subject, the first agent is a first AI agent, and the first response information includes: Regarding the calling information of the third agent, the function of the third agent matches the event included in the first instruction.
6. The method according to any one of claims 1 to 5, characterized in that The method further comprises: In response to a second input operation on the input field, obtaining a second instruction, wherein the second instruction instructs displaying a data stream of a target event, the target event being an event that has been executed or is being executed by the agent system, and the language of the second instruction is a natural language; Displays the data stream of the target event.
7. The method according to claim 6, characterized in that Before displaying the data stream of the target event, the method further includes: Displaying third response information on the conversation interface, where the third response information is for the second instruction; The data stream displaying the target event includes: In response to the touch operation for the third response information, the data stream of the target event is displayed.
8. The method according to claim 6 or 7, characterized in that The data stream displaying the target event includes: The data stream of the target event is displayed on the session interface, or the data stream of the target event is displayed on an interface other than the session interface.
9. The method according to any one of claims 6 to 8, characterized in that The data stream of the target event includes: a real-time data stream of the target event, a historical data stream of the target event, a first-perspective data stream of the target event, or a third-perspective data stream of the target event; The first perspective is the perspective of the agent executing the target event, and the third perspective is the perspective of the agent system.
10. The method according to any one of claims 1 to 9, characterized in that The method further comprises: In response to a third input operation on the input field, obtaining a third instruction, wherein the third instruction is in a natural language and instructs to build a target simulation environment; Based on the third instruction, the target simulation environment is displayed on the simulation interface.
11. The method according to claim 10, characterized in that The method further comprises: In response to a fourth input operation on the input field, obtaining a fourth instruction, wherein the fourth instruction is in a natural language and instructs generating data of a target type; Based on the fourth instruction, a fourth response message is displayed on the conversation interface, where the fourth response message indicates data of the target type.
12. The method according to claim 10 or 11, characterized in that The method further comprises: In response to a fifth input operation on the input field, obtaining a fifth instruction, wherein the fifth instruction is in a natural language and the event included in the fifth instruction matches the function of the virtual embodied intelligent agent; The fourth response information replied by the virtual embodied intelligent body on the conversation interface is displayed, wherein the fourth response information includes the status of the virtual embodied intelligent body and / or the event status corresponding to the fourth instruction.
13. The method according to any one of claims 1 to 12, characterized in that The method further comprises: In response to a sixth input operation on the input field, obtaining a sixth instruction, wherein the sixth instruction is in a natural language and instructs generating a skill code for a fourth agent; Based on the sixth instruction, the skill code of the fourth agent is displayed on the code interface.
14. The method according to claim 13, characterized in that The method further comprises: testing a skill code of the fourth embodied intelligent agent; Display the test results of the skill code of the fourth embodied intelligent agent.
15. An information processing device for an intelligent agent, characterized in that: The invention comprises modules for implementing the method according to any one of the preceding claims 1 to 14.
16. A computer equipment cluster, characterized in that: comprising at least one computing device, each computing device including a processor and a memory; The processor of the at least one computing device is configured to execute instructions stored in a memory of the at least one computing device, so as to enable the computing device cluster to perform the method according to any one of claims 1 to 14.
17. A computer-readable storage medium, characterized in that The computer-readable storage medium includes computer program instructions, which, when executed by a computing device cluster, cause the computing device cluster to perform the method according to any one of claims 1 to 14.
18. A computer program product comprising instructions, characterized in that When the instructions are executed by a computer device cluster, the computer device cluster is caused to perform the method according to any one of claims 1 to 14.
Citation Information
Patent Citations
Intelligent agent information processing method and related equipment
CN120597927A
Message processing method, device and readable storage medium
CN113569037A
Message processing method and device, equipment and storage medium
CN113885757A
Workflow calling method and device, computer equipment and storage medium
CN114971505A
Session message processing method and device, computer equipment and storage medium
CN116415008A