Human-computer interaction method, device and embodied intelligent agent based on embodied intelligent agent
By introducing object memory bank and action matching mechanisms into the embodied agent, responding to user description text, determining interactive tasks and target objects, and matching candidate actions, the embodied agent's lack of correlation between objects and tasks in complex scenarios is solved, and efficient and accurate interactive tasks are achieved.
Patent Information
- Application Number
- CN202411961508.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-30
- Publication Date
- 2025-05-16
- Estimated Expiration
- 2044-12-30
AI Technical Summary
The limited understanding of the correlation between objects and tasks in complex scenarios of existing embodied agents makes it difficult to meet actual needs for the accuracy and efficiency of interactive tasks.
A human-computer interaction method based on an embodied agent is proposed. By responding to the user's description text, the interactive task and the target object are determined, and the target perception data is retrieved in the pre-established object memory bank, and the candidate actions are matched to complete the task.
It significantly improves the task execution efficiency and adaptability of the embodied intelligent body, improves the completion and accuracy of interactive tasks, can dynamically adapt to scene changes, accurately locate target objects, and meet users' diverse task needs.
Smart Images

Figure CN119376549B_ABST
Abstract
Description
Technical Field
[0001] The present application belongs to the field of artificial intelligence, and in particular, relates to a human-computer interaction method, device and embodied intelligent agent based on embodied intelligent agent. Background Art
[0002] In recent years, with the continuous development of artificial intelligence technology and natural language processing technology, the application of intelligent systems in scene understanding and task execution has become increasingly widespread. Especially in the fields of service robots, smart homes and industrial automation, users hope to command the system to complete complex scene-based interactive tasks through natural language. Traditional intelligent systems usually use fixed instructions or predefined operation sequences to complete tasks, but in practical applications, the flexibility and robustness of these systems are often limited.
[0003] In the existing technology, embodied agents are an intelligent system with environmental perception and physical interaction capabilities, and are widely used in a variety of complex scenarios. However, current embodied agents can only perform tasks based on simple instructions, but their understanding of the relationship between objects and tasks in complex scenes is relatively limited, resulting in the accuracy and efficiency of interactive tasks being difficult to meet actual needs. Summary of the invention
[0004] The present application aims to solve at least one of the technical problems existing in the related art. To this end, the present application proposes a human-computer interaction method, device and embodied intelligent agent based on embodied intelligent agent to efficiently and accurately realize interaction with users.
[0005] In a first aspect, the present application provides a human-computer interaction method based on an embodied intelligent agent, the method comprising:
[0006] In response to a description text sent by a user for a target scene, determining an interactive task indicated by the description text;
[0007] determining a target object associated with the interactive task;
[0008] Searching a pre-established object memory library to obtain target perception data corresponding to the target object; the object memory library is used to store the perception data of all objects pre-identified from the target scene;
[0009] Based on the target perception data, determining at least one target action from a plurality of pre-configured candidate actions;
[0010] The target action is performed on the target object to complete the interactive task.
[0011] In the above technical scheme, by responding to the description text sent by the user for a target scene, the interactive task indicated by the description text is determined, the user's multimodal input can be responded to and analyzed, and the user's interactive experience can be improved; by determining the target object associated with the interactive task, and searching the pre-established object memory library, the target perception data corresponding to the target object is obtained and introduced into the object memory library, and the object memory library is introduced to persistently remember the objects in the target scene, so that the embodied intelligent body can dynamically adapt to scene changes and accurately locate the target object, thereby improving the completion and accuracy of the interactive task; then, based on the target perception data, at least one target action is determined from a plurality of pre-configured candidate actions, and the target action is performed on the target object to complete the interactive task. Through the action matching mechanism, the embodied intelligent body can adapt to the diverse task requirements in different scenarios, significantly improving the task execution efficiency and adaptability of the embodied intelligent body, thereby better meeting user needs.
[0012] According to an embodiment of the present application, determining the target object associated with the interactive task includes:
[0013] Performing semantic understanding on the description text to determine a first target object indicated by the description text; the first target object is a target object for a target action to be performed;
[0014] Determining whether the description text indicates a second target object associated with the first target object; the second target object includes a target object that provides a location clue for the first target object to perform the target action; the location clue is included in the description text;
[0015] If the description text indicates a second target object, a target object associated with the interactive task is determined based on the first target object and the second target object.
[0016] In a second aspect, the present application provides a human-computer interaction device based on an embodied intelligent agent, the device comprising:
[0017] A response module, configured to respond to a description text sent by a user for a target scene and determine an interactive task indicated by the description text;
[0018] A determination module, used to determine a target object associated with the interactive task;
[0019] A retrieval module, used to search in a pre-established object memory library to obtain target perception data corresponding to the target object; the object memory library is used to store the perception data of all objects pre-identified from the target scene;
[0020] a planning module, configured to determine at least one target action from a plurality of pre-configured candidate actions based on the target perception data;
[0021] An execution module is used to execute the target action on the target object to complete the interactive task.
[0022] In the above technical scheme, by responding to the description text sent by the user for a target scene, the interactive task indicated by the description text is determined, the user's multimodal input can be responded to and analyzed, and the user's interactive experience can be improved; by determining the target object associated with the interactive task, and searching the pre-established object memory library, the target perception data corresponding to the target object is obtained and introduced into the object memory library, and the object memory library is introduced to persistently remember the objects in the target scene, so that the embodied intelligent body can dynamically adapt to scene changes and accurately locate the target object, thereby improving the completion and accuracy of the interactive task; then, based on the target perception data, at least one target action is determined from a plurality of pre-configured candidate actions, and the target action is performed on the target object to complete the interactive task. Through the action matching mechanism, the embodied intelligent body can adapt to the diverse task requirements in different scenarios, significantly improving the task execution efficiency and adaptability of the embodied intelligent body, thereby better meeting user needs.
[0023] In a third aspect, the present application provides an embodied intelligent body, wherein the embodied intelligent body includes the human-computer interaction device based on the embodied intelligent body as described in the second aspect.
[0024] In a fourth aspect, the present application provides a computer device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the computer program, the human-computer interaction method based on embodied intelligent body as described in the first aspect above is implemented.
[0025] In a fifth aspect, the present application provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the human-computer interaction method based on an embodied intelligent agent as described in the first aspect above.
[0026] In a sixth aspect, the present application provides a chip, comprising a processor and a communication interface, wherein the communication interface is coupled to the processor, and the processor is used to run programs or instructions to implement the human-computer interaction method based on embodied intelligent body as described in the first aspect above.
[0027] In a seventh aspect, the present application provides a computer program product, including a computer program, which, when executed by a processor, implements the human-computer interaction method based on embodied intelligent body as described in the first aspect above.
[0028] Additional aspects and advantages of the present application will be given in part in the description below, and in part will become apparent from the description below, or will be learned through the practice of the present application. BRIEF DESCRIPTION OF THE DRAWINGS
[0029] The above and / or additional aspects and advantages of the present application will become apparent and easily understood from the description of the embodiments in conjunction with the following drawings, in which:
[0030] Figure 1 is a schematic diagram of an application scenario of a human-computer interaction method based on an embodied intelligent agent provided in some embodiments of the present application;
[0031] Figure 2 is a flowchart of a human-computer interaction method based on an embodied intelligent agent provided in some embodiments of the present application;
[0032] Figure 3 is a schematic diagram of the principles of the interaction scenarios provided in some embodiments of the present application;
[0033] Figure 4 is a schematic diagram of a process of constructing an object memory library provided in some other embodiments of the present application;
[0034] Figure 5 is a schematic diagram of the principle of constructing an object memory bank provided in some other embodiments of the present application;
[0035] Figure 6 It is a schematic diagram of another interactive scenario provided in some embodiments of the present application;
[0036] Figure 7 is a schematic diagram of the structure of a human-computer interaction device based on an embodied intelligent body provided in some embodiments of the present application;
[0037] Figure 8 It is a schematic diagram of the structure of a computer device provided in some embodiments of the present application. DETAILED DESCRIPTION
[0038] The following will be combined with the drawings in the embodiments of the present application to clearly describe the technical solutions in the embodiments of the present application. Obviously, the described embodiments are part of the embodiments of the present application, rather than all the embodiments. All other embodiments obtained by ordinary technicians in this field based on the embodiments in the present application belong to the scope of protection of this application.
[0039] Unless otherwise defined, all technical and scientific terms used in this application have the same meanings as those commonly understood by technicians in the technical field of this application; the terms used in the specification of this application are only for the purpose of describing specific embodiments and are not intended to limit this application; the terms "including" and "having" in the specification and claims of this application and the above-mentioned drawings and any variations thereof are intended to cover non-exclusive inclusions. The terms "first", "second", etc. in the specification and claims of this application or the above-mentioned drawings are used to distinguish different objects, rather than to describe a specific order or a primary and secondary relationship.
[0040] Reference to "embodiments" in this application means that a particular feature, structure, or characteristic described in conjunction with the embodiments may be included in at least one embodiment of the present application. The appearance of the phrase in various locations in the specification does not necessarily refer to the same embodiment, nor is it an independent or alternative embodiment that is mutually exclusive with other embodiments. It is explicitly and implicitly understood by those skilled in the art that the embodiments described in this application may be combined with other embodiments.
[0041] In the description of this application, it should be noted that, unless otherwise clearly specified and limited, the terms "installed", "connected", "connected", and "attached" should be understood in a broad sense, for example, it can be a fixed connection, a detachable connection, or an integral connection; it can be a direct connection, or an indirect connection through an intermediate medium, or it can be the internal communication of two elements. For ordinary technicians in this field, the specific meanings of the above terms in this application can be understood according to specific circumstances.
[0042] The term "and / or" in this application is only a description of the association relationship of associated objects, indicating that there can be three relationships. For example, A and / or B can represent: A exists alone, A and B exist at the same time, and B exists alone. In addition, the character " / " in this application generally indicates that the associated objects before and after are in an "or" relationship.
[0043] The term "multiple" as used in the present application refers to more than two (including two). Similarly, the term "multiple groups" refers to more than two groups (including two groups), and the term "multiple sheets" refers to more than two sheets (including two sheets).
[0044] In the existing technology, embodied agents still have certain deficiencies in task understanding and execution due to the following reasons: First, the ability to associate tasks with objects is insufficient: embodied agents are usually unable to accurately understand task requirements and the objects involved through natural language descriptions input by users, resulting in inefficient task decomposition and execution. Second: limitations of scene perception: in dynamic scenes, the location information and attributes of objects may change over time, and the existing technology has limited perception and memory capabilities for scenes, making it difficult to provide real-time and comprehensive object information support. Third: insufficient versatility in action selection: for different task requirements, the action planning of embodied agents often relies on predefined rules, lacks flexibility and adaptability, and is difficult to meet diverse scene requirements.
[0045] In view of this, the embodiments of the present application provide a human-computer interaction method, device and embodied intelligent agent based on embodied intelligent agent, which can dynamically analyze user instructions, flexibly match interactive tasks with target objects, and realize accurate action execution based on scene memory. By introducing the object memory library and candidate action matching mechanism, the task execution efficiency and adaptability of the intelligent system can be significantly improved, thereby better meeting user needs.
[0046] Among them, embodied intelligent agents refer to intelligent agents that are equipped with artificial intelligence and have the ability to perceive and interact with the outside world. They can interact with the environment in real time through perception and interaction, and can perform various tasks in virtual environments or real physical environments.
[0047] The embodiment of the present application also proposes an embodied video agent (Embodied VideoAgent), which can process videos to explore scenes, make decisions and plans, and perform actions.
[0048] The human-computer interaction method, device and embodied intelligent body based on embodied intelligent body provided in the embodiments of the present application can be applied in various fields, such as smart home, security monitoring, smart assistant, automatic driving and games.
[0049] In the following, in conjunction with the accompanying drawings, the human-computer interaction method based on embodied intelligent body provided by the embodiment of the present application is described in detail through specific embodiments and their application scenarios.
[0050] The human-computer interaction method, device and embodied intelligent agent provided in the embodiments of the present application can be applied to Figure 1In the application environment shown. Among them, the embodied intelligent body 10 is placed in a scene, which can be a real physical scene or a virtual scene, such as a simulator scene, a game scene, etc. The embodied intelligent body 10 can be provided with a camera and a depth sensor, etc., for detecting the external scene. The embodied intelligent body 10 can explore the scene by itself to identify and remember the objects in the scene. In some embodiments, the embodied intelligent body 10 can also analyze based on a given video recording of the scene to identify and remember the objects in the scene. For example, the embodied intelligent body 10 senses the objects in the target scene by acquiring the target video and the collected embodied sensor data, and obtains the perception data of at least one object. Then, an object memory library corresponding to the target scene is established, and the perception data of at least one object is stored in the object memory library by entries. In this way, the initial recognition and memory of each object in the scene is achieved. Of course, the embodied intelligent body 10 can also explore the environment by moving, and continuously take images / record videos in the process, and then identify and memorize objects based on each image / video frame. When the target scene changes, for example, the embodied intelligent agent 10 moves, causing the video frame captured by the camera to change, or the object in the target scene is moved, the embodied intelligent agent 10 determines the target object related to the target scene change and updates the object memory based on the target object.
[0051] Based on the constructed object memory library, the embodied intelligent agent 10 is able to perform interactions with the user. When the embodied intelligent agent 10 receives a description text sent by the user for a target scene, the embodied intelligent agent 10 responds to the description text, determines the interactive task indicated by the description text, and determines the target object associated with the interactive task. The embodied intelligent agent 10 searches the pre-established object memory library corresponding to the target scene, obtains the target perception data corresponding to the target object, and based on the target perception data, determines at least one target action from a plurality of pre-configured candidate actions. Then, the embodied intelligent agent 10 performs the target action on the target object to complete the interactive task.
[0052] The embodiment of the present application provides a human-computer interaction method based on an embodied intelligent agent. The execution subject of the human-computer interaction method based on an embodied intelligent agent can be an embodied intelligent agent or a functional module or functional entity in the embodied intelligent agent that can implement the human-computer interaction method based on an embodied intelligent agent. The human-computer interaction method based on an embodied intelligent agent provided in the embodiment of the present application is described below by taking the embodied intelligent agent as an execution subject.
[0053] like Figure 2 As shown, the human-computer interaction method based on embodied intelligent body includes: steps 210 to 250.
[0054] Step 210: In response to a description text sent by a user for a target scene, determine an interactive task indicated by the description text.
[0055] The target scene refers to the spatial range recorded by the target video and perceived by the embodied agent, including objects, structures and their relationships within the range. The target scene is the core background for the embodied agent to analyze and make decisions. The embodied agent builds an object memory based on the perception data in the target scene, thereby realizing dynamic monitoring of the scene and object management. The target scene can vary depending on the specific application, for example, it can be a scene in a room, a section of road, or a storage area.
[0056] Description text refers to the content input by the user in natural language, which is used to indicate the interactive task that the user wants to perform in a specific scene. Description text usually contains information about the interactive task, target object or expected result. For example, in a home scene, the user inputs "put the red cup on the table into the cupboard" through voice or text, and this sentence is the description text. The embodied intelligent agent extracts the interactive task, target object and its associated interactive requirements indicated by the user by parsing the description text. The interactive task is the specific operation requirement generated after the description text is parsed, which indicates the function or action that the user wants the intelligent system to perform in the target scene. Interaction tasks are usually associated with specific target objects and actions. For example, the embodied intelligent agent parses the description text "put the red cup on the table into the cupboard" and identifies that the interactive task it contains is an interactive task of the type "place an object in a specified position".
[0057] Specifically, the embodied intelligent agent can be equipped with a variety of sensors or voice collection devices to monitor in real time whether the user has transmitted the description text. The user can send the description text to the embodied intelligent agent in the form of voice or input text for a target scene. The embodied intelligent agent parses the description text to determine the interactive task indicated by the description text, that is, the task that the user wants the embodied intelligent agent to perform.
[0058] The interactive task can be a question-and-answer task, in which the user asks a question and the embodied agent gives a corresponding answer to the question based on the target scenario. For example, if the user enters a description text: "Where is my water cup?", the embodied agent determines the location of the "water cup" by retrieving the object memory based on the target scenario and gives a corresponding answer.
[0059] Interaction tasks can also be execution tasks, in which the user gives instructions and the embodied agent performs corresponding actions based on the instructions. For example, if the user enters the description text: "Please put the bowl next to the pool", the embodied agent will retrieve the locations of "bowl" and "pool" in the object memory based on the target scene and perform corresponding actions, such as going to the location of the "bowl" and operating the robotic arm to pick up the "bowl", then going to the location of the "pool" and operating the robotic arm to place the "bowl" next to the "pool", and so on.
[0060] Step 220: Determine the target object associated with the interactive task.
[0061] That is, the embodied agent identifies the target object that the user wants to operate in the interactive task parsed from the description text. The target object is usually the core operation object of the task, and its characteristics (such as name, location, color, etc.) are the key input for the subsequent task execution.
[0062] Specifically, the embodied agent can parse the description text input by the user and extract text features through the semantic understanding model, and then determine the target object related to the interactive task based on the text features. For example, for the description text "put the book on the table", the target object extracted by the embodied agent is, for example, "book", or it can also be "book" and "table". In this process, the embodied agent can also combine the task reasoning mechanism to more accurately understand the user's intention and determine the target object.
[0063] Exemplarily, the semantic understanding model includes, but is not limited to, one or more of a large language model, a visual understanding model, a Transformer model, and the like.
[0064] Step 230: Search in a pre-established object memory library to obtain target perception data corresponding to the target object; the object memory library is used to store the perception data of all objects pre-identified from the target scene.
[0065] Among them, the object memory is a data structure dedicated to storing and managing information related to objects in the target scene. Its storage content includes the object's perception data (such as location, shape, size and other attributes) and the object's historical records (such as appearance time, interaction records, dynamic changes, etc.). The memory is usually organized in the form of entries, each entry corresponds to an object or a class of objects, and a unique identifier is assigned to each object.
[0066] Exemplarily, each entry includes but is not limited to the following fields: object identification, object category, object state description, positional relationship between objects, three-dimensional boundary data of the object, and visual data of the object.
[0067] Object identifiers are usually automatically assigned numbers by embodied intelligent agents, or generated by the object's own attributes (for example, its color, shape, label and other features can be combined) to ensure the distinction between different objects in the scene and the management of corresponding memory, and to provide a reference for subsequent dynamic tracking and retrieval interactions of objects.
[0068] Object categories are classifications of object properties, usually based on the function, form, or purpose of the object. In some embodiments, object categories can be automatically identified by a machine learning algorithm carried by the embodied intelligent agent.
[0069] The state description of an object refers to a comprehensive expression of the state of the object at a certain point in time, which may include but is not limited to position changes, physical forms, etc. For example, the state description of a refrigerator may include "normal", "refrigerator door open" or "refrigerator door closed", and the state description of an object may include "on the table", "in hand" or "not in hand", etc. Exemplarily, the state description of an object may be set to "normal" by default at the initial time (for example, when it is first recognized), and may be updated according to real-time changes in the scene.
[0070] The positional relationship between objects refers to the spatial description of the relative positions of different objects in the target scene, including distance, orientation, contact relationship, etc. For example, object A is placed on the upper surface of object B, object C is located inside object D, or object E is located on the right side of object F and the distance is 1 meter, etc. The positional relationship between the objects can be recorded and stored in the form of a list, for example.
[0071] The three-dimensional boundary data of an object is an accurate description of the shape, size, and spatial occupancy range of the object in three-dimensional space, and is usually represented by a three-dimensional coordinate point cloud or voxel data.
[0072] The visual data of an object refers to the feature information of the object extracted from the image / video frame, including but not limited to the image / video frame or a part thereof (such as an image area where the object is located) and feature vectors.
[0073] Of course, it is not limited to the above data. The embodied intelligent agent can store different types of perception data according to the needs of the actual scenario.
[0074] The embodied agent accesses the object memory associated with the target scene and retrieves the pre-stored object perception data to obtain the perception data of the target object. If there is no directly matching object in the object memory, the system may prompt the user for further confirmation. If multiple candidate objects are retrieved from the object memory, the embodied agent can accurately match them through multi-feature matching or reasoning. In this process, the embodied agent can also combine the perception data of the target object to further confirm the target object indicated by the description text.
[0075] For example, when a user sends a description text "pass me the red cup on my table", the embodied agent may retrieve the perception data corresponding to multiple red cups in the object memory bank, but combined with the location clue of "on the table", the embodied agent can determine which red cup it is.
[0076] For example, when a user sends a description text "put the cup away", the embodied agent can reason based on the temporary memory buffer or scene context, for example, the cup may be the cup that the user just put aside a few seconds ago, rather than an unused cup in the cupboard.
[0077] The temporary memory buffer includes an action buffer and an object buffer. The action buffer is used to record the action detected by the embodied intelligent agent and the timestamp of the action, the name of the action, the object as the action target (for example, the object identifier), and the visual data of the current frame when the action is detected (including object features and context features). The object buffer is used to record each recognized object and its recognition timestamp, object identifier, and three-dimensional boundary data. As a result, temporary memory can be quickly referenced in subsequent tasks, further enhancing the embodied intelligent agent's understanding of dynamic scenes and improving task execution efficiency.
[0078] Step 240: Determine at least one target action from a plurality of pre-configured candidate actions based on the target perception data.
[0079] Candidate actions are a set of standardized operation templates pre-configured by the system to cope with common task requirements in different scenarios. Exemplarily, the executable candidate actions pre-configured by the embodied agent include, but are not limited to, one or more of: search, goto, open, close, pick up, or place. Associated with each candidate action is one or more physical operations, such as "move", "rotate the robotic arm", etc. The target action is an operation step determined from the candidate actions and specifically performed on the target object to complete the interactive task. For example, for the description text of "grab the red cup", the target action can be, for example, "pick up".
[0080] Specifically, the embodied intelligent agent can preliminarily select one or more target actions that match the task requirements from the candidate actions based on the perception data of the target object, such as matching the "pick up" action for the cup, matching the "open" or "close" action for the refrigerator door, etc. The embodied intelligent agent also determines the location of the target object in the target scene based on the perception data of the target object, and matches actions such as "go to" or "navigate" to go to the corresponding location and perform the matched "pick up" action. Therefore, by utilizing multimodal perception data, the accuracy of action planning can be improved, and the specific target action to be performed can be dynamically adjusted through real-time perception of the environment.
[0081] Step 250: Execute the target action on the target object to complete the interactive task.
[0082] That is, the embodied agent completes the interactive task specified by the user's instructions by performing the determined target actions on the target object based on the results of previous parsing and planning.
[0083] For example, for question-answering tasks, if the user asks the question "Where is my water cup?", the embodied agent determines the location of the "water cup" by retrieving the object memory based on the target scene and gives a corresponding answer. The target action is the "chat" action, etc.
[0084] For example, for execution tasks, the user gives the instruction "Please put the bowl next to the sink", then the embodied intelligent agent, based on the target scene, retrieves the locations of "bowl" and "sink" in the object memory bank, determines the target action as "go to the location of the "bowl", "pick up the "bowl", then "go to the location of the "sink", and "place the "bowl" next to the sink".
[0085] For example, Figure 3 As shown, the user can send instructions by voice or manually inputting description text. The embodied intelligent agent analyzes the description text and determines the target object associated with the description text. For example, for the user's description text "What's on the microwave oven", the embodied intelligent agent can determine that the target object is "microwave oven"; for another example, for the user's description text "Can you put the bowl by the sink", the embodied intelligent agent can determine that the target objects are "bowl" and "sink". Then, the embodied intelligent agent searches the object memory library and obtains the perception data corresponding to one or more target objects, which are called target perception data for distinction.
[0086] As a result, the embodied agent can perform the target action on the target object based on the retrieved target perception data. For example, when the embodied agent performs the interactive task of "putting the bowl next to the pool", it determines the location of each according to the perception data of the "bowl" and "pool" in the object memory, and performs path planning and navigation from the current location to the location of the "bowl". After that, the embodied agent performs the "pick" operation on the object "bowl", moves it to the side of the "pool", and returns the corresponding task completion feedback to the user.
[0087] The human-computer interaction method based on embodied intelligent agent provided in the embodiment of the present application can respond to and analyze the user's multimodal input by responding to the description text sent by the user for a target scene, and can improve the user's interactive experience; by determining the target object associated with the interactive task, and searching in a pre-established object memory library, the target perception data corresponding to the target object is obtained and introduced into the object memory library, and the object memory library is introduced to persistently remember the objects in the target scene, so that the embodied intelligent agent can dynamically adapt to scene changes and accurately locate the target object, thereby improving the completion and accuracy of the interactive task; and then based on the target perception data, at least one target action is determined from a plurality of pre-configured candidate actions, and the target action is executed on the target object to complete the interactive task, and the action matching mechanism enables the embodied intelligent agent to adapt to the diverse task requirements in different scenarios, significantly improving the task execution efficiency and adaptability of the embodied intelligent agent, thereby better meeting user needs.
[0088] In scenarios where embodied agents interact with users, how to accurately understand the user's intentions and determine the target object to be operated is the key to improving task execution efficiency and interaction quality. Especially in complex scenarios, user instructions may contain multiple clues related to the target object (such as location, color, status, etc.), and even involve locating the target object through the orientation clues provided by the associated object.
[0089] To this end, in some embodiments, determining the target object associated with the interactive task includes: performing semantic understanding on the description text to determine the first target object indicated by the description text; the first target object is the target object for the target action to be performed; determining whether the description text indicates a second target object associated with the first target object; the second target object includes a target object that provides a location clue for the first target object to perform the target action; the location clue is included in the description text; if the description text indicates a second target object, then based on the first target object and the second target object, determining the target object associated with the interactive task.
[0090] Specifically, the embodied intelligent agent parses the first target object indicated in the description text based on the description text input by the user, so as to perform semantic understanding based on a single object. For example, if the user inputs the description text "Where is my coat", the embodied intelligent agent parses the first target object in the description text as "coat", and does not parse the second target object associated with the "coat", then the embodied intelligent agent takes the first target object as the final target object.
[0091] If the embodied agent parses out that the description text also indicates a second target object associated with the first target object, then the first target object and the second target object are both taken as target objects. For example, the user inputs the description text "bring me the blue cup on the table", then the embodied agent parses out that the first target object in the description text is "blue cup", the second target object is "table", and the second target object provides the location clue of the first target object "placed on..." Therefore, the embodied agent takes both "blue cup" and "table" as target objects associated with the interactive task.
[0092] In the above embodiment, by combining the first target object and the second target object related to it, the target object can be accurately determined in a complex scene, the accuracy of object recognition is improved, and the efficiency and accuracy of the execution of the interactive task are improved; and by using the clues provided by the second target object for auxiliary positioning, the adaptability of the embodied intelligent body to complex dynamic scenes is enhanced.
[0093] In intelligent interactive systems, in scenarios based on multi-object perception and action planning and execution, the positional relationship between multiple target objects is crucial for action decision-making. Traditional methods often only focus on the characteristics of a single object and ignore the relative relationship between multiple objects, which may lead to inefficient action decision-making or execution errors. Especially in complex tasks involving multi-target operations, if there is no clear action planning and execution sequence, it may cause task failure or waste of resources.
[0094] To this end, in some embodiments, based on the target perception data, at least one target action is determined from a plurality of pre-configured candidate actions, including: determining the positional relationship between the first target object and the second target object based on the first perception data and the second perception data; determining a first target action corresponding to the first target object and an action execution order from a plurality of pre-configured candidate actions based on the positional relationship and the first perception data; determining a second target action corresponding to the second target object and an action execution order from a plurality of pre-configured candidate actions based on the positional relationship and the second perception data; and executing the first target action and the second target action in accordance with the action execution order.
[0095] Among them, the target perception data includes first perception data corresponding to the first target object and second perception data corresponding to the second target object. As mentioned in the above embodiments, the perception data in the object memory library includes feature data (object features and context features), object state description and positional relationship with other objects, etc. Regardless of whether there is a second target object to provide orientation clues, the embodied intelligent agent can determine the existence status of the target object in the target scene based on these data, for example, indexing and determining the position of the target object in the target scene based on the feature data and positional relationship, and confirming the state of the target object based on the object state description, etc. If there is a second target object to provide orientation clues, the embodied intelligent agent can further confirm and judge the existence status of the first target object in the target scene in combination with the second perception data.
[0096] Thus, the embodied intelligent agent determines the first target action and the action execution order corresponding to the first target object, and the second target action and the action execution order corresponding to the second target object, respectively, based on the first perception data and the second perception data, and the positional relationship between the first target object and the second target object, taking into account the object relationship and the temporal dependency of the action. If there is a logical dependency between the actions (such as grabbing first and then placing), the embodied intelligent agent will give priority to executing the action that meets the dependency condition. If the actions can be performed in parallel, the embodied intelligent agent can perform parallel operations to improve efficiency.
[0097] There may be one or more first target actions corresponding to the first target object, and there may also be one or more second target actions corresponding to the second target object. Each target action corresponds to an action execution order, such as first executing the first target action A1, then executing the second target action B1, and then executing the first target action A2, etc.
[0098] Thus, the embodied agent gradually performs corresponding operations on the target object according to the planned target actions and execution sequence. During the execution process, the task status is monitored in real time through sensors, and the action path or sequence is dynamically adjusted according to the feedback signal to ensure the completion of the task. If a deviation is detected during the execution (such as a change in the position of the target object or execution failure), the embodied agent can, for example, recalculate the action plan and iterate until the task is completed.
[0099] In the above embodiment, by utilizing the multimodal perception data in the object record library, combined with the pre-configured action library, action selection and planning are performed for task execution, and the target action is executed in the planned order, it is possible to achieve efficient completion of interactive tasks in the case of multiple objects being associated, and it can be applicable to complex dynamic scenes, thereby improving the execution accuracy and efficiency of interactive tasks.
[0100] Similar to the above embodiment, based on the target perception data, at least one target action is determined from a plurality of pre-configured candidate actions, and further includes: based on the target perception data, determining a target action corresponding to the target object and an action execution order from a plurality of pre-configured candidate actions; and executing the target action according to the action execution order. The specific steps can be referred to the above embodiment, and will not be repeated here.
[0101] The object memory plays a key role in task execution. In some embodiments, Figure 4 As shown, the method further includes steps 410 to 430:
[0102] Step 410: Acquire a target video and embodied sensor data corresponding to the target video; the target video is used to present a target scene, including at least one video frame;
[0103] Step 420: Determine, based on at least one video frame and the embodied sensor data, the perception data of the embodied intelligent agent for at least one object in the target scene;
[0104] Step 430: Establish an object memory library corresponding to the target scene, and store the perception data of at least one object in the object memory library.
[0105] The target video is the video data used to present the target scene, which is usually composed of multiple consecutive frames of images. A video frame is the basic unit of a video, and each frame is a static image, representing a picture captured by the video at a certain moment.
[0106] In an embodiment of the present application, the target video provides visual information of the target scene. By processing multiple video frames, it is possible to extract information such as object features, spatial layout, dynamic changes, etc. in the scene, thereby assisting the embodied intelligent body in environmental understanding and decision-making.
[0107] Exemplarily, the target video may be a first-person video. The embodied agent may be an embodied video agent.
[0108] Embodied sensor data refers to the sensory data collected by various sensors on the embodied intelligent body, including but not limited to physical, chemical or spatial attribute information, such as distance, temperature, vibration, sound or smell, etc. Embodied sensor data can be combined with video data to provide multimodal data support, thereby improving the accuracy of object recognition and feature analysis in the target scene, and helping the embodied intelligent body to interact with the environment and perceive dynamic changes, such as detecting the movement of objects.
[0109] Exemplarily, the embodied sensing data includes depth data and camera 6D pose data. Depth data refers to information reflecting the distance between an object or surface in a scene and a sensor, usually expressing the depth value of each point in the form of pixels. Depth data can be obtained by a depth camera (such as LiDAR, ToF camera, or structured light camera), or calculated from multi-view images by a stereo vision algorithm. The camera 6D pose data describes the comprehensive information of the position (3 degrees of freedom) and direction (3 degrees of freedom) of the camera in three-dimensional space, including the three-dimensional position of the camera in a reference coordinate system, usually represented by a three-dimensional vector; and the posture of the camera in the reference coordinate system, that is, its shooting direction and rotation state, usually represented by a rotation matrix, Euler angles, or quaternions. The reference coordinate system is, for example, the world coordinate system.
[0110] The embodied intelligent agent obtains the target video by real-time shooting / recording through its own visual acquisition device, or by receiving a given video. The visual acquisition device includes but is not limited to a camera.
[0111] Exemplarily, the embodied intelligent agent extracts a historically recorded video from a local storage space as a target video, downloads a video from the Internet as a target video, or receives a video transmitted by a user or other device as a target video, etc.
[0112] The embodied agent recognizes objects in the target scene based on the video frames in the target video and the embodied sensory data corresponding to each video frame, and extracts the sensory data of the objects in the target scene. The objects may be biological objects or non-biological objects, such as tables, refrigerators, cats, humans, or other robots.
[0113] The embodied sensor data corresponding to each video frame refers to, for example, the embodied sensor data collected at the time corresponding to the corresponding video frame.
[0114] For example, in an indoor environment, the embodied intelligent agent can identify objects such as tables, chairs, and cups through video frames, and further identify the positions and sizes of these objects, as well as the positional relationships between these objects by combining video frames and embodied sensor data, for example, the chair is next to the table and the cup is placed on the table.
[0115] The embodied agent associates these data with the corresponding objects and stores them in the object memory library in separate entries to achieve persistent memory of each object in the scene.
[0116] In the above embodiment, by acquiring the target video and the collected embodied sensor data, and performing multimodal data processing based on the target video and the embodied sensor data, each object in the target scene is identified, and the corresponding perception data is extracted, and then stored in an object memory library. By constructing an object memory library, persistent memory of objects in the scene is achieved, so that multimodal perception data can be stored and managed in a structured manner, and long-term tracking of objects can be achieved through dynamic maintenance, which facilitates the efficient execution of subsequent analysis and decision-making; when the target scene changes, the object memory library can be updated in real time based on the target objects related to the change, so that the embodied intelligent agent can quickly identify and respond to changes in the target scene, thereby enhancing the adaptability of the embodied intelligent agent in complex scenes.
[0117] In a specific example, Figure 5 As shown, taking the intelligent robot as an embodied video agent as an example, the embodied video agent is equipped with a camera and a depth detection device, etc. The embodied video agent can combine the depth data, the camera 6D posture and the video data for multimodal processing, identify each object in the target scene, and extract the perception data of each object. For example, the embodied video agent performs dimensionality-upgrading processing on the two-dimensional bounding box obtained by object detection from the video frame by combining the depth data and the camera 6D posture data to obtain the three-dimensional bounding box of the object. The specific processing flow can refer to the above embodiment, which will not be repeated here. Thus, the embodied video agent stores the perception data of each object in the object memory. Exemplarily, the object memory is distinguished by the object identification (ID), and the state description (STATE) of the object is "normal" at the beginning. The state of the object can be updated synchronously based on the real-time detection of dynamic changes in the scene. For example, the state description of the object "microwave oven" is updated from "normal" to "closed" to indicate that the door of the microwave oven is in a closed state. In addition, the object memory also stores the positional relationships (RO) between objects. For example, if paper roll O0 is placed on microwave oven O1, the corresponding RO is (on, O1); and if microwave oven O1 supports paper roll O0, the corresponding RO is (hold, O0). When an action is detected in the scene, the embodied video agent also updates the object memory according to the objects related to the action, thereby accurately tracking dynamic scenes or dynamic changes of objects.
[0118] In some embodiments, when the target scene changes, the embodied intelligent agent determines the target object related to the target scene change and updates the object memory based on the target object.
[0119] The change in the target scene may be a change in an object within the target scene, such as an object being moved or a new object being added; or it may be a change caused by a change in visual acquisition, such as the movement of an embodied intelligent body or a change in the position of a visual acquisition device carried by it.
[0120] When the target scene changes, the embodied intelligent agent can collect new images / video frames and identify the objects in the new images / video frames, determine whether they are memorized objects or new objects, or update the existence status of memorized objects in the scene.
[0121] Specifically, the embodied agent determines the target object associated with the change by comparing the newly collected data with the data in the object memory, and updates the object memory. For example, when a chair is detected to move from a corner of the room to the side of the table, the target object associated with the change is the chair, and the embodied agent will update the position information of the chair to reflect the latest status.
[0122] For example, the embodied intelligent agent can detect the difference between frames by comparing the continuous video frames in the target video, identify the area in the scene that has changed, and thus determine the target object related to the change of the target scene. Alternatively, the embodied intelligent agent can also analyze the changes in the relative position, posture or distance between objects through other sensor information such as depth data and camera posture data, so as to determine the target object related to the change of the target scene. For another example, the embodied intelligent agent can determine which objects have changed their state due to the interaction by monitoring the interaction between objects (such as collision, contact, occlusion, etc.).
[0123] In some embodiments, the embodied intelligent agent can understand and analyze the scene through machine learning or deep learning models, thereby identifying dynamic objects in the scene and determining whether the perception data of these objects in the object memory has changed, such as whether its state description has changed.
[0124] In some embodiments, based on at least one video frame and embodied sensor data, determining perception data of the embodied intelligent agent for at least one object in the target scene includes steps 510 to 540:
[0125] Step 510: for any video frame, identify at least one object in the targeted video frame, and obtain feature data, attribute data and two-dimensional boundary data of the at least one object;
[0126] Step 520: Perform dimensionality upscaling on the two-dimensional boundary data using the embodied sensor data corresponding to the targeted video frame to obtain three-dimensional boundary data corresponding to at least one object;
[0127] Step 530: determining a positional relationship between at least one object based on the three-dimensional boundary data corresponding to the at least one object;
[0128] Step 540: Based on the feature data, attribute data, three-dimensional boundary data and positional relationship, obtain the perception data of the embodied intelligent agent for at least one object in the target scene.
[0129] Among them, feature data includes visual features. Attribute data includes but is not limited to object identification, object category, object state description, etc. Two-dimensional boundary data is, for example, coordinate data of a two-dimensional detection box of an object. Specifically, for each video frame, the embodied intelligent agent detects the object included in the video frame by performing feature extraction, and extracts its feature data, attribute data, and two-dimensional boundary data through image processing algorithms, etc. Exemplarily, the embodied intelligent agent can extract image features in the video frame through its onboard target detection algorithm, etc., and classify based on the image features, thereby obtaining the object identification and object category of the object in the video frame.
[0130] In order to more accurately identify the existence status of objects in the target scene, the embodied intelligent agent also obtains the embodied sensor data corresponding to the video frame, including but not limited to depth maps and camera 6D pose data.
[0131] Furthermore, based on the embodied sensor data, the embodied intelligent body can map the position of the object in the two-dimensional image to the three-dimensional physical space, thereby determining the existence state of the object in the three-dimensional scene. Specifically, the embodied intelligent body performs dimensionality upscaling on the two-dimensional boundary data through the embodied sensor data to obtain the three-dimensional boundary data of the object. The three-dimensional boundary data is, for example, the coordinate data of the three-dimensional detection box of the object.
[0132] In some embodiments, the embodied intelligent agent can extract the two-dimensional detection frame of the object from the video frame through a machine learning model or a deep learning model. On this basis, the two-dimensional boundary data is combined with the depth information in combination with the depth map and the camera 6D pose data, and the pixel points in each two-dimensional boundary box are depth matched to obtain the three-dimensional space coordinates of each pixel. For example, the embodied intelligent agent can map the two-dimensional coordinates to the three-dimensional space using the perspective projection formula through the intrinsic and extrinsic parameters of the camera and the depth value of each pixel. Furthermore, the embodied intelligent agent can generate the three-dimensional boundary data of the object by calculating the position of each pixel point in the three-dimensional space.
[0133] Therefore, for the video frame, the embodied intelligent agent can judge the spatial relationship between the object and other objects based on the three-dimensional boundary data of the object and the three-dimensional boundary data of other objects, such as the distance between object A and object B, the upper and lower position relationship between object C and object D, etc.
[0134] Finally, the embodied intelligent agent can generate the perceptual data of the object based on the above data for memory and storage. For example, the embodied intelligent agent can directly memorize the above data as perceptual data, or the embodied intelligent agent can further process the above data, such as deduplication and noise reduction, and then memorize the processed data as perceptual data. For example, the three-dimensional boundary data obtained after dimensionality increase is often affected by noise or errors. The embodied intelligent agent can process the three-dimensional boundary data through geometric optimization algorithms (such as point cloud fitting, boundary smoothing, etc.) to make it more consistent with the actual shape of the object, and so on.
[0135] In the above embodiments, by combining video frames and embodied sensor data and performing processing based on multimodal data, a more accurate and comprehensive perception of objects is achieved, thereby improving the embodied intelligent agent's perception of space and objects in the space, enabling it to better understand and interact with its environment, thereby performing more intelligent tasks.
[0136] In the process of feature extraction of video frames, the embodied agent can not only identify the position of each object in the image, but also fully understand the nature and environment of the target object through the combination of object features and context features, thereby perceiving the object more accurately.
[0137] To this end, in some embodiments, for any video frame, identifying at least one object in the targeted video frame and obtaining feature data of the at least one object includes steps 610 to 640:
[0138] Step 610: for any video frame, identifying the image region where at least one object in the targeted video frame is located;
[0139] Step 620: extract features from the image regions where at least one object is located, and obtain object features corresponding to the at least one object.
[0140] Step 630: extract features from the targeted video frame to obtain context features corresponding to at least one object;
[0141] Step 640: Obtain feature data of at least one object in the targeted video frame based on the object features and the context features.
[0142] Specifically, for any video frame, the embodied agent uses an object detection model (such as YOLO or Faster R-CNN based on convolutional neural networks) to detect objects in the video frame and obtain at least one object and its image region in the video frame, such as a part of the video frame. In the image region where the object is located, the embodied agent uses a visual feature extraction algorithm (such as ResNet, CNN, etc.) to extract the visual features of the object, which are called object features.
[0143] In addition to the features of the object itself, the embodied intelligent agent also extracts features from the entire video frame to obtain contextual features related to the object. Contextual features can reflect the position of the object in the scene, etc. Thus, the embodied intelligent agent can generate complete feature data of the target object based on the object features and contextual features. These data combine the information of the object itself and its relationship information in the scene, which is convenient for subsequent storage, analysis and decision-making. In some embodiments, the embodied intelligent agent uses the object features and contextual features as feature data of the object and stores them in the object memory.
[0144] In the above embodiment, by performing multi-level feature extraction and fusion of the visual information and scene context information of the object in the target video frame, the target object can be accurately identified and its rich feature data can be obtained, thereby improving the accuracy of object recognition; and, by combining object features with context features, the embodied intelligent agent can more comprehensively perceive the object's ontological properties and its relative relationship in the scene, thereby improving the accuracy and robustness of object recognition.
[0145] It should be noted that context features can reflect the position and state of an object in the world, and can provide reliable data support when the embodied agent performs interactive tasks. For example, the embodied agent can answer the user's questions such as "Where is object A?" based on the object features and context features.
[0146] When the target scene changes, the embodied agent can understand the changed scene based on the object memory that has been built. When the target scene changes, the embodied agent can not only timely identify unknown objects and classify them, but also update the existing data in the object memory to ensure the accuracy and completeness of the object perception data. At the same time, relying on the object memory, the embodied agent can also accurately identify objects that have moved.
[0147] To this end, in some embodiments, when the target scene changes, a target object related to the target scene change is determined, and the object memory is updated based on the target object, including steps 710 to 740:
[0148] Step 710: when an unknown object is identified in the target scene, the unknown object is re-identified to obtain a re-identification result;
[0149] Step 720: if the re-identification result indicates that the unknown object is the same as any identified object, the perception data corresponding to the identified object in the object memory is updated;
[0150] Step 730: If the re-identification result indicates that the unknown object is not any identified object, determine the perception data corresponding to the unknown object;
[0151] Step 740: Store the perception data of the unknown object into the object memory.
[0152] When the target scene changes, if the embodied agent detects an unknown object, it first uses re-identification technology to determine whether it matches the identified object. If the match is successful, it means that the unknown object is an identified object, and its position, shape, etc. may have changed. The embodied agent updates the perception data of the identified object in the object memory; if the match fails, it means that the unknown object may be a newly added object, and the perception data of the object is added as a new entry to the object memory. Therefore, through the dynamic update mechanism, the embodied agent can adapt to changes in the target scene and continue to maintain a comprehensive perception of the scene.
[0153] Specifically, the embodied agent can identify the set of objects in the target scene through the target detection algorithm. If an unknown object (i.e., an object that is not in the object memory) is detected, the re-identification process is triggered. For example, there was originally a red cup on the table in the target scene. When someone puts a green cup on the table, the embodied agent identifies the red cup and the green cup, and determines that the red cup is a recognized object, while the green cup is an unknown object.
[0154] For unknown objects, the embodied intelligent agent can re-identify the unknown object by judging the similarity between the visual features and spatial position of the unknown object and the recognized objects, thereby determining whether the unknown object is a recognized object or a new object.
[0155] If the re-identification result indicates that the unknown object is the same as an identified object, the perception data of the object in the object memory is updated, such as the object's position, state, or other attribute information. For example, if the embodied agent recognizes that the green cup matches a memorized green cup, the position and state description of the green cup is updated, such as its position is moved from the refrigerator to the table.
[0156] If the re-identification result shows that the unknown object does not match the object in the memory bank, it is added as a new entry to the object memory bank, and its perception data is stored at the same time, including feature data, attribute data, three-dimensional boundary data, and the positional relationship between the object and other objects.
[0157] In the above embodiment, through real-time recognition and updating, the embodied intelligent agent can quickly respond to the addition, removal or state change of objects in the target scene, ensuring the accuracy of the perceived data; and, by introducing a re-recognition algorithm, the embodied intelligent agent can distinguish between known objects and unknown objects, avoid repeated recording or omission of important information in the scene, and achieve persistent object memory and accurate tracking of the object state.
[0158] In the process of re-identifying objects, the embodiment of the present application also improves the accuracy of re-identifying unknown objects by constructing a re-identification mechanism of stereo similarity comparison. In addition, for unknown objects in the target scene, the identity of unknown objects and identified objects can be efficiently determined through differentiated processing strategies under static and dynamic scene conditions. To this end, in some embodiments, re-identifying unknown objects and obtaining re-identification results include steps 810 to 850:
[0159] Step 810, obtaining first three-dimensional boundary data of an unknown object, and extracting second three-dimensional boundary data corresponding to each identified object from an object memory library;
[0160] Step 820: Based on the first three-dimensional boundary data and each second three-dimensional boundary data, respectively perform a stereo similarity comparison to obtain a stereo similarity result;
[0161] Step 830: Determine whether the unknown object is a static object or a dynamic object;
[0162] Step 840: When the unknown object is a static object, if the stereo similarity result satisfies the first similarity condition, a first recognition result is obtained;
[0163] Step 850: When the unknown object is a dynamic object, if the stereo similarity result satisfies a second similarity condition, a second recognition result is obtained; wherein the first recognition result and the second recognition result indicate that the unknown object is the same object as a recognized object.
[0164] That is, the embodied agent obtains the first three-dimensional boundary data of the unknown object and compares it with the second three-dimensional boundary data in the object memory library to match the object. Specifically, the matching degree of the unknown object is judged by the stereo similarity of the three-dimensional shape. In addition, the embodied agent also uses different similarity conditions for re-identification according to the dynamic properties of the object (static object or dynamic object). Therefore, in the process of re-identifying the object, not only the shape characteristics of the object are considered, but also its dynamic properties are considered, which can enhance the adaptability of the specific agent to changing objects in dynamic scenes.
[0165] Specifically, the embodied intelligent agent obtains the three-dimensional boundary data of the unknown object. The specific steps can refer to the above embodiment. For the convenience of distinction, the three-dimensional boundary data of the unknown object is called the first three-dimensional boundary data, and the three-dimensional boundary data of the identified object in the object memory is called the second three-dimensional boundary data. The embodied intelligent agent then compares the first three-dimensional boundary data with each second three-dimensional boundary data in three-dimensional and four-dimensional dimensions to obtain a similarity result.
[0166] In addition, the embodied agent also determines whether the unknown object is a static object or a dynamic object. In some embodiments, the embodied agent can observe the historical motion trajectory of the position object based on the historical multi-frame video frame, and determine whether it is a static object or a dynamic object. For example, when the embodied agent detects that the red cup has not changed position in the past 10 seconds, it is determined to be a static object.
[0167] For unknown objects with different dynamic attributes, the embodied intelligent agent makes judgments based on different similarity conditions. That is, if the unknown object is a static object, the embodied intelligent agent determines whether the stereo similarity meets the first similarity condition. If so, a first recognition result is obtained, which indicates that the unknown object is the same object as a recognized object.
[0168] If the unknown object is a dynamic object, the embodied intelligent agent determines whether the stereoscopic similarity satisfies a second similarity condition. If so, a second recognition result is obtained, which indicates that the unknown object is the same object as an already recognized object.
[0169] It is easy to understand that if the unknown object is a static object and the stereoscopic similarity does not meet the first similarity condition, the embodied intelligent agent obtains a third-level recognition result, which indicates that the unknown object is not the same object as a recognized object. If the unknown object is a dynamic object and the stereoscopic similarity does not meet the second similarity condition, the embodied intelligent agent obtains a fourth-level recognition result, which indicates that the unknown object is not the same object as a recognized object.
[0170] Therefore, the embodied agent updates the sensory data in the object memory based on the re-identification results, including the location, state, and other dynamic attributes. If any similarity conditions cannot be met, the unknown object is treated as a new object and its sensory data is added to the object memory.
[0171] In the above embodiment, by comparing the stereo similarity based on the three-dimensional boundary data, it is possible to effectively distinguish unknown objects from identified objects to avoid misidentification or missed identification; and by introducing the classification processing of static and dynamic objects, the embodied intelligent agent can have a deeper understanding of the changes and interaction relationships of objects in the scene, thereby improving the accuracy and robustness of object recognition.
[0172] In some embodiments, for static objects, the first similarity condition includes: the degree of overlap exceeds a first threshold, or the maximum inclusion ratio exceeds a second threshold and the object categories are the same.
[0173] The Intersection over Union (IoU) represents the overlap between the 3D bounding boxes of an unknown object and an identified object. The greater the overlap, the more likely it is that the unknown object and the identified object are the same object. For example, the overlap can be expressed by the following formula (1):
[0174] (1)
[0175] in, is the intersection of their three-dimensional bounding boxes, is the union of their 3D bounding boxes.
[0176] The Maximum Ratio of Intersection over Subsets (MaxIoS) represents the inclusion relationship between the 3D bounding box of the unknown object and the 3D bounding box of the identified object. When the two bounding boxes show a strong inclusion relationship, MaxIoS will be close to its maximum value of 1, which means that the unknown object is likely to be the same object as the identified object. For example, the maximum inclusion ratio can be expressed by the following formula (2):
[0177] (2)
[0178] in, represents the volume of the unknown object, Indicates that an object has been recognized.
[0179] Assumptions and are all detected as "table", where The volume is one tenth of The 3D bounding box of If the object is within the 3D bounding box, MaxIoS will reach 1, while IoU will be only 0.1. By introducing the maximum inclusion ratio and combining it with the object category for discrimination, it is possible to re-identify some observable objects in the case of occlusion. For example, may be Some observations (such as is part of the table), since they have overlapping bounding boxes and belong to the same object category.
[0180] Since objects are changing dynamically, their 3D detection frames may change significantly in space, and may be blocked by other objects at certain moments, so static overlap and inclusion are not applicable. Therefore, in some embodiments, for dynamic objects, the second similarity condition includes: the volume similarity exceeds the third threshold, or the visual features match.
[0181] Volume similarity (Bounding Box Volume Similarity, Vol_Sim) characterizes the volume similarity between the unknown object and the identified object. When two 3D bounding boxes have similar volumes, the value of Vol_Sim will be larger. The embodied agent determines the volume of the unknown object and the volume of each identified object based on the first 3D boundary data and each second 3D boundary data, and determines the volume similarity. The volume similarity can be expressed by the following formula (3):
[0182] (3)
[0183] When the volume similarity between the unknown object and the volume of a recognized object exceeds a third threshold, it indicates that the unknown object and the recognized object are the same object.
[0184] If the visual features of the unknown object match the visual features of an identified object, the embodied agent can also determine that the unknown object is the same object as the identified object.
[0185] Exemplarily, whether the visual features of the unknown object match the visual features of the identified object can be determined based on the respective object features of the two, such as comparing the similarity between the visual features.
[0186] In the above embodiment, by comparing the stereo similarity based on the three-dimensional boundary data, unknown objects and identified objects can be effectively distinguished, thereby improving the accuracy of object recognition and further enhancing the embodied intelligent agent's understanding of the changes and interaction relationships of objects in dynamic scenes; in addition, according to the scene requirements, the similarity conditions of static objects and dynamic objects can be dynamically adjusted to adapt to scenes of different complexities.
[0187] Among them, in some embodiments, determining whether an unknown object is a static object or a dynamic object includes: obtaining two-dimensional boundary data of the unknown object, and determining a first image area in the targeted video frame based on the two-dimensional boundary data; performing feature extraction on the first image area to obtain a first object feature corresponding to the unknown object; obtaining a previous video frame, and determining a second image area in the previous video frame that has the same position as the first image area; performing feature extraction on the second image area to obtain a second object feature; if the difference between the first object feature and the second object feature is less than a preset threshold, determining that the unknown object is a static object; if the difference between the first object feature and the second object feature is not less than a preset threshold, determining that the unknown object is a dynamic object.
[0188] That is, the embodied intelligent agent locates the image area of the unknown object in the current video frame and the previous video frame based on the two-dimensional boundary data, extracts the visual features of the image area, and compares the feature differences between the two frames. Then, the embodied intelligent agent can determine whether the unknown object is static or dynamic based on the size of the difference and the preset threshold. As a result, the embodied intelligent agent can quickly identify the dynamic properties of objects in the target scene, providing a reliable basis for the subsequent identification of unknown objects.
[0189] Specifically, the embodied intelligent agent extracts the two-dimensional boundary data of the unknown object from the current video frame, such as a two-dimensional bounding box, and determines the first image area accordingly. The first image area is the pixel range where the unknown object is located in the current video frame. For example, in the current video frame, the xy coordinates of the upper left corner of the two-dimensional bounding box of the water bottle and the width and height of the two-dimensional bounding box are (200, 300, 50, 150) respectively, and the corresponding first image area is the pixel data within the two-dimensional bounding box.
[0190] The embodied intelligent agent may, for example, extract first object features of the unknown object from the first image region by using a feature extraction algorithm (such as a convolutional neural network).
[0191] Furthermore, the embodied intelligent agent determines the second image area with the same position in the previous video frame according to the position of the first image area, and extracts the second object feature of the image area. For example, in the previous video frame, the embodied intelligent agent also extracts features of the pixel data in the image area defined by the two-dimensional bounding box of (200, 300, 50, 150) to obtain the second object feature. Thus, the embodied intelligent agent determines whether the difference between the visual feature of the first object and the visual feature of the second object (such as Euclidean distance or cosine similarity, etc.) is less than a preset threshold. If the difference between the first object feature and the second object feature is less than the preset threshold, the embodied intelligent agent determines that the unknown object is a static object; otherwise, it means that the posture of the object in two adjacent frames has changed significantly, and the embodied intelligent agent determines that the unknown object is a dynamic object.
[0192] In the above embodiment, by comparing the visual features of the object image regions in the consecutive video frames, the dynamic properties of the object can be accurately determined, and then the object can be accurately re-identified, which has strong robustness.
[0193] When an unknown object is re-identified as an identified object, in some embodiments, the embodied intelligent agent can update the three-dimensional boundary data, object features, and context features corresponding to the identified object in the object memory by means of moving average update, and re-determine the positional relationship between the identified object and other objects to update the positional relationship corresponding to the identified object in the object memory. The moving average update method can effectively smooth errors and eliminate noise in a single detection, ensuring the stability and reliability of object data.
[0194] In addition to the changes in the target scene in the above embodiments, the target scene may also change according to the actions of animals, humans or other robots in the scene, thereby causing the state of the object to change. For example, in a first-person video, when a user picks up an object or moves an object, the perceptual data of the object will change due to the interactive action, so the corresponding perceptual data in the object memory needs to be updated. However, updating the state changes of an object caused by an interactive action is a key challenge, especially in the case of visual occlusion. Therefore, in some embodiments, the object recognition method in a dynamic scene provided by an embodiment of the present application also includes: when an action is detected in any video frame, the embodied intelligent agent determines the object associated with the action, and retrieves the perceptual data corresponding to the object in the object memory. If the object associated with the action includes multiple objects belonging to the same object category, the embodied intelligent agent determines the perceptual data corresponding to each object respectively.
[0195] For each object, the embodied agent renders the corresponding 3D boundary data in the object memory to the current frame, and uses the Vision Language Model (VLM) to determine whether the object defined by the 3D boundary data (e.g., the object within the 3D bounding box) is the target of the action. If so, the embodied agent updates the perception data of the object in the object memory that is the target of the action. For example, if the object was originally placed on the table, and the action indicates that the object is picked up, the embodied agent updates the state description of the object to "in hand", and so on.
[0196] Furthermore, when faced with a user's question-and-answer task regarding an action, the embodied agent can answer based on the sensory data corresponding to the object associated with the action. Figure 6As shown in the figure, there are jars A and B placed in the target scene respectively. When the user asks "Which jar is currently picked up?", the embodied intelligent agent can determine the perception data corresponding to jars A and B associated with the action. For example, based on the three-dimensional boundary data corresponding to jars A and B, the visual language model is used to determine whether jars A and B are the targets of the action. Then, the embodied intelligent agent outputs the corresponding answer, such as "The jar currently picked up is jar A". Therefore, through this visual prompting method, the action can be associated with the object and the object memory, the scene can be more accurately identified, and a more accurate answer to the question can be output.
[0197] The embodiment of the present application provides a human-computer interaction method based on embodied intelligent agents, and the execution subject can be a human-computer interaction device based on embodied intelligent agents. In the embodiment of the present application, the human-computer interaction method based on embodied intelligent agents is performed by a human-computer interaction device based on embodied intelligent agents as an example to illustrate the human-computer interaction device based on embodied intelligent agents provided in the embodiment of the present application.
[0198] The present application also provides a human-computer interaction device based on an embodied intelligent body, which is applied to an embodied intelligent body. Figure 7 As shown, the human-computer interaction device based on embodied intelligent body includes a response module 701, a determination module 702, a retrieval module 703, a planning module 704 and an execution module 705. Among them:
[0199] The response module 701 is used to respond to a description text sent by a user for a target scene and determine an interactive task indicated by the description text.
[0200] The determination module 702 is used to determine the target object associated with the interactive task.
[0201] The retrieval module 703 is used to search the pre-established object memory library to obtain the target perception data corresponding to the target object; the object memory library is used to store the perception data of all objects pre-identified from the target scene.
[0202] The planning module 704 is used to determine at least one target action from a plurality of pre-configured candidate actions based on the target perception data.
[0203] The execution module 705 is used to execute the target action on the target object to complete the interactive task.
[0204] According to the human-computer interaction device based on embodied intelligent body provided in the embodiment of the present application, by responding to the description text sent by the user for a target scene, determining the interaction task indicated by the description text, it is possible to respond to and parse the user's multimodal input, and improve the user's interaction experience; by determining the target object associated with the interaction task, and searching in a pre-established object memory library, the target perception data corresponding to the target object is obtained and introduced into the object memory library, and the object memory library is introduced to perform persistent memory on the objects in the target scene, so that the embodied intelligent body can dynamically adapt to scene changes and accurately locate the target object, thereby improving the completion and accuracy of the interaction task; and then based on the target perception data, at least one target action is determined from a plurality of pre-configured candidate actions, and the target action is executed on the target object to complete the interaction task, and the action matching mechanism enables the embodied intelligent body to adapt to the diverse task requirements in different scenarios, significantly improving the task execution efficiency and adaptability of the embodied intelligent body, thereby better meeting user needs.
[0205] In some embodiments, the determination module is also used to perform semantic understanding of the description text, determine the first target object indicated by the description text; the first target object is the target object for the target action to be performed; determine whether the description text indicates a second target object associated with the first target object; the second target object includes a target object that provides orientation clues for the first target object to perform the target action; the orientation clues are included in the description text; if the description text indicates a second target object, then based on the first target object and the second target object, determine the target object associated with the interactive task.
[0206] In some embodiments, the target perception data includes first perception data corresponding to the first target object and second perception data corresponding to the second target object; the determination module is also used to determine the positional relationship between the first target object and the second target object based on the first perception data and the second perception data; based on the positional relationship and the first perception data, determine the first target action corresponding to the first target object and the order of action execution among multiple pre-configured candidate actions; based on the positional relationship and the second perception data, determine the second target action corresponding to the second target object and the order of action execution among multiple pre-configured candidate actions; and execute the first target action and the second target action according to the action execution order.
[0207] In some embodiments, the above-mentioned device also includes a memory module, which is used to obtain a target video and embodied sensor data corresponding to the target video; the target video is used to present a target scene, including at least one video frame; based on at least one video frame and the embodied sensor data, the perception data of the embodied intelligent agent for at least one object in the target scene is determined respectively; an object memory library corresponding to the target scene is established, and the perception data of at least one object is stored in the object memory library as items.
[0208] In some embodiments, the memory module is also used to identify at least one object in any video frame, and obtain feature data, attribute data and two-dimensional boundary data of at least one object; perform dimensionality upscaling on the two-dimensional boundary data through the embodied sensor data corresponding to the video frame, and obtain three-dimensional boundary data corresponding to at least one object; determine the positional relationship between at least one object based on the three-dimensional boundary data corresponding to at least one object; based on the feature data, attribute data, three-dimensional boundary data and positional relationship, obtain the perception data of the embodied intelligent agent for at least one object in the target scene.
[0209] In some embodiments, the memory module is also used to, for any video frame, if an unknown object is identified in the targeted video frame, re-identify the unknown object to obtain a re-identification result; if the re-identification result indicates that the unknown object is the same object as any identified object, update the perception data corresponding to the identified object in the memory bank; if the re-identification result indicates that the unknown object is not any identified object, determine the perception data corresponding to the unknown object; and store the perception data of the unknown object in the object memory bank.
[0210] The human-computer interaction device based on embodied intelligent body in the embodiment of the present application can be a computer device, or a component in the computer device, such as an integrated circuit or a chip. The computer device can be an embodied intelligent body. Exemplarily, the computer device can be a mobile phone, a tablet computer, a laptop computer, a PDA, a vehicle-mounted computer device, a mobile Internet device (Mobile Internet Device, MID), an augmented reality (Augmented Reality, AR) / virtual reality (Virtual Reality, VR) device, a robot, a wearable device, an ultra-mobile personal computer (Ultra-mobile Personal Computer, UMPC), a netbook or a personal digital assistant (Personal Digital Assistant, PDA), etc. It can also be a server, a network attached storage (Network Attached Storage, NAS), a personal computer (Personal Computer, PC), a television (Television, TV), a teller machine or a self-service machine, etc., which is not specifically limited in the embodiment of the present application.
[0211] The embodiment of the present application also provides an embodied intelligent agent, which includes: Figure 7 The human-computer interaction device based on embodied intelligent agent is shown.
[0212] The human-computer interaction device based on embodied intelligent body in the embodiment of the present application may be a device with an operating system. The operating system may be a Microsoft (Windows) operating system, an Android (Android) operating system, an IOS operating system, or other possible operating systems, which are not specifically limited in the embodiment of the present application.
[0213] The human-computer interaction device based on embodied intelligent body provided in the embodiment of the present application can realize Figure 2 or Figure 4 To avoid repetition, the various processes implemented by the method embodiment are not described here.
[0214] In some embodiments, Figure 8 As shown, an embodiment of the present application also provides a computer device 800, including a processor 801, a memory 802, and a computer program stored in the memory 802 and executable on the processor 801. When the program is executed by the processor 801, each process of the above-mentioned method embodiments is implemented, and the same technical effect can be achieved. To avoid repetition, it will not be described here.
[0215] It should be noted that the computer device in the embodiment of the present application includes the mobile computer device and the non-mobile computer device mentioned above.
[0216] An embodiment of the present application also provides a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the various processes of the above-mentioned human-computer interaction method embodiment based on embodied intelligent body are implemented, and the same technical effect can be achieved. To avoid repetition, it will not be repeated here.
[0217] The processor is the processor in the computer device described in the above embodiment. The readable storage medium includes a computer readable storage medium, such as a computer read-only memory ROM, a random access memory RAM, a magnetic disk or an optical disk.
[0218] An embodiment of the present application also provides a computer program product, including a computer program, which, when executed by a processor, implements the above-mentioned human-computer interaction method based on embodied intelligent body.
[0219] The processor is the processor in the computer device described in the above embodiment. The readable storage medium includes a computer readable storage medium, such as a computer read-only memory ROM, a random access memory RAM, a magnetic disk or an optical disk.
[0220] An embodiment of the present application further provides a chip, which includes a processor and a communication interface, wherein the communication interface is coupled to the processor, and the processor is used to run programs or instructions to implement the various processes of the above-mentioned human-computer interaction method embodiment based on embodied intelligent body, and can achieve the same technical effect. To avoid repetition, it will not be repeated here.
[0221] It should be understood that the chip mentioned in the embodiments of the present application can also be called a system-level chip, a system chip, a chip system or a system-on-chip chip, etc.
[0222] It should be noted that, in this article, the terms "comprise", "include" or any other variant thereof are intended to cover non-exclusive inclusion, so that the process, method, article or device including a series of elements includes not only those elements, but also includes other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of further restrictions, the elements defined by the sentence "comprise one..." do not exclude the presence of other identical elements in the process, method, article or device including the element. In addition, it should be pointed out that the scope of the methods and devices in the embodiments of the present application is not limited to performing functions in the order shown or discussed, and may also include performing functions in a substantially simultaneous manner or in reverse order according to the functions involved, for example, the described method may be performed in an order different from that described, and various steps may also be added, omitted, or combined. In addition, the features described with reference to certain examples may be combined in other examples.
[0223] Through the description of the above implementation methods, those skilled in the art can clearly understand that the above-mentioned embodiment methods can be implemented by means of software plus a necessary general hardware platform, and of course by hardware, but in many cases the former is a better implementation method. Based on such an understanding, the technical solution of the present application, or the part that contributes to the relevant technology, can be embodied in the form of a computer software product, which is stored in a storage medium (such as ROM / RAM, disk, CD), and includes a number of instructions for a terminal (which can be a mobile phone, computer, server, or network device, etc.) to execute the methods described in each embodiment of the present application.
[0224] The embodiments of the present application are described above in conjunction with the accompanying drawings, but the present application is not limited to the above-mentioned specific implementation methods. The above-mentioned specific implementation methods are merely illustrative and not restrictive. Under the guidance of the present application, ordinary technicians in this field can also make many forms without departing from the purpose of the present application and the scope of protection of the claims, all of which are within the protection of the present application.
[0225] In the description of this specification, the description with reference to the terms "one embodiment", "some embodiments", "illustrative embodiments", "examples", "specific examples", or "some examples" means that the specific features, structures, materials, or characteristics described in conjunction with the embodiment or example are included in at least one embodiment or example of the present application. In this specification, the schematic representation of the above terms does not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described may be combined in any one or more embodiments or examples in a suitable manner.
[0226] If not otherwise specified, all embodiments and optional embodiments of the present application can be combined with each other to form a new technical solution.
[0227] Unless otherwise specified, all technical features and optional technical features of this application can be combined with each other to form a new technical solution.
[0228] If there is no special explanation, all steps of the present application can be performed sequentially or randomly, preferably sequentially. For example, the method includes steps (a) and (b), which means that the method may include steps (a) and (b) performed sequentially, or may include steps (b) and (a) performed sequentially. For example, it is mentioned that the method may also include step (c), which means that step (c) can be added to the method in any order. For example, the method may include steps (a), (b) and (c), or may include steps (a), (c) and (b), or may include steps (c), (a) and (b), etc.
[0229] The above are only preferred embodiments of the present application and are not intended to limit the present application. For those skilled in the art, the present application may have various modifications and variations. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application shall be included in the protection scope of the present application.
Claims
1. A human-computer interaction method based on embodied intelligent agent, characterized in that: The method comprises: In response to a description text sent by a user for a target scene, determining an interactive task indicated by the description text; Performing semantic understanding on the description text to determine a first target object indicated by the description text; the first target object is a target object for a target action to be performed; Determining whether the description text indicates a second target object associated with the first target object; the second target object includes a target object that provides a location clue for the first target object to perform the target action; the location clue is included in the description text; If the description text indicates that there is a second target object, determining a target object associated with the interactive task based on the first target object and the second target object; Searching a pre-established object memory library to obtain target perception data corresponding to the target object; the object memory library is used to store the perception data of all objects pre-identified from the target scene; Based on the target perception data, determining at least one target action and an action execution order corresponding to each target action from a plurality of pre-configured candidate actions; The target action is performed on the target object according to the action execution sequence to complete the interactive task.
2. The method according to claim 1, characterized in that The target perception data includes first perception data corresponding to the first target object and second perception data corresponding to the second target object; and determining at least one target action and an action execution sequence corresponding to each target action from a plurality of pre-configured candidate actions based on the target perception data includes: Determining a positional relationship between the first target object and the second target object based on the first perception data and the second perception data; Based on the positional relationship and the first perception data, determining a first target action corresponding to the first target object and an action execution order from a plurality of pre-configured candidate actions; Based on the positional relationship and the second perception data, a second target action corresponding to the second target object and an order of action execution are determined from a plurality of pre-configured candidate actions.
3. The method according to claim 1 or 2, characterized in that: The method further comprises: Acquire a target video and embodied sensor data corresponding to the target video; the target video is used to present the target scene and includes at least one video frame; Determining, based on the at least one video frame and the embodied sensory data, perception data of the embodied intelligent agent for at least one object in the target scene; An object memory library corresponding to the target scene is established, and the perception data of the at least one object is stored in the object memory library.
4. The method according to claim 3, characterized in that The determining, based on the at least one video frame and the embodied sensor data, respectively, the perception data of the embodied intelligent agent for at least one object in the target scene comprises: For any video frame, identifying at least one object in the targeted video frame, and obtaining feature data, attribute data and two-dimensional boundary data of the at least one object; Performing dimensionality upscaling processing on the two-dimensional boundary data using the embodied sensor data corresponding to the targeted video frame to obtain three-dimensional boundary data corresponding to the at least one object; Determining a positional relationship between the at least one object based on the three-dimensional boundary data respectively corresponding to the at least one object; Based on the feature data, the attribute data, the three-dimensional boundary data and the positional relationship, perception data of the embodied intelligent agent for at least one object in the target scene is obtained.
5. The method according to claim 4, characterized in that The method further comprises: For any video frame, if an unknown object is identified in the targeted video frame, the unknown object is re-identified to obtain a re-identification result; If the re-identification result indicates that the unknown object is the same object as any identified object, updating the perception data corresponding to the identified object in the memory bank; If the re-identification result indicates that the unknown object is not any identified object, determining the perception data corresponding to the unknown object; The perception data of the unknown object is stored in the object memory bank.
6. A human-computer interaction device based on embodied intelligent agent, characterized in that: The device comprises: A response module, configured to respond to a description text sent by a user for a target scene and determine an interactive task indicated by the description text; A determination module is used to perform semantic understanding on the description text, determine a first target object indicated by the description text; the first target object is a target object for a target action to be performed; determine whether the description text indicates a second target object associated with the first target object; the second target object includes a target object that provides a location clue for the first target object to perform the target action; the location clue is included in the description text; if the description text indicates a second target object, determine a target object associated with the interactive task based on the first target object and the second target object; A retrieval module, used to search in a pre-established object memory library to obtain target perception data corresponding to the target object; the object memory library is used to store the perception data of all objects pre-identified from the target scene; A planning module, configured to determine at least one target action and an action execution sequence corresponding to each target action from a plurality of pre-configured candidate actions based on the target perception data; An execution module is used to execute the target action on the target object according to the action execution sequence to complete the interaction task.
7. An embodied intelligent agent, characterized in that: It includes a scene understanding device based on embodied intelligent agent as described in claim 6.
8. A computer device comprising a memory, a processor and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the program, the human-computer interaction method based on embodied intelligent agent as described in any one of claims 1-5 is implemented.
9. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the human-computer interaction method based on embodied intelligent agent as described in any one of claims 1 to 5 is implemented.
Citation Information
Patent Citations
Object tracking method, tracking model training method, device, equipment and medium
CN115909173A
Intelligent agent behavior determination method, computer equipment and storage medium
CN117828039A
Large-model-driven intelligent agent scene exploration and memory management method and system
CN117854059A