Intelligent agent control method, device and intelligent agent
By constructing and updating the three-dimensional scene map of the agent and using the prediction model to predict the target position, the problem of low task execution efficiency in an unknown environment is solved, and efficient and economical task execution is achieved.
Patent Information
- Application Number
- CN202411985389.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-31
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2044-12-31
AI Technical Summary
When existing agents perform tasks in unknown environments, they need to explore the environment first to obtain information, resulting in increased time and cost, low task execution efficiency and low success rate, affecting user experience.
By constructing a three-dimensional scene map of the target scene, and using the prediction model to predict the location of the target task-related objects, update the three-dimensional scene map, thereby controlling the agent to perform tasks, which is suitable for task execution in unknown scenarios.
Execute tasks without pre-exploring the environment, significantly reducing time and cost, improving task execution efficiency and success rate, having a wide range of application scenarios, and improving user experience.
Smart Images

Figure CN119379964B_ABST
Abstract
Description
Technical Field
[0001] The present application belongs to the field of intelligent agents, and in particular, relates to a control method, device and intelligent agent. Background Art
[0002] As people's living standards improve, intelligent agents are increasingly used in daily life. In related technologies, when intelligent agents perform tasks in unknown scenarios, they often need to explore the overall environment first to obtain the information required for the task, and then perform operations after the exploration is completed, which increases time and cost, and the task execution effect in unknown environments is not good, resulting in low task execution efficiency and low task execution success rate, affecting the user experience. Summary of the invention
[0003] The present application aims to solve at least one of the technical problems existing in the prior art. To this end, the present application proposes a control method, device and intelligent agent, which can perform target tasks without exploring the overall environment in advance, significantly reducing time and cost, improving task execution efficiency, and having good task execution effect and a wide range of application scenarios.
[0004] In a first aspect, the present application provides a method for controlling an intelligent agent, the method comprising:
[0005] Based on the target scene, generating a three-dimensional scene graph corresponding to the target scene;
[0006] Based on the target task and the three-dimensional scene graph, a prediction model is used to predict a target position of a prediction object related to the target task in the target scene;
[0007] Based on the predicted object and the target position, updating the three-dimensional scene graph;
[0008] Based on the three-dimensional scene graph, the agent is controlled.
[0009] According to the control method of the intelligent agent of the present application, by constructing an initial three-dimensional scene graph corresponding to the target scene, a prediction model is used to predict the target position of the prediction object related to the target task in the target scene to update the initial three-dimensional scene graph, and based on the updated three-dimensional scene graph, the intelligent agent is controlled to perform corresponding actions to complete the task. It is suitable for task execution in unknown scenes, and the target task can be executed without exploring the overall environment in advance, which significantly reduces time and cost, improves task execution efficiency, has good task execution effect and a wide range of application scenarios, thereby improving user experience.
[0010] According to an embodiment of the present application, generating a three-dimensional scene graph corresponding to the target scene based on the target scene includes:
[0011] Nodes are constructed based on objects in the target scene, edges are constructed based on spatial position relationships between the objects, and task dependency attributes corresponding to the nodes are constructed based on the target tasks to generate the three-dimensional scene graph.
[0012] According to an embodiment of the present application, the node includes at least one of: an identifier corresponding to the object, semantic information corresponding to the object, geometric shape information corresponding to the object, and three-dimensional contour information corresponding to the object.
[0013] According to one embodiment of the present application, the edge includes: a first node, a second node, a spatial position relationship between the first node and the second node, and a semantic description; wherein the first node and the second node are nodes with a supporting relationship.
[0014] According to an embodiment of the present application, the task dependency attribute includes: support attribute information of an object corresponding to the node, and the support attribute information is used to indicate whether the object has a support attribute.
[0015] According to one embodiment of the present application, the spatial position relationship is determined in the following manner:
[0016] The target scene is taken as the root node, the sub-areas included in the target scene are taken as the first-level nodes, the objects with supporting attributes in the sub-areas are taken as the second-level nodes, and the objects corresponding to the target tasks are taken as the third-level nodes, and the spatial position relationship between the objects is determined.
[0017] According to an embodiment of the present application, updating the three-dimensional scene graph based on the predicted object and the target position includes:
[0018] Controlling the intelligent agent to move to an area corresponding to the target position;
[0019] In the case where it is determined that the target position prediction is wrong based on the environmental information perceived by the agent in the area corresponding to the target position, based on the target task and the three-dimensional scene graph, using the prediction model, predict a new target position corresponding to the prediction object;
[0020] When it is determined that the target position prediction is correct based on the environmental information perceived by the agent in the area corresponding to the target position, the three-dimensional scene graph is supplemented based on the predicted object and the target position.
[0021] According to an embodiment of the present application, the supplementing of the three-dimensional scene graph based on the predicted object and the target position includes:
[0022] Determine the spatial position relationship corresponding to the prediction object and the task attribute corresponding to the prediction object based on the environmental information;
[0023] The prediction object is taken as a node, the spatial position relationship corresponding to the prediction object is taken as an edge, and the task attribute corresponding to the prediction object is taken as a task dependency attribute, which are added to the three-dimensional scene graph.
[0024] In a second aspect, the present application provides a control device for an intelligent body, the device comprising:
[0025] A first processing module, used for generating a three-dimensional scene graph corresponding to the target scene based on the target scene;
[0026] A second processing module is used to predict a target position of a prediction object related to the target task in the target scene using a prediction model based on the target task and the three-dimensional scene graph;
[0027] A third processing module, configured to update the three-dimensional scene graph based on the predicted object and the target position;
[0028] A fourth processing module is used to control the intelligent agent based on the three-dimensional scene graph.
[0029] According to the control device of the intelligent agent of the present application, by constructing an initial three-dimensional scene graph corresponding to the target scene, a prediction model is used to predict the target position of the prediction object related to the target task in the target scene to update the initial three-dimensional scene graph, and based on the updated three-dimensional scene graph, the intelligent agent is controlled to perform corresponding actions to complete the task. It is suitable for task execution in unknown scenes, and the target task can be executed without exploring the overall environment in advance, which significantly reduces time and cost, improves task execution efficiency, has good task execution effect and a wide range of application scenarios, thereby improving user experience.
[0030] In a third aspect, the present application provides an intelligent agent that performs a target task based on the control method of the intelligent agent as described in the first aspect.
[0031] In a fourth aspect, the present application provides an electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the computer program, the control method of the intelligent body as described in the first aspect above is implemented.
[0032] In a fifth aspect, the present application provides a non-transitory computer-readable storage medium having a computer program stored thereon, and when the computer program is executed by a processor, the control method of the intelligent agent as described in the first aspect above is implemented.
[0033] In a sixth aspect, the present application provides a computer program product, including a computer program, which, when executed by a processor, implements the control method of the intelligent agent as described in the first aspect above. BRIEF DESCRIPTION OF THE DRAWINGS
[0034] The above and / or additional aspects and advantages of the present application will become apparent and easily understood from the description of the embodiments in conjunction with the following drawings, in which:
[0035] Figure 1 It is one of the flow charts of the control method of the intelligent body provided in the embodiment of the present application;
[0036] Figure 2 This is the second flow chart of the control method of the intelligent body provided in the embodiment of the present application;
[0037] Figure 3 It is one of the intermediate result schematic diagrams of the control method of the intelligent body provided in the embodiment of the present application;
[0038] Figure 4 This is the second intermediate result schematic diagram of the control method of the intelligent body provided in the embodiment of the present application;
[0039] Figure 5 This is the third intermediate result schematic diagram of the control method of the intelligent body provided in the embodiment of the present application;
[0040] Figure 6 This is the fourth intermediate result schematic diagram of the control method of the intelligent body provided in the embodiment of the present application;
[0041] Figure 7 It is one of the structural schematic diagrams of the control device of the intelligent body provided in the embodiment of the present application;
[0042] Figure 8 This is the second structural diagram of the control device of the intelligent body provided in the embodiment of the present application;
[0043] Fig. 9 It is a schematic diagram of the structure of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0044] The following will be combined with the drawings in the embodiments of the present application to clearly describe the technical solutions in the embodiments of the present application. Obviously, the described embodiments are part of the embodiments of the present application, rather than all the embodiments. All other embodiments obtained by ordinary technicians in this field based on the embodiments in the present application belong to the scope of protection of this application.
[0045] The terms "first", "second", etc. in the specification and claims of the present application are used to distinguish similar objects, and are not used to describe a specific order or sequence. It should be understood that the data used in this way can be interchangeable under appropriate circumstances, so that the embodiments of the present application can be implemented in an order other than those illustrated or described here, and the objects distinguished by "first", "second", etc. are generally of one type, and the number of objects is not limited. For example, the first object can be one or more. In addition, "and / or" in the specification and claims represents at least one of the connected objects, and the character " / " generally indicates that the objects associated with each other are in an "or" relationship.
[0046] In conjunction with the accompanying drawings, the control method of the intelligent body, the control device of the intelligent body, the electronic device and the readable storage medium provided in the embodiments of the present application are described in detail through specific embodiments and their application scenarios.
[0047] The control method of the intelligent agent may be applied to a terminal, and may be specifically executed by hardware or software in the terminal.
[0048] The control method of an intelligent body provided in an embodiment of the present application may be executed by a functional module or functional entity in the intelligent body that can implement the control method of the intelligent body, or a server or electronic device that is communicatively connected to the intelligent body. The electronic devices mentioned in the embodiment of the present application include but are not limited to mobile phones, tablet computers, computers, cameras, and wearable devices. The control method of an intelligent body provided in an embodiment of the present application is described below using an electronic device as an example of an execution subject.
[0049] like Figure 1 As shown, the control method of the intelligent agent includes: step 110, step 120, step 130 and step 140.
[0050] Step 110: Based on the target scene, generate a three-dimensional scene graph corresponding to the target scene;
[0051] In this step, the target scene is the task execution scene in which the agent performs the target task. It can be a real scene or a virtual scene. The agent can be an embodied agent or a virtual object in a virtual scene, etc.
[0052] The target scene can be a known scene or an unknown scene.
[0053] Objects are other objects in the target scene except the agent, including movable objects and immovable objects.
[0054] For example, if the target scene is a house, the agent is a robot, and the target task is "the robot needs to put the banana on the table", the objects in the target scene may include but are not limited to: tables, chairs, refrigerators, sofas, beds, cabinets, doors, bananas, and apples. Among these objects, some are related to the target task, such as tables and bananas; in some cases, bottles may also be related to the target task.
[0055] In the case where the target scene is an unknown scene, the positions of some objects are unknown and variable, such as movable objects such as bananas, apples, water cups or chairs.
[0056] The three-dimensional scene graph is information used to characterize the intelligent agent's perception of the layout corresponding to the target scene, the categories of objects in the target scene, the object distribution, and the relative positional relationships between objects. It can be understood that the three-dimensional scene graph is a scene graph that is dynamically updated in real time based on the movement of the intelligent agent.
[0057] Before executing the task, the three-dimensional scene graph obtained by the agent is the initial scene graph, which is used to represent the initial perception environment in the target scene and to integrate known and unknown objects and their relationships. The initial scene graph can be constructed based on prior knowledge, including object positions, relationships, and spatial constraints, while taking into account the uncertainty in areas that have not yet been explored by the agent, and is acquired using the common sense knowledge of a pre-trained model.
[0058] It is understandable that the initial scene graph may lack objects related to the target task and the corresponding position information of the objects; Figure 4 As shown in the figure, taking the target task "the robot needs to put the banana on the table" as an example, before executing the task, the banana may not be within the robot's perception range, and the robot needs to find the banana through exploration and put it on the table; in this case, the initial scene graph obtained by the robot may not include the banana and its location information.
[0059] In some embodiments, step 110 may include:
[0060] Nodes are constructed based on the objects in the target scene, edges are constructed based on the spatial position relationship between objects, and task dependency attributes corresponding to nodes are constructed based on the target tasks to generate a three-dimensional scene graph.
[0061] In this embodiment, the following expression can be used to represent the three-dimensional scene graph: ;in, Used to represent nodes, To characterize the edge, It is used to characterize the task dependency properties associated with a node, and the task dependency properties are used to define the potential interactions with each node.
[0062] In some embodiments, the node may include at least one of: an identifier corresponding to the object, semantic information corresponding to the object, geometric shape information corresponding to the object, and three-dimensional contour information corresponding to the object.
[0063] In this embodiment, the identifier is a unique identifier of the object, which may be a name, ID or number, etc., and is used to distinguish other objects in the target scene.
[0064] The geometry information may be represented as a set of geometric primitives used to describe the complete geometry of the object.
[0065] The 3D contour information can be expressed as a 3D bounding box.
[0066] A node can be represented as ;in, is the i-th node in the target scene, is the identifier of the object corresponding to the i-th node; It is the semantic information corresponding to the object; It is a set of geometric primitives used to describe the complete geometric shape of an object. is the 3D bounding box of the object; i is a positive integer, and .
[0067] Through this node information, basic information such as the type, size, and shape of the object can be fully described.
[0068] In some embodiments, an edge may include: a first node, a second node, a spatial position relationship between the first node and the second node, and a semantic description; wherein the first node and the second node are nodes having a supporting relationship.
[0069] In this embodiment, the first node and the second node may be any two nodes having a supporting relationship in the target scene.
[0070] The supporting relationship includes: the first node is located above the second node, or the first node is located inside the second node; corresponding to an object with a supporting attribute class and / or a container class.
[0071] Edges are directed edges used to represent the support relationship between nodes. Include from parent node To child node Space transformation and semantic description of supporting relationships ;For example ; where vi is the i-th node, is the j-th node, i and j are positive integers respectively.
[0072] Among them, the parent node is a supporting or including node. For example, if a is placed on b, then a is a child node and b is a parent node; for another example, if a is placed inside b, then a is a child node and b is a parent node.
[0073] Space Transformation Node Relative to the node spatial position relationship.
[0074] Semantic Description Used to describe nodes Is it located at the node Above or within.
[0075] In some embodiments, the task dependency attribute may include: support attribute information of an object corresponding to the node, where the support attribute information is used to indicate whether the object has a support attribute.
[0076] In this embodiment, the task dependency attribute is used to characterize the attributes assigned to each node, and the support attribute information is used to characterize whether the object has a support attribute. For example, an object includes six faces, each face represents a support attribute, which is used to characterize that the object has six faces that can support other objects.
[0077] In some embodiments, for container-type objects, such as drawers and cabinets, the task dependency attributes may also include container opening and closing status information, such as whether a cabinet door is open or closed.
[0078] In some embodiments, the task dependency attribute can be expressed as ;in, is the support attribute information corresponding to the i-th node; It is the container opening and closing status information corresponding to the i-th node.
[0079] According to the control method of the intelligent agent provided in the embodiment of the present application, nodes are constructed based on objects, edges are constructed based on spatial position relationships, and task dependency attributes corresponding to nodes are determined based on target tasks to construct a three-dimensional scene graph. The physical scene in the target scene can be abstracted and summarized as a three-dimensional scene, providing a hierarchical semantic abstraction of the environment, simplifying the calculation process, improving calculation efficiency, and facilitating a better description of the dynamic changes of the target scene and the dynamic update of unknown scenes.
[0080] In some embodiments, the spatial position relationship may be determined by:
[0081] The target scene is taken as the root node, the sub-areas included in the target scene are the first-level nodes, the objects with supporting attributes in each sub-area are the second-level nodes, and the objects corresponding to the target tasks are the third-level nodes to determine the spatial position relationship between the objects.
[0082] In this embodiment, taking the target scene as a house as an example, the house can be used as the root node, such as Figure 3 Shown in medium grey circles.
[0083] The sub-areas can be the rooms in the house, such as Figure 3 As shown in the blue circle, the first-level nodes may include: living room, kitchen, bathroom, and bedrooms, etc.
[0084] Objects with support properties are objects that can support or hold other objects, such as containers and surfaces. Figure 3 As shown in the green circle, secondary nodes may include but are not limited to: tables, coffee tables, and beds.
[0085] The object corresponding to the target task is the object related to the target task. For example, if the target task is "the robot needs to put the banana on the table", the banana is the object corresponding to the target task, that is, the third-level node. Figure 3 Shown in pink circle.
[0086] According to the control method of the intelligent body provided in the embodiment of the present application, by layering the target scene and constructing a hierarchical mapping based on the hierarchical information to characterize the spatial position relationship between objects, it is possible to support downstream tasks, improve the subsequent updating effect and efficiency of the three-dimensional scene graph, and improve the task execution planning effect.
[0087] Step 120: Based on the target task and the three-dimensional scene graph, a prediction model is used to predict a target position of a prediction object related to the target task in the target scene;
[0088] In this step, we continue to take the target scene as a house, the intelligent agent as a robot, and the target task as "the robot needs to put a banana on the table" as an example. The prediction object related to the target task can be a banana; the target position is the position where the predicted object is most likely to be located.
[0089] The prediction model is a model obtained through pre-training and is used to predict the possible locations of objects related to the target task. During the training process, sample objects and the locations of sample objects in sample scenes can be used as training samples to train the prediction model.
[0090] It is understandable that in the actual execution process, the position of the banana is variable. For the intelligent agent, the position of the banana is unknown. In the initial three-dimensional scene graph obtained, there may be a lack of nodes and information related to the banana. In this case, a prediction model can be used to predict the information related to the possible location of the banana, thereby realizing intelligent exploration of unknown objects according to task requirements.
[0091] In some embodiments, the target position may be the position of the predicted object itself; or may be the position of a parent node corresponding to the predicted object, such as the position of a container where the predicted object may be placed.
[0092] Step 130: Update the three-dimensional scene graph based on the predicted object and the target position;
[0093] In this step, the predicted objects and the target positions corresponding to the predicted objects are used to complete the missing nodes in the three-dimensional scene graph and the edges and task dependency attributes corresponding to the nodes, thereby updating the three-dimensional scene graph. The updated three-dimensional scene graph includes objects related to the target task and the position information corresponding to the objects, which facilitates the subsequent generation of task execution plans to control the robot to perform the target task.
[0094] In some embodiments, it is also possible to determine whether the predicted target position is correct based on the feedback results of the intelligent agent during the task execution planning process, so as to correct the prediction results and update the three-dimensional scene graph with the corrected target position.
[0095] During the actual execution process, the intelligent agent can maintain the three-dimensional scene graph while performing the task. For example, during the execution of the task, based on the real-time perception of the environment, it determines whether the predicted object exists at the target position, corrects the target position based on the feedback result, and completes the three-dimensional scene graph based on the corrected target position to obtain a correct and relatively complete three-dimensional scene graph corresponding to the target scene.
[0096] In some embodiments, step 130 may include:
[0097] Control the agent to move to the area corresponding to the target position;
[0098] In the case where the target position prediction error is determined based on the environmental information perceived by the agent in the area corresponding to the target position, a prediction model is used to predict a new target position corresponding to the predicted object based on the three-dimensional scene graph and the target task;
[0099] When it is determined that the target position prediction is correct based on the environmental information perceived by the agent in the area corresponding to the target position, the three-dimensional scene graph is supplemented based on the predicted object and the target position.
[0100] In this embodiment, the perceived environmental information can be observed by sensors installed on the robot, and the perceived environmental information can be expressed as ,in, Indicates the nodes corresponding to visible objects. Represents the spatial position relationship between objects with supporting relationships.
[0101] In this application, it can be approximated that whenever the robot enters a room, it can perceive all objects and their relationships except those in closed containers.
[0102] The target position prediction error may include: the predicted object is not at the target position, and the predicted object is not at all other positions that the intelligent agent can perceive within the area corresponding to the target position.
[0103] Correct prediction of the target position may include: the predicted object is at the target position; or the predicted object is not at the target position but is at another position within the area corresponding to the target position that can be perceived by the intelligent agent.
[0104] During the actual execution process, after obtaining the updated three-dimensional scene graph, a corresponding task execution plan can be generated based on the three-dimensional scene graph. The task execution plan can include multiple sub-actions to indicate the movement path of the intelligent agent and the actions to be performed when moving to the corresponding node.
[0105] Before executing each action, a task execution plan is generated through the motion planner. If the motion planner throws an exception, the action will not be executed; otherwise, the action will be executed in the target environment.
[0106] For example, for nodes with incorrectly estimated target positions and the corresponding wrong target position , that is, the robot cannot find the expected object at the target location, and the object is not included in the new observation, that is, Not included in , the prediction model is used to predict the new position and use replace , to update the 3D scene graph .
[0107] If the robot recognizes the predicted object at the target position, or the position of the predicted object is included in the new observation, the position prediction is considered correct, and the node and position corresponding to the predicted object are added to the three-dimensional environment map; and the robot is controlled to perform the corresponding action based on the action corresponding to the node under the task execution plan. After that, the robot receives the new observation. .
[0108] In some embodiments, supplementing the three-dimensional scene graph based on the predicted object and target position may include:
[0109] Determine the spatial position relationship corresponding to the prediction object and the task attribute corresponding to the prediction object based on the environmental information;
[0110] The prediction object is taken as a node, the spatial position relationship corresponding to the prediction object is taken as an edge, and the task attribute corresponding to the prediction object is taken as a task dependency attribute, which are added to the three-dimensional scene graph.
[0111] In this embodiment, in order to complete the three-dimensional scene graph , the scene graph corresponding to the correct predicted object can be obtained from the prediction Add task-related nodes and properties , and edges added based on predicted relationships For each edge , predict conversion is unknown; among them, and For two nodes, For slave nodes To Node spatial transformation; A semantic description of the supporting relationship.
[0112] Assumptions Information about missing nodes is provided in ;in, is the predicted node; The identifier of the predicted object corresponding to the predicted node; To predict the semantic information of the object; is a set of geometric primitives used to describe the complete geometry of the predicted object; is the three-dimensional bounding box of the predicted object, p is a positive integer, and p∉ .
[0113] The updated 3D scene graph is finally described as: ;
[0114] in, is the updated 3D scene graph; ( , , ) is the initial three-dimensional scene graph; ( , , ) is the missing scene information in the initial 3D scene graph obtained based on the target task prediction .
[0115] According to the control method of the intelligent agent provided in the embodiment of the present application, the three-dimensional scene graph is updated by the predicted target position, and the intelligent agent is controlled to perform corresponding movements based on the updated three-dimensional scene graph. Whether the predicted target position is correct is judged according to the actual environmental information perceived by the intelligent agent during the movement, so as to correct the prediction result. The three-dimensional scene graph is updated by the corrected target position, so as to realize real-time updating and improvement of the three-dimensional scene graph, realize dynamic maintenance of the three-dimensional scene graph while executing the task, improve the comprehensiveness and accuracy of the three-dimensional scene graph, and improve the accuracy and success rate of task execution on the basis of improving the efficiency of task execution.
[0116] Step 140: Control the agent based on the three-dimensional scene graph;
[0117] In this step, after the three-dimensional scene graph is updated based on step 130, the agent can be controlled to perform corresponding actions based on the updated three-dimensional scene graph to complete the task.
[0118] In the actual execution process, the graph edit distance algorithm can be used to generate a set of graph editing operations, namely, the task execution plan, to control the agent to perform the task based on the generated graph editing operations.
[0119] In some embodiments, at least one of a global planning algorithm and a local planning algorithm may be used to optimize the generated graph editing operation to control the agent to perform tasks based on the optimized graph editing operation.
[0120] During the research and development process, the inventors discovered that in the real world, robots often lack a complete understanding of the environment and find it difficult to detect environmental changes caused by human activities in a timely manner. In this case, in related technologies, robots need to first explore to obtain the information required for the task and then operate. However, separating exploration from operation will lead to reduced task execution efficiency, increased time and cost, and fail to effectively integrate exploration and sequential operations. In addition, the state of the environment is often uncertain before it is perceived, and robots face the challenge of how to effectively perform tasks in unknown states. For example, the robot may not know the exact location of the target object of the task, so it must prioritize searching in the place where the object is most likely to be found. How to strike a balance between exploration and operation to achieve efficient task execution is a major problem in the prior art.
[0121] In the related art, there are some methods that attempt to solve the above problems, but they rely on artificially designed heuristic rules and are difficult to deal with dynamic changes and emergencies in complex and long-term tasks. These methods lack versatility and flexibility, especially when facing unknown environments.
[0122] According to the control method of the intelligent agent provided in the embodiment of the present application, by constructing an initial three-dimensional scene graph corresponding to the target scene, a prediction model is used to predict the target position of the prediction object related to the target task in the target scene to update the initial three-dimensional scene graph, and based on the updated three-dimensional scene graph, the intelligent agent is controlled to perform corresponding actions to complete the task. It is suitable for task execution in unknown and complex scenes, and does not require prior exploration of the overall environment or reliance on human control to execute the target task. It has a high degree of intelligence, significantly reduces time and cost, improves task execution efficiency, has good task execution effects and a wide range of application scenarios, thereby improving user experience.
[0123] In some embodiments, the three-dimensional scene graph can be dynamically updated based on the feedback results of the intelligent agent during the task execution planning process. When the three-dimensional scene graph changes, a new task execution plan can be generated based on the changed three-dimensional scene graph to control the intelligent agent based on the new task execution plan.
[0124] The implementation of step 140 is described below.
[0125] The embodiment of the present application also provides a method for controlling an intelligent agent.
[0126] like Figure 2 As shown, the control method of the intelligent agent includes: step 141 and step 142.
[0127] Step 141: Generate a target task execution plan based on the three-dimensional scene graph corresponding to the target scene and the target scene graph;
[0128] In this step, the three-dimensional scene graph is dynamically updated based on the feedback results of the agent during the execution planning process of the target task.
[0129] The feedback results may include, but are not limited to: environmental information perceived by the agent, related abnormal feedback based on the environmental information, and action feedback of the agent after performing a sub-action.
[0130] At the beginning stage of the intelligent agent's task execution, the three-dimensional scene graph is obtained by adding the target position of the predicted object related to the target task in the target scene predicted by the prediction model based on the target task and the three-dimensional scene graph to the initial three-dimensional scene graph.
[0131] During the process of the intelligent agent performing a task, the three-dimensional scene graph may be obtained by updating the last updated three-dimensional scene graph according to the action that has been completed by the intelligent agent.
[0132] The target scene graph is given, which is the final scene graph corresponding to the target scene after the agent determined in advance based on the target task completes the target task. For example, in the target scene graph, the banana should be placed on the table.
[0133] The target task execution plan is used to convert the current three-dimensional scene graph into the target scene graph, and corresponds to one or more sub-actions and displacements of the intelligent agent involved.
[0134] In the actual implementation process, Figure 5 As shown, the graph edit distance (GED) can be called after updating the 3D scene graph. Given the current 3D scene graph (initial graph) and the target scene graph (target graph), a task plan is generated by topological sorting to generate one or more groups of graph editing operations (i.e., action sequences) to transform the current 3D scene graph into the target scene graph.
[0135] In some embodiments, step 141 may include:
[0136] Obtaining at least one candidate task execution plan corresponding to the agent in a process of converting the three-dimensional scene graph into a target scene graph;
[0137] Based on the cost of executing the agent corresponding to each candidate task execution plan, a target task execution plan is obtained by screening at least one candidate task execution plan.
[0138] In this embodiment, the candidate task execution plan may be expressed as an action sequence consisting of a group of sub-actions.
[0139] The target task execution plan can be obtained through the following formula:
[0140]
[0141] in, Execute planning for target tasks; is the current three-dimensional scene graph; is the target scene graph; Used to indicate that Convert to A set of sub-actions, k is the number of sub-actions, and k is a positive integer; is the cost corresponding to executing the i-th sub-action, .
[0142] In some embodiments, the cost can be expressed as the displacement or time consumed by the agent to perform the sub-action.
[0143] Through the above formula, a set of sub-actions with the lowest cost can be determined as the target task execution plan.
[0144] According to the control method of the intelligent agent provided in the embodiment of the present application, through a global planning algorithm, a set of sub-actions with the minimum cost required for the intelligent agent to complete the task is used as the target task execution plan, which can save task execution time and improve task execution efficiency.
[0145] In some embodiments, based on the cost of executing the agent corresponding to each candidate task execution plan, selecting at least one candidate task execution plan to obtain a target task execution plan may include:
[0146] The computing agent updates the sum of costs corresponding to the nodes, edges, and task dependency attributes in the three-dimensional scene graph during the execution planning of the candidate tasks;
[0147] The candidate task execution plan corresponding to the minimum total cost is determined as the target task execution plan.
[0148] In this embodiment, the following categories of editing operations can be defined, corresponding to the actions of the agent:
[0149] Delete edge delete( ) pick( , ): used to characterize slave nodes Pick Node The corresponding object.
[0150] Add edge insert( )→place( , ): used to characterize the node The corresponding objects are placed in superior.
[0151] Replace attribute substitute( )→Open( ): used to represent open nodes The corresponding object's door makes the node Objects within the corresponding object are editable and observable.
[0152] Replace attribute substitute( )→Close( ): used to represent closed nodes The corresponding object's door makes the node Objects within the corresponding object are not editable and not observable.
[0153] Based on the above editing operations, the graph edit distance algorithm can be used to obtain a set of graph editing operations required between the 3D scene graph and the target scene graph, generating a set of sub-actions. , that is, the candidate task execution plan.
[0154] According to the control method of the intelligent agent provided in the embodiment of the present application, the calculation process is simplified and the calculation efficiency and accuracy are improved by converting the actions of the intelligent agent in the process of executing the task into operations on the nodes and edges in the three-dimensional scene graph.
[0155] In some embodiments, obtaining at least one candidate task execution plan corresponding to the agent in the process of converting the three-dimensional scene graph into the target scene graph may include:
[0156] Generate multiple sub-actions involved in the process of converting the three-dimensional scene graph into a target scene graph;
[0157] Based on the time sequence constraints of multiple sub-actions, a candidate task execution plan is obtained.
[0158] In this embodiment, a set of multiple sub-actions without order can be generated first, and then sorted according to the rules to obtain an action sequence, that is, a candidate task execution plan, so that a set of sub-actions corresponding to the same candidate task execution plan have time dependence, forming a partially ordered set.
[0159] For example, for the node , action pick( , ) must be in the action place( , ), these time dependencies create a set of constraints , where each pair of actions Must meet the conditions , to ensure that the actions are performed in the correct order, respecting the time sequence required to complete the task.
[0160] In the actual implementation process, The task execution planning problem on The topological sorting problem on , to generate a sorted sequence of actions . For example, define a search node ,in, , used to represent unexplored sub-actions, represents the remaining time constraint, is the child action that leads to the current node, and Is to execute sub-action The updated 3D scene graph. Initially, is generated by the graph edit distance algorithm, and Determined by time sequence, and As a starting action. Include sub-actions that do not depend on previous actions, making them available for exploration in subsequent search steps.
[0161] According to the control method of the intelligent body provided in the embodiment of the present application, by imposing time constraints on each sub-action, the rationality and feasibility of the generated candidate task execution plan can be improved, which helps to improve the success rate of task execution.
[0162] In some embodiments, based on the cost of executing the agent corresponding to each candidate task execution plan, selecting at least one candidate task execution plan to obtain a target task execution plan may include:
[0163] Inserting walking actions into the candidate task execution plan;
[0164] Based on the cost corresponding to executing each candidate task execution plan and the displacement of the agent when executing the walking action, a target task execution plan is obtained by screening from at least one candidate task execution plan.
[0165] In this embodiment, a walking action can be inserted into the action sequence so that the agent first moves to the vicinity of its parent container node, and then executes the sub-action corresponding to the child node, such as first controlling the robot to move from the current node to the refrigerator (parent node), and then controlling the robot to execute the "open door" sub-action to open the refrigerator, and then execute the sub-action of picking up a banana (child node), thereby solving the problem that the location of certain objects is unknown to the agent.
[0166] In some embodiments, the The algorithm is used to estimate the moving distance corresponding to the walking action to take into account the movement cost of the robot and seek to minimize the moving distance.
[0167] In some embodiments, a pruning strategy may be used to continuously update the upper limit of the movement cost, effectively reducing the search space and optimizing the path planning.
[0168] For example, before executing each sub-action, the agent will calculate the minimum moving distance to execute the sub-action. During the next search process, if the moving distance exceeds the current minimum moving distance, a pruning operation will be performed.
[0169] According to the control method of the intelligent agent provided in the embodiment of the present application, on the basis of calculating the action execution cost, the moving distance required by the intelligent agent before and after performing each action is further considered. By comprehensively considering the action execution cost and the moving distance, the search space is effectively reduced, the path planning effect is improved, and thus the task execution efficiency is improved.
[0170] Step 142: Execute planning based on the target task and control the agent.
[0171] In this step, after obtaining the optimal target task execution plan, the intelligent agent can be controlled to run based on the target task execution plan and run to each node in sequence according to the action sequence corresponding to the target task execution plan to perform corresponding actions, such as controlling the robot to move to the refrigerator, perform the opening action, take the banana in the refrigerator, and then close the refrigerator, control the robot to move to the table, perform the placement action, and place the banana on the table, thereby completing the target task "the robot needs to put the banana on the table".
[0172] In some embodiments, step 142 may further include:
[0173] Based on the target task execution plan, control the agent to move to the area corresponding to the target location;
[0174] In the case where the target position prediction error is determined based on the environmental information perceived by the agent in the area corresponding to the target position, a prediction model is used to predict a new target position corresponding to the predicted object based on the three-dimensional scene graph and the target task;
[0175] When it is determined that the target position prediction is correct based on the environmental information perceived by the agent in the area corresponding to the target position, the three-dimensional scene graph is supplemented based on the predicted object and the target position;
[0176] Generate a new target task execution plan based on the updated three-dimensional scene graph and the target scene graph;
[0177] Plan and control the agent based on the new target task execution.
[0178] In this embodiment, the updating method of the three-dimensional scene graph has been described in the above embodiment and will not be repeated here.
[0179] It can be understood that in the initial stage, the initial scene graph can be updated based on the target position of the predicted object in the target scene predicted by the prediction model, so as to generate a target task execution plan based on the updated three-dimensional scene graph and the target scene graph through a global planning algorithm to control the operation of the intelligent agent; before the intelligent agent executes the sub-actions included in the target task execution plan, it can determine whether the target position is correct based on the perceived environmental information, and continue to execute the corresponding sub-actions if it is correct; in the case of an error, the prediction model is used to predict the new target position to update the current three-dimensional scene graph, and based on the latest three-dimensional scene graph and the target scene graph, a global planning algorithm is used to generate a new target task execution plan, and the cycle is repeated until the target task is completed, thereby achieving simultaneous environmental exploration and task execution.
[0180] According to the control method of the intelligent agent provided in the embodiment of the present application, global planning is performed by dynamically updating the three-dimensional scene graph and the target scene graph based on the feedback results of the intelligent agent in the process of executing the target task planning, and a dynamic target task planning is generated to control the intelligent agent to perform related tasks based on the dynamic target task planning, and in turn optimize and update the three-dimensional scene graph based on the actual perception of the intelligent agent in the process of executing the task, so as to optimize the target task planning, better adapt to the unknown environment, improve the task execution effect in the unknown environment, and thus improve the user experience.
[0181] In some embodiments, step 141 may include:
[0182] When the intelligent agent is controlled to execute the sub-action corresponding to the first target node based on the target task execution plan, the three-dimensional scene graph is updated based on the feedback result;
[0183] Based on the updated three-dimensional scene graph and the target scene graph, the target task execution plan is updated.
[0184] In this embodiment, the first target node can be any node in the target task execution plan. After the intelligent agent executes the sub-action corresponding to the node, the environmental state corresponding to the node will change, and the change should be synchronously reflected in the three-dimensional scene graph.
[0185] For example, when the robot finds a banana in the refrigerator and picks it up, the banana will be temporarily removed from the 3D scene graph, which changes the state of the environment.
[0186] After the environmental state changes, the target task execution plan should be updated accordingly based on the updated three-dimensional scene graph and target scene graph to control the agent to perform subsequent actions.
[0187] In the actual execution process, the target task can be represented by the tuple P of the planning problem. Representation is used to solve the problem of sequential operation planning in an environment with unknown object states, a limited number of actions and deterministic transitions. is a three-dimensional scene graph; is the target scene graph. Among them, given an environment state , is a set of environmental states; it can be derived from the set of applicable actions Select a sub-action ; Transfer function Defines the dynamics of the environment, expressed in state Execute sub-actions Will lead to a new state Observed value , is the environmental information perceived by the agent when it reaches a new state.
[0188] The solution to the planning problem P is a sequence of actions , which can transform the status Transition to target state .
[0189] Considering that the movement frequency of rooms and some larger objects is low, but the movement frequency of smaller objects is high, the planning problem P assumes that the agent does not know the location of the objects involved in the task. is a set of environmental states, i.e., the three-dimensional scene graph proposed in this application ; The initial three-dimensional scene graph usually only includes houses, rooms, containers or surfaces (i.e., root nodes, first-level nodes, and second-level nodes), as well as the edges between them, but does not include the nodes of the objects corresponding to the target task; is the target scene graph, including all task-related nodes and their relationships, that is, including the nodes of the objects corresponding to the target task.
[0190] During the actual execution process, the three-dimensional scene graph can be updated based on the perceived environmental information and the actions performed under the current environmental state to update the environmental state, thereby solving the action sequence under the new environmental state, and while executing actions based on the new action sequence, continue to perceive the environmental information and update the three-dimensional scene graph, and repeat this cycle until the target scene at this time presents the environmental state corresponding to the target state.
[0191] According to the control method of the intelligent agent provided in the embodiment of the present application, the planning problem is defined based on the three-dimensional environment map, the environmental information perceived by the intelligent agent during operation, and the action execution status, the association and interaction between the elements are strengthened, and the dynamic update between the elements is realized, so that the intelligent agent can generate the target task plan by dynamically updating the three-dimensional scene map based on the feedback results in the process of executing the target task planning. While controlling the intelligent agent to perform related tasks based on the target task planning, the three-dimensional scene map is in turn optimized and updated based on the actual perception of the intelligent agent to optimize the target task planning, realize the synchronization of task execution and environmental exploration, enable it to better adapt to the unknown environment, and improve the task execution effect in the unknown environment.
[0192] In some embodiments, step 142 may further include:
[0193] When the agent runs to the second target node based on the target task execution plan, the perceived environmental information is obtained;
[0194] In the case where a planning anomaly is determined based on the environmental information, a solution strategy is generated based on the anomaly category;
[0195] Based on the solution strategy, control the agent.
[0196] In this embodiment, if Figure 6 As shown, the abnormal categories may include: blocking, unreachable, collision, and instability, etc.
[0197] Abnormal categories can be identified during the motion planning process.
[0198] In some embodiments, the identified abnormality category can be input into a pre-trained policy generation model to obtain a solution strategy output by the policy generation model to guide the intelligent agent to handle abnormal situations.
[0199] Continue to refer Figure 4 For example, when the robot moves to the parent node "table" corresponding to the banana, it recognizes that the banana is blocked by the bottle. The solution strategy "take the bottle away and then pick up the banana" can be obtained according to the strategy generation model. Based on the solution strategy, the sub-action "pick up the bottle" is inserted into the target task execution plan to realize the update of the target task execution plan. Based on the updated target task execution plan, the robot is controlled to take the bottle away first and then execute the sub-action "pick up the banana".
[0200] It is understandable that in the actual action execution process, the agent may still not be able to fully understand the target scene. The target task execution plan generated by the global planner may be difficult to deal with these unknown contents in the actual action execution process, which is uncertain. On the basis of global planning, further local planning is carried out, and the local planning algorithm is used to deal with abnormal categories to obtain a solution strategy to deal with unexpected situations caused by motion planning failure, further improving the success rate of task execution.
[0201] In this application, a two-layer planning scheme is adopted, using a global planner for global planning and a local planner for local planning. For the global planner, the environment is represented by a three-dimensional scene graph, integrating known and unknown objects and their relationships. The initial three-dimensional scene graph is constructed based on prior knowledge, including object positions, relationships, and spatial constraints, while taking into account the uncertainty in areas that the robot has not yet explored, using common sense knowledge from the model. After new observations (i.e., perceived environmental information) are collected during the robot's execution, the three-dimensional scene graph will be dynamically updated to support the use of graph editing algorithms and topological sorting for long-distance task planning. The process will be iterative until the task is completed.
[0202] If an exception occurs during execution that prevents the robot from completing an action, the model-based local planner recursively solves the encountered problem until the action is completed.
[0203] According to the control method of the intelligent agent provided in the embodiment of the present application, by locally optimizing the target task planning based on the actual performance of the intelligent agent in the process of executing the task on the basis of global planning, the understanding of the target scene is further improved, and the task execution effect and success rate are improved.
[0204] The control method of the intelligent body provided in the embodiment of the present application can be executed by the control device of the intelligent body. In the embodiment of the present application, the control device of the intelligent body executing the control method of the intelligent body is taken as an example to illustrate the control device of the intelligent body provided in the embodiment of the present application.
[0205] An embodiment of the present application also provides a control device for an intelligent body.
[0206] like Figure 7 As shown, the control device of the intelligent agent includes: a first processing module 710, a second processing module 720, a third processing module 730 and a fourth processing module 740.
[0207] A first processing module 710 is used to generate a three-dimensional scene graph corresponding to the target scene based on the target scene;
[0208] The second processing module 720 is used to predict the target position of the prediction object related to the target task in the target scene using the prediction model based on the target task and the three-dimensional scene graph;
[0209] A third processing module 730, for updating the three-dimensional scene graph based on the predicted object and the target position;
[0210] The fourth processing module 740 is used to control the intelligent agent based on the three-dimensional scene graph.
[0211] According to the control device of the intelligent agent provided in the embodiment of the present application, by constructing an initial three-dimensional scene graph corresponding to the target scene, a prediction model is used to predict the target position of the prediction object related to the target task in the target scene to update the initial three-dimensional scene graph, and based on the updated three-dimensional scene graph, the intelligent agent is controlled to perform corresponding actions to complete the task. It is suitable for task execution in unknown scenes, and the target task can be executed without exploring the overall environment in advance, which significantly reduces time and cost, improves task execution efficiency, has good task execution effect and a wide range of application scenarios, thereby improving user experience.
[0212] In some embodiments, the first processing module 710 may also be used to:
[0213] Nodes are constructed based on the objects in the target scene, edges are constructed based on the spatial position relationship between objects, and task dependency attributes corresponding to nodes are constructed based on the target tasks to generate a three-dimensional scene graph.
[0214] In some embodiments, the first processing module 710 may also be used to:
[0215] The target scene is taken as the root node, the sub-areas included in the target scene are the first-level nodes, the objects with supporting attributes in each sub-area are the second-level nodes, and the objects corresponding to the target tasks are the third-level nodes to determine the spatial position relationship between the objects.
[0216] In some embodiments, the third processing module 730 may also be used to:
[0217] Control the agent to move to the area corresponding to the target position;
[0218] In the case where the target position prediction error is determined based on the environmental information perceived by the agent in the area corresponding to the target position, a prediction model is used to predict a new target position corresponding to the predicted object based on the three-dimensional scene graph and the target task;
[0219] When it is determined that the target position prediction is correct based on the environmental information perceived by the agent in the area corresponding to the target position, the three-dimensional scene graph is supplemented based on the predicted object and the target position.
[0220] In some embodiments, the third processing module 730 may also be used to:
[0221] Determine the spatial position relationship corresponding to the prediction object and the task attribute corresponding to the prediction object based on the environmental information;
[0222] The prediction object is taken as a node, the spatial position relationship corresponding to the prediction object is taken as an edge, and the task attribute corresponding to the prediction object is taken as a task dependency attribute, which are added to the three-dimensional scene graph.
[0223] An embodiment of the present application also provides a control device for an intelligent body.
[0224] like Figure 8 As shown, the control device of the intelligent agent includes: a fifth processing module 810 and a sixth processing module 820.
[0225] A fifth processing module 810, configured to generate a target task execution plan based on the three-dimensional scene graph corresponding to the target scene and the target scene graph;
[0226] The sixth processing module 820 is used to control the intelligent agent based on the target task execution plan; wherein,
[0227] The three-dimensional scene graph is dynamically updated based on the feedback results of the agent during the execution planning process of the target task, and the target scene graph is determined based on the target task.
[0228] According to the control device of the intelligent agent provided in the embodiment of the present application, global planning is performed by dynamically updating the three-dimensional scene graph and the target scene graph based on the feedback results of the intelligent agent in the process of executing the target task planning, and a dynamic target task planning is generated to control the intelligent agent to perform related tasks based on the dynamic target task planning, and in turn optimize and update the three-dimensional scene graph based on the actual perception of the intelligent agent in the process of executing the task, so as to optimize the target task planning, better adapt to the unknown environment, improve the task execution effect in the unknown environment, and thus improve the user experience.
[0229] In some embodiments, the fifth processing module 810 may also be used to:
[0230] Obtaining at least one candidate task execution plan corresponding to the agent in a process of converting the three-dimensional scene graph into a target scene graph;
[0231] Based on the cost of executing the agent corresponding to each candidate task execution plan, a target task execution plan is obtained by screening at least one candidate task execution plan.
[0232] In some embodiments, the fifth processing module 810 may also be used to:
[0233] Generate multiple sub-actions involved in the process of converting the three-dimensional scene graph into a target scene graph;
[0234] Based on the time sequence constraints of multiple sub-actions, a candidate task execution plan is obtained.
[0235] In some embodiments, the fifth processing module 810 may also be used to:
[0236] The computing agent updates the sum of costs corresponding to the nodes, edges, and task dependency attributes in the three-dimensional scene graph during the execution planning of the candidate tasks;
[0237] The candidate task execution plan corresponding to the minimum total cost is determined as the target task execution plan.
[0238] In some embodiments, the fifth processing module 810 may also be used to:
[0239] Inserting walking actions into the candidate task execution plan;
[0240] Based on the cost corresponding to executing each candidate task execution plan and the displacement of the agent when executing the walking action, a target task execution plan is obtained by screening from at least one candidate task execution plan.
[0241] In some embodiments, the fifth processing module 810 may also be used to:
[0242] When the intelligent agent is controlled to execute the sub-action corresponding to the first target node based on the target task execution plan, the three-dimensional scene graph is updated based on the feedback result;
[0243] Based on the updated three-dimensional scene graph and the target scene graph, the target task execution plan is updated.
[0244] In some embodiments, the sixth processing module 820 may also be used to:
[0245] When the agent runs to the second target node based on the target task execution plan, the perceived environmental information is obtained;
[0246] In the case where a planning anomaly is determined based on the environmental information, a solution strategy is generated based on the anomaly category;
[0247] Based on the solution strategy, control the agent.
[0248] The control device of the intelligent body in the embodiment of the present application can be an electronic device, or a component in the electronic device, such as an integrated circuit or a chip. The electronic device can be a terminal, or it can be other devices other than a terminal. Exemplarily, the electronic device can be a mobile phone, a tablet computer, a laptop computer, a PDA, a vehicle-mounted electronic device, a mobile Internet device (Mobile Internet Device, MID), an augmented reality (augmented reality, AR) / virtual reality (virtual reality, VR) device, a robot, a wearable device, an ultra-mobile personal computer (ultra-mobile personal computer, UMPC), a netbook or a personal digital assistant (personal digital assistant, PDA), etc. It can also be a server, a network attached storage (Network Attached Storage, NAS), a personal computer (personal computer, PC), a television (television, TV), a teller machine or a self-service machine, etc., which is not specifically limited in the embodiment of the present application.
[0249] The control device of the intelligent body in the embodiment of the present application may be a device having an operating system. The operating system may be an Android operating system, an IOS operating system, or other possible operating systems, which are not specifically limited in the embodiment of the present application.
[0250] The control device of the intelligent body provided in the embodiment of the present application can achieve Figures 1 to 6 To avoid repetition, the various processes implemented by the method embodiment are not described here.
[0251] The embodiment of the present application also provides an intelligent agent.
[0252] The intelligent agent performs the target task based on the intelligent agent control method described in any of the above embodiments.
[0253] The intelligent agent may be an embodied intelligent agent or a virtual intelligent agent, which is not limited in this application.
[0254] According to the intelligent agent provided in the embodiment of the present application, by constructing an initial three-dimensional scene graph corresponding to the target scene, a prediction model is used to predict the target position of the prediction object related to the target task in the target scene to update the initial three-dimensional scene graph, and based on the updated three-dimensional scene graph, the intelligent agent is controlled to perform corresponding actions to complete the task. It is suitable for task execution in unknown scenes, and the target task can be executed without exploring the overall environment in advance, which significantly reduces time and cost, improves task execution efficiency, has good task execution effect and a wide range of application scenarios, thereby improving user experience.
[0255] In some embodiments, Fig. 9 As shown, an embodiment of the present application also provides an electronic device 900, including a processor 901, a memory 902, and a computer program stored in the memory 902 and executable on the processor 901. When the program is executed by the processor 901, each process of the above-mentioned intelligent body control method embodiment is implemented, and the same technical effect can be achieved. To avoid repetition, it will not be described here.
[0256] It should be noted that the electronic devices in the embodiments of the present application include the mobile electronic devices and non-mobile electronic devices mentioned above.
[0257] An embodiment of the present application also provides a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the various processes of the above-mentioned intelligent body control method embodiment are implemented, and the same technical effect can be achieved. To avoid repetition, it will not be repeated here.
[0258] The processor is the processor in the electronic device described in the above embodiment. The readable storage medium includes a computer readable storage medium, such as a computer read-only memory ROM, a random access memory RAM, a magnetic disk or an optical disk.
[0259] An embodiment of the present application also provides a computer program product, including a computer program, which implements the above-mentioned control method of the intelligent agent when executed by a processor.
[0260] The processor is the processor in the electronic device described in the above embodiment. The readable storage medium includes a computer readable storage medium, such as a computer read-only memory ROM, a random access memory RAM, a magnetic disk or an optical disk.
[0261] An embodiment of the present application further provides a chip, which includes a processor and a communication interface, wherein the communication interface is coupled to the processor, and the processor is used to run programs or instructions to implement the various processes of the above-mentioned intelligent body control method embodiment, and can achieve the same technical effect. To avoid repetition, it will not be repeated here.
[0262] It should be understood that the chip mentioned in the embodiments of the present application can also be called a system-level chip, a system chip, a chip system or a system-on-chip chip, etc.
[0263] It should be noted that, in this article, the terms "comprise", "include" or any other variant thereof are intended to cover non-exclusive inclusion, so that the process, method, article or device including a series of elements includes not only those elements, but also includes other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of further restrictions, the elements defined by the sentence "comprise one..." do not exclude the presence of other identical elements in the process, method, article or device including the element. In addition, it should be pointed out that the scope of the methods and devices in the embodiments of the present application is not limited to performing functions in the order shown or discussed, and may also include performing functions in a substantially simultaneous manner or in reverse order according to the functions involved, for example, the described method may be performed in an order different from that described, and various steps may also be added, omitted, or combined. In addition, the features described with reference to certain examples may be combined in other examples.
[0264] Through the description of the above implementation methods, those skilled in the art can clearly understand that the above-mentioned embodiment methods can be implemented by means of software plus a necessary general hardware platform, and of course by hardware, but in many cases the former is a better implementation method. Based on such an understanding, the technical solution of the present application, or the part that contributes to the prior art, can be embodied in the form of a computer software product, which is stored in a storage medium (such as ROM / RAM, a magnetic disk, or an optical disk), and includes a number of instructions for enabling a terminal (which can be a mobile phone, a computer, a server, or a network device, etc.) to execute the methods described in each embodiment of the present application.
[0265] The embodiments of the present application are described above in conjunction with the accompanying drawings, but the present application is not limited to the above-mentioned specific implementation methods. The above-mentioned specific implementation methods are merely illustrative and not restrictive. Under the guidance of the present application, ordinary technicians in this field can also make many forms without departing from the purpose of the present application and the scope of protection of the claims, all of which are within the protection of the present application.
[0266] In the description of this specification, the description with reference to the terms "one embodiment", "some embodiments", "illustrative embodiments", "examples", "specific examples", or "some examples" means that the specific features, structures, materials, or characteristics described in conjunction with the embodiment or example are included in at least one embodiment or example of the present application. In this specification, the schematic representation of the above terms does not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described may be combined in any one or more embodiments or examples in a suitable manner.
[0267] Although the embodiments of the present application have been shown and described, those skilled in the art will appreciate that various changes, modifications, substitutions and variations may be made to the embodiments without departing from the principles and spirit of the present application, and that the scope of the present application is defined by the claims and their equivalents.
Claims
1. A control method for an intelligent agent, characterized in that: include: Based on the target scene, generating a three-dimensional scene graph corresponding to the target scene; Based on the target task and the three-dimensional scene graph, a prediction model is used to predict a target position of a prediction object related to the target task in the target scene; Based on the predicted object and the target position, updating the three-dimensional scene graph; Based on the three-dimensional scene graph, controlling the intelligent agent; The updating of the three-dimensional scene graph based on the predicted object and the target position includes: Controlling the intelligent agent to move to an area corresponding to the target position; In the case where it is determined that the target position prediction is wrong based on the environmental information perceived by the agent in the area corresponding to the target position, based on the target task and the three-dimensional scene graph, using the prediction model, predict a new target position corresponding to the prediction object; When it is determined that the target position prediction is correct based on the environmental information perceived by the agent in the area corresponding to the target position, the three-dimensional scene graph is supplemented based on the predicted object and the target position.
2. The control method of an intelligent agent according to claim 1, characterized in that: The step of generating a three-dimensional scene graph corresponding to the target scene based on the target scene includes: Nodes are constructed based on objects in the target scene, edges are constructed based on spatial position relationships between the objects, and task dependency attributes corresponding to the nodes are constructed based on the target tasks to generate the three-dimensional scene graph.
3. The control method of an intelligent agent according to claim 2, characterized in that: The node includes at least one of: an identifier corresponding to the object, semantic information corresponding to the object, geometric shape information corresponding to the object, and three-dimensional contour information corresponding to the object.
4. The control method of an intelligent agent according to claim 2, characterized in that: The edge includes: a first node, a second node, a spatial position relationship between the first node and the second node, and a semantic description; wherein the first node and the second node are nodes having a supporting relationship.
5. The control method of an intelligent agent according to claim 2, characterized in that: The task dependency attribute includes: support attribute information of the object corresponding to the node, and the support attribute information is used to indicate whether the object has a support attribute.
6. The control method of an intelligent agent according to claim 2, characterized in that: The spatial position relationship is determined by: The target scene is taken as the root node, the sub-areas included in the target scene are taken as the first-level nodes, the objects with supporting attributes in the sub-areas are taken as the second-level nodes, and the objects corresponding to the target tasks are taken as the third-level nodes, and the spatial position relationship between the objects is determined.
7. The control method of an intelligent agent according to claim 1, characterized in that: The supplementing of the three-dimensional scene graph based on the predicted object and the target position includes: Determine the spatial position relationship corresponding to the prediction object and the task attribute corresponding to the prediction object based on the environmental information; The prediction object is taken as a node, the spatial position relationship corresponding to the prediction object is taken as an edge, and the task attribute corresponding to the prediction object is taken as a task dependency attribute, which are added to the three-dimensional scene graph.
8. A control device for an intelligent agent, characterized in that: include: A first processing module, used for generating a three-dimensional scene graph corresponding to the target scene based on the target scene; A second processing module is used to predict a target position of a prediction object related to the target task in the target scene using a prediction model based on the target task and the three-dimensional scene graph; A third processing module, configured to update the three-dimensional scene graph based on the predicted object and the target position; a fourth processing module, configured to control the agent based on the three-dimensional scene graph; The third processing module is used to: control the intelligent agent to move to the area corresponding to the target position; In the case where it is determined that the target position prediction is wrong based on the environmental information perceived by the agent in the area corresponding to the target position, based on the target task and the three-dimensional scene graph, using the prediction model, predict a new target position corresponding to the prediction object; When it is determined that the target position prediction is correct based on the environmental information perceived by the agent in the area corresponding to the target position, the three-dimensional scene graph is supplemented based on the predicted object and the target position.
9. An intelligent agent, characterized in that: The intelligent agent performs the target task based on the intelligent agent control method as described in any one of claims 1-7.
10. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the control method of the intelligent agent as described in any one of claims 1-7 is implemented.
11. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, the control method of the intelligent agent as described in any one of claims 1-7 is implemented.
Citation Information
Patent Citations
Grabbing control method and device, server, electronic equipment and storage medium
CN115213890A
Three-dimensional point cloud scene target positioning method based on natural language instruction
CN118229782A