Robot operation method, robot, medium, apparatus, and program product
By utilizing target mission information and visual input information in the robot operating system for task planning and spatial constraint planning of standardized interaction primitives, combined with closed-loop planning and execution, the versatility problem of fine-grained operations of the robot operating system in the real-world environment is solved, and more stable and universal operation control is achieved.
Patent Information
- Application Number
- CN202510117537.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2024-12-23
- Filing Date
- 2025-01-23
- Publication Date
- 2025-10-17
- Estimated Expiration
- 2045-01-23
AI Technical Summary
Existing robot operating systems lack the spatial understanding capability of fine-grained, low-level manipulation tasks in real-world environments, resulting in low versatility of manipulation control. Existing methods also lack task relevance and stability when generating manipulation primitives.
By utilizing target task information and visual input information for task planning, determining the subtasks and target objects of the stage, using canonical interaction primitives for spatial constraint planning, and combining closed-loop planning and execution, the robot's end-effector operation is controlled.
It improves the versatility and stability of robot operation control, ensures the consistency and reusability of operation strategies in different scenarios, and reduces dependence on high-cost data collection.
Smart Images

Figure CN119871410B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of robot operation, and in particular to a robot operation method, a robot, a medium, equipment and a program product. BACKGROUND
[0002] Due to the complexity and diversity of real-world environments, developing a general-purpose robot operating system has always been a challenging task. Inspired by the rapid development of Large Language Models (LLM) and Vision Language Models (VLM), researchers have utilized vast amounts of internet data to acquire rich common-sense knowledge and explored its application in the field of robots. Most current research focuses on using this knowledge for high-level task planning, lacking the ability to understand space, which is crucial for fine-grained, low-level operation tasks, resulting in low generalizability of robot operation control.
[0003] Based on this, the present application provides a robot operation method, a robot, a medium, equipment and a program product to improve related technologies. SUMMARY
[0004] The purpose of the present application is to provide a robot operation method, a robot, a medium, equipment and a program product to improve the generalizability of robot operation control.
[0005] The purpose of the present application is achieved by the following technical solutions:
[0006] In a first aspect, the present application provides a robot operation method, the method comprising: performing task planning using target task information and visual input information to determine at least one stage of sub-tasks and target objects, the target objects of the stage including a main object and a passive object; for at least one stage, performing spatial constraint planning according to the normative interaction primitives of the sub-tasks and target objects of the stage to obtain spatial constraint information based on normative interaction primitives; and controlling the operation of the end effector of the robot according to the sub-tasks and spatial constraint information of the stage.
[0007] In some embodiments, the main object and the passive object corresponding to the same stage satisfy the following conditions: the passive object is one of the at least one task object; the main object is the end effector, or the main object is one of the at least one task object and is different from the passive object.
[0008] In some embodiments, the obtaining process of the at least one task object comprises: processing the visual input information to label at least one input object; and screening the at least one task object corresponding to the target task information from the at least one input object.
[0009] In some embodiments, the obtaining process of the canonical interaction primitive of the target object comprises: extracting the canonical interaction primitive of the target object using the canonical representation information of the target object and the subtask of the stage to obtain the canonical interaction primitive of the target object; and the canonical representation information comprises a three-dimensional model and a pose.
[0010] In some embodiments, the three-dimensional model of the target object is obtained by three-dimensional structure reconstruction of the target object using the visual input information, and / or the pose of the target object is obtained by pose estimation of the target object using the visual input information.
[0011] In some embodiments, the canonical interaction primitive of the target object comprises an interaction point and / or an interaction direction.
[0012] In some embodiments, the interaction point comprises at least one of: a first interaction point that is visible and touchable; and a second interaction point that is invisible or untouchable.
[0013] In some embodiments, the obtaining process of the interaction point comprises: superimposing a Cartesian network on a target image containing the target object to obtain a superimposed image; the target image is determined according to the visual input information; and the interaction point is located using the superimposed image according to the subtask of the stage.
[0014] In some embodiments, the target image corresponding to the first interaction point comprises an input image obtained by processing the visual input information; and / or the target image corresponding to the first interaction point comprises at least one view of the three-dimensional model of the target object, and the at least one view comprises at least one of six views.
[0015] In some embodiments, the locating the interaction point using the superimposed image according to the subtask of the stage comprises: locating the interaction point using the superimposed image according to the task type of the subtask of the stage and the object type of the target object; and the object type comprises active and passive.
[0016] In some embodiments, the extraction process of the interaction point comprises: extracting a canonical interaction primitive based on the canonical representation information of the target object to obtain a plurality of candidate interaction points; processing the plurality of candidate interaction points according to the subtasks of the stage to generate an interaction point heat map; and determining at least one interaction point from the plurality of candidate interaction points using the interaction point heat map.
[0017] In some embodiments, the extraction process of the interaction direction comprises: extracting at least one candidate interaction direction of the target object according to the canonical representation information of the target object; generating semantic description information corresponding to the candidate interaction direction and calculating a relevance score of the semantic description information and the subtasks of the stage for one or more candidate interaction directions; and sorting the corresponding candidate interaction directions according to the relevance scores to determine at least one interaction direction.
[0018] In some embodiments, the spatial constraint planning based on the subtasks of the stage and the canonical interaction primitive of the target object to obtain spatial constraint information based on the canonical interaction primitive comprises: performing spatial constraint planning based on the subtasks of the stage and the canonical interaction primitive of the target object to obtain at least one unverified constraint information, the at least one unverified constraint information including distance constraint information and / or angle constraint information; verifying the unverified constraint information for one or more unverified constraint information to obtain a verification result corresponding to the unverified constraint information; and adding the unverified constraint information to the spatial constraint information if the verification result corresponding to the unverified constraint information is successful.
[0019] In some embodiments, the verification of the unverified constraint information to obtain a verification result corresponding to the unverified constraint information comprises: rendering an interaction image corresponding to the unverified constraint information; and verifying the interaction image to obtain the verification result corresponding to the unverified constraint information.
[0020] In some embodiments, the method further comprises: verifying a next unverified constraint information if the verification result corresponding to the unverified constraint information is failed; or resampling based on the current canonical interaction primitive to adjust the canonical interaction primitive and performing spatial constraint planning based on the adjusted canonical interaction primitive to obtain at least one unverified constraint information if the verification result corresponding to the unverified constraint information is optimized.
[0021] In some embodiments, the controlling the operation of the end effector of the robot according to the subtask and the space constraint information of the stage comprises: calculating a loss value of a target loss function according to the subtask and the space constraint information of the stage to determine a target pose of the end effector satisfying a target optimization condition; the target loss function comprises one or more loss terms of a constraint loss term, a collision loss term and a path loss term, and the loss value of at least one loss term is calculated according to the pose of the end effector; and performing trajectory planning on the end effector by using the target pose to obtain trajectory information of the end effector, and the trajectory information is used to control the operation of the end effector.
[0022] In some embodiments, the constraint loss term is calculated according to the pose of the target object of the stage, the pose of the main object is determined according to the pose of the end effector, and the method further comprises: performing pose tracking on the target object of the stage to update the pose of the target object; updating the loss value of at least one loss term in the target loss function based on the updated pose of the target object to realize the update of the loss value of the target loss function; and re-determining the target pose of the end effector satisfying the target optimization condition according to the updated loss value of the target loss function.
[0023] In a second aspect, the present application provides a robot operation system, comprising: a task planning module configured to perform task planning by using target task information and visual input information to determine a subtask and a target object of at least one stage, the target object of the stage comprising a main object and a target object; a constraint planning module configured to perform space constraint planning according to a specification interaction primitive of the subtask and the target object of at least one stage to obtain space constraint information based on the specification interaction primitive; and an operation control module configured to control the operation of an end effector of the robot according to the subtask and the space constraint information of at least one stage.
[0024] In a third aspect, the present application provides a robot comprising a robot operation system and an end effector, wherein the robot operation system is configured to perform any of the above methods to control the operation of the end effector.
[0025] In a fourth aspect, the present application provides a computer readable storage medium storing a computer program, wherein the computer program is executed by a processor to implement any of the above methods.
[0026] In a fifth aspect, the present application provides a computer device comprising a memory and a processor, wherein the memory stores a computer program, and the processor is configured to execute the computer program to implement any of the above methods.
[0027] In a sixth aspect, the present application provides a computer program product comprising a computer program which, when executed by a processor, implements any of the above methods.
[0028] The present application provides a robot operation method, a robot, a medium, an apparatus and a program product. The method comprises: performing task planning by using target task information and visual input information to determine at least one stage of sub-tasks and target objects, the target objects of a stage comprising a main object and a manipulated object; performing spatial constraint planning according to a canonical interaction primitive of the target objects of the stage and the sub-tasks of the stage to obtain spatial constraint information based on the canonical interaction primitive for at least one stage; and controlling the operation of an end effector of the robot according to the sub-tasks of the stage and the spatial constraint information. The present application can use the canonical interaction primitive of an object to perform spatial constraint planning, thereby improving the versatility of robot operation control. BRIEF DESCRIPTION OF DRAWINGS
[0029] The present application will be further described below in conjunction with the accompanying drawings and specific embodiments.
[0030] Figure 1a is a schematic diagram of a robot operation system framework provided by an embodiment of the present application.
[0031] Figure 1b is a schematic diagram of a robot operation method provided by an embodiment of the present application.
[0032] Figure 2 is a schematic diagram of an interaction point generation process provided by an embodiment of the present application.
[0033] Figure 3 is a schematic diagram of an interaction direction extraction process provided by an embodiment of the present application.
[0034] Figure 4 is a schematic diagram of a stability analysis of an interaction primitive provided by an embodiment of the present application.
[0035] Figure 5 is a schematic diagram of a qualitative analysis of the impact of a visual angle on performance provided by an embodiment of the present application.
[0036] Figure 6 is a schematic diagram of a closed-loop planning provided by an embodiment of the present application.
[0037] Figure 7 are two typical failure cases without closed-loop execution provided by an embodiment of the present application.
[0038] Figure 8 is a structural block diagram of a robot provided by an embodiment of the present application.
[0039] Figure 9 is a structural block diagram of a computer device provided by an embodiment of the present application. DETAILED DESCRIPTION
[0040] The technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the drawings in the present application. Obviously, the described embodiments are only part of the embodiments of the present application, rather than all the embodiments of the present application. Based on the embodiments in the present application, all other embodiments obtained by a person skilled in the art without creative work fall within the scope of protection of the present application.
[0041] In the description of the embodiments of the present application, it should be understood that the terms "first", "second" are used only for the purpose of description, and cannot be understood as indicating or implying relative importance or implicitly indicating the number of the indicated technical features. Therefore, the features defined with "first", "second" can explicitly or implicitly include one or more of the features. In the description of the embodiments of the present application, the meaning of "multiple" is two or more, unless otherwise specifically limited.
[0042] Developing a general-purpose robotic operating system has been a challenging task, mainly due to the complexity and diversity of real-world environments. Inspired by the rapid development of large language models (LLMs) and visual language models (VLMs), researchers have recently started exploring their applications in robotics by leveraging the vast amount of internet data to acquire rich common-sense knowledge. Most current research focuses on using this knowledge for high-level task planning, such as semantic reasoning. Despite these advances, current VLMs are primarily trained on large-scale 2D visual data and lack the ability to understand precise 3D space, which is crucial for fine-grained, low-level manipulation tasks. Therefore, they still struggle with fine-grained manipulation tasks in unstructured environments.
[0043] One approach to overcome this limitation is to fine-tune VLMs on large-scale robotic datasets, transforming them into Vision-Language-Action (VLA) models. However, this approach faces two major challenges. First, acquiring diverse, high-quality robotic data is costly and time-consuming; second, fine-tuning VLMs into VLAs generates representations that are specific to certain robots, limiting their generalizability. A promising alternative approach is to abstract robot actions into interaction primitives (e.g., points or vectors) and leverage VLMs for reasoning to define spatial constraints of these primitives, while leaving execution to relevant planning algorithms. However, related approaches have several limitations in defining and using primitives: the process of generating primitive proposals is task-agnostic, which can result in a lack of suitable proposals. Additionally, relying on hand-designed rules to post-process proposals also introduces instability.
[0044] In the application of vision-language models (VLMs), base models in the robotics domain excel in environmental understanding and high-level common-sense reasoning, demonstrating potential for controlling robots to perform general-purpose tasks in novel and unstructured environments. In some embodiments, VLA models capable of outputting robot trajectories are created by fine-tuning VLMs on robotic datasets, but these embodiments are limited by high data acquisition costs and insufficient generalizability. In some embodiments, operational primitives can be extracted from vision base models, which are then used as visual or linguistic cues for VLMs for high-level common-sense reasoning, combined with motion planners for low-level control. However, these embodiments are subject to two major limitations: one is the ambiguity of compressing 3D primitives into 2D images or 1D text required by VLMs; the other is the hallucination tendency of VLMs themselves. These limitations make it difficult to ensure that the high-level plans generated by VLMs are accurate.
[0045] Some related technologies reveal the unique advantages of open-vocabulary manipulation models (e.g., OmniManip models) in addressing these challenges, particularly in fine-grained 3D understanding and mitigating large model hallucinations. Structured representations for manipulation determine the capabilities and effectiveness of manipulation methods. Among various representation methods, keypoint representation methods are a popular choice due to their flexibility, generalizability, and ability to model diversity. However, these keypoint-based methods require manual task-specific annotation to generate actions. To enable zero-shot open-world manipulation, keypoints can be converted into visual cues for vision-language models (VLMs), facilitating the automatic generation of high-level planning results. However, while keypoints have advantages, they are also susceptible to occlusions and present challenges in the extraction and selection of specific keypoints.
[0046] Some related technologies use a rigid body pose representation method, which can efficiently define long-range dependencies between objects and provide certain robustness in occlusion cases. However, these related technologies need to pre-model geometric relationships, and due to the sparsity of poses, it is difficult to provide fine-grained geometric information. This limitation can cause the operation strategy to fail across different objects, especially in cases where the intra-class variation is large.
[0047] Referring to Figure 1a and Figure 1b , Figure 1a is a schematic diagram of a robot operating system framework provided by an embodiment of the present application, Figure 1b is a schematic diagram of a robot operating method provided by an embodiment of the present application.
[0048] To improve related technologies, an embodiment of the present application provides a robot operating method, which includes steps S101-S102.
[0049] Step S101: task planning using target task information and visual input information to determine at least one stage of subtasks and target objects, the target objects of the stage including a main object and a secondary object.
[0050] Step S102: for at least one stage, spatial constraint planning according to the canonical interaction primitive of the subtasks and target objects of the stage to obtain spatial constraint information based on the canonical interaction primitive; and controlling the operation of the end effector of the robot according to the subtasks and spatial constraint information of the stage.
[0051] In some embodiments, the method can be executed on a robot operating system. As an example, the robot operating system can be integrated with a robot.
[0052] For each stage, in addition to determining the subtasks and target objects, actions can also be determined. In some embodiments, the task planning using target task information and visual input information to determine at least one stage of subtasks and target objects can include task planning using target task information and visual input information to determine at least one stage of subtasks, actions, and target objects.
[0053] In the above embodiments, the target objects corresponding to the same stage include two objects, i.e., a main object and a subordinate object. In some embodiments, the main object and the subordinate object corresponding to the same stage can satisfy the following conditions: the subordinate object is one of the at least one task object; and the main object is the end effector, or the main object is one of the at least one task object and is different from the subordinate object. In other words, the main object can be the end effector or the task object, but the subordinate object can only be the task object. Among them, the task object is a collection of objects related to the task.
[0054] The above embodiments do not limit the visual input information, and in some embodiments, the visual input information can include one or more input images. The input image may, for example, include an RGB color image, a grayscale image, a depth image, etc.
[0055] In some embodiments, the acquisition process of the at least one task object can include: processing the visual input information to annotate at least one input object; and screening at least one task object corresponding to the target task information from the at least one input object. Since the task object is screened from the input object, the task object can be regarded as a subset of the input object.
[0056] In some embodiments, the processing of the visual input information to annotate at least one input object can include: processing the visual input information using a visual base model to obtain visual prompt information, the visual prompt information being used to annotate at least one input object.
[0057] The above embodiments do not limit the visual base model, which may, for example, include GroundingDINO (a target detection model), SAM (Segment Anything Model, a segmentation model), etc. As an example, GroundingDINO can annotate the foreground objects in the scene and their positions, and SAM can annotate the masks (which can be regarded as segmentation results) corresponding to the foreground objects.
[0058] In some embodiments, the screening of at least one task object corresponding to the target task information from the at least one input object can include: processing the visual prompt information and the RGB-D observation information of the at least one input object using a visual language model to screen at least one task object corresponding to the target task information. Among them, RGB (Red-Green-Blue) in RGB-D represents color information, and D (Depth) in RGB-D represents depth information. The RGB-D observation information of the input object can be observed using an RGB-D camera.
[0059] For example, visual input information can be input into a visual fundamental model, which is then used to process the visual input information to obtain visual cue information. For example, the visual cue information can include labeling instructions, which are used to label one or more foreground objects (these labeled foreground objects can be used as input objects). The labeling instructions and RGB-D observation information from the visual fundamental model (VFM) are input into a VLM visual language model, which is used to filter out task objects related to the task and divide the task into one or more stages.
[0060] For example, Figure 1a As shown, it is assumed that an operation task T is given (for example, pouring tea into a cup), and it is assumed that the visual input information includes an input image containing a teapot and a cup. First, after receiving the target task information corresponding to the operation task, the operation task can be decomposed. As an example, the two visual basis models (VFMs) GroundingDINO and SAM can be used to process the input image, annotate all foreground objects in the scene, and obtain visual cue information. Then, the target task information and visual cue information are input into the VLM, and the VLM is used to screen objects related to the task (i.e., task objects), and the task is decomposed into one or more stages S = {S1, S2, ..., S n}, n is a positive integer. Each stage S i Can be formalized as Where i represents the stage number, i is a positive integer, A i indicates an action to be performed (e.g., grabbing, pouring liquid), while and They represent the object that initiates the interaction (called the active object) and the object being manipulated (called the passive object). It can be seen that in the first stage of grasping the teapot, the subtask is, for example, to grasp the handle of the teapot with a gripper, the action is, for example, grasping, the active object is, for example, the robot's end effector (for example, the gripper), and the passive object is, for example, the teapot. In the second stage of pouring tea into a cup, the subtask is, for example, to pour tea from a teapot into a cup (Pour tea from a teapot into a cup), the action is, for example, pouring, the active object is, for example, the teapot, and the passive object is, for example, the cup. It can be seen that the teapot is a passive object in the first stage and an active object in the second stage. In other words, in the first stage, the object type of the teapot is passive; in the second stage, the object type of the teapot is active.
[0061] In some embodiments, the obtaining process of the canonical interaction primitive of the target object can include: using the subtask of the stage and canonical representation information of the target object to extract the canonical interaction primitive, to obtain the canonical interaction primitive of the target object; and the canonical representation information includes a three-dimensional model and a pose.
[0062] In some embodiments, the three-dimensional model of the target object can be obtained by three-dimensional structure reconstruction of the target object using the visual input information. As an example, a three-dimensional generation model such as a monocular 3D generation network can be used to process the visual input information to generate an object mesh (e.g., a three-dimensional model) of the target object. The three-dimensional model may, for example, adopt a 3D mesh model. The 3D generation model may, for example, adopt a 3D AIGC model. AIGC (Artificial Intelligence Generated Content) refers to generative artificial intelligence.
[0063] In some embodiments, the pose of the target object can be obtained by pose estimation of the target object using the visual input information. As an example, a 6D pose estimation model can be used to process the visual input information to estimate the pose of the target object. The 6D pose may, for example, include three-dimensional position information and three-dimensional pose information. That is, a general 6D pose estimation model can be used to normalize the object and describe the rigid body transformation of the object in the interaction process.
[0064] In some embodiments, the canonical interaction primitive of the target object can include an interaction point and / or an interaction direction. The canonical interaction primitive may, for example, refer to an interaction primitive in a canonical space.
[0065] The above embodiments implement an object-centric canonical interaction primitive. Specifically, the above embodiments propose a novel object-centric representation method, which describes the interaction manner of an object in an operation task using a canonical interaction primitive. As an example, the interaction primitive of an object (e.g., a target object, including a primary object and a secondary object) can be characterized by an interaction point and an interaction direction in its canonical space. As an example, the interaction point p ∈ R 3 may represent a key position on the object where interaction occurs, and the interaction direction v ∈ R 3 may represent a main axis related to the task. Together, they form an interaction primitive O = {p, v}, which encapsulates the basic geometric and functional properties required to meet the task constraints. R 3 denotes a three-dimensional space. Since these canonical interaction primitives are defined with respect to the canonical space of the object, they remain consistent in different scenarios, making the operation strategy more versatile and reusable.
[0066] In some embodiments, the interaction point can include at least one of: a first interaction point that is visible and touchable; a second interaction point that is invisible or untouchable. Here, touchable can be regarded as physically touchable, and untouchable can be regarded as physically untouchable.
[0067] In some embodiments, the extraction process of the interaction point can include: superimposing a Cartesian network on a target image containing the target object to obtain a superimposed image; and locating the interaction point using the superimposed image according to a subtask of the stage; the target image being determined according to the visual input information.
[0068] In some embodiments, the locating the interaction point using the superimposed image according to a subtask of the stage can include: locating the interaction point using the superimposed image according to a task type of the subtask of the stage and an object type of the target object; the object type including, for example, active and passive.
[0069] For example, assuming that the target object is a teapot, and assuming that the task type includes one or more of a grasping task, a pouring task, and a placing task. In the case of the task type being the grasping task and the object type being passive (i.e., the teapot being a passive object), the handle of the teapot can be located as the interaction point; in the case of the task type being the pouring task and the object type being active (i.e., the teapot being an active object), the spout can be located as the interaction point; in the case of the task type being the pouring task or the placing task and the object type being passive (i.e., the teapot being a passive object), the center of the opening of the teapot can be located as the interaction point. It can be seen that for the same object (e.g., the teapot), even if the task type is the same (e.g., the pouring task), if the object type is different, the extracted interaction point can be different. In the above example, in the case of the teapot being an active object, the located interaction point is the spout; in the case of the teapot being a passive object, the located interaction point is the center of the opening of the teapot.
[0070] In some embodiments, the target image corresponding to the first interaction point can include an input image; the input image being obtained by processing the visual input information. Assuming that the visual input information is in the form of an image, as an example, the image corresponding to the visual input information can be cropped to obtain the input image. Next, a Cartesian network can be superimposed on the input image to obtain a superimposed image; and the first interaction point can be located in an image plane corresponding to the superimposed image according to a subtask of the stage.
[0071] In some embodiments, the target image corresponding to the second interaction point comprises at least one view of a three-dimensional model of the target object, for example. In some embodiments, the at least one view can comprise at least one of six views. The six views can comprise, for example, a front view, a back view, a left view, a right view, a top view, a bottom view, and the like. As an example, the views of different perspectives of the three-dimensional model can be rendered from the three-dimensional model. Next, a Cartesian network can be superimposed on the at least one view to obtain a superimposed image; and the second interaction point can be located in the image plane corresponding to the superimposed image according to the subtask of the stage.
[0072] The extraction process of the first interaction point and the second interaction point is illustrated below. As shown in Figure 1a , first, a single view Figure 3 D network (as an example of a 3D generation model) can be used to obtain the 3D mesh model of the main object and the object related to the task, and then Omni6DPose (as an example of a 6D pose estimation model) can be used for pose estimation of the object to achieve a normalized representation. Next, the normalized interaction primitives related to the task can be extracted.
[0073] Referring to Figure 2 , Figure 2 is a schematic diagram of a generation process of an interaction point provided by an embodiment of the present application. In order to realize the positioning of the interaction point, as shown in Figure 2 , the interaction point can be divided into two categories: the first interaction point (for example, the handle of the teapot) which is visible and touchable; and the second interaction point (for example, the center of the opening of the teapot) which is invisible or touchable. The above embodiment can use the SCAFFOLD visual cue mechanism to enhance the positioning ability of the VLM for the interaction point. As an example, this mechanism can superimpose a Cartesian grid on the input image, and the corresponding two-dimensional coordinates (for example, (x1, y1), (x1, y2), (x2, y1), (x2, y2)) are as shown in Figure 2 . The first interaction point can be directly located in the image plane, for example, while the second interaction point can be inferred by multi-view reasoning based on the normalized object representation, for example. As an example, the reasoning can start from the main perspective, and if there is ambiguity, switch to the orthogonal perspective to deal with the problem. The above embodiment uses different positioning methods to locate the two types of interaction points, namely the first interaction point and the second interaction point, which can make the positioning of the interaction point more flexible and reliable.
[0074] In some embodiments, the extraction process of the interaction point can include: extracting a canonical interaction primitive based on the canonical representation information of the target object to obtain a plurality of candidate interaction points; processing the plurality of candidate interaction points according to the subtask of the stage to generate an interaction point heat map; and determining at least one interaction point from the plurality of candidate interaction points using the interaction point heat map. In this way, the robustness of interaction point positioning can be improved by generating a heat map from a plurality of candidate interaction points and then determining an interaction point using the interaction point heat map.
[0075] In actual applications, the interaction point heat map can be drawn in combination with the task type and the object type to achieve accurate interaction point positioning. In some embodiments, the processing of the plurality of candidate interaction points according to the subtask of the stage to generate an interaction point heat map can include processing the plurality of candidate interaction points according to the task type of the subtask and the object type of the target object to generate the interaction point heat map. As an example, assuming that the target object is a teapot, in the case of a task type of grasping and an object type of passive, the interaction point heat map can be as shown in FIG. 8B. As can be seen, the position of the handle is highlighted in the interaction point heat map. Figure 2
[0076] In the canonical space, the principal axis of an object is mostly related to its function. In some embodiments, the extraction process of the interaction direction can include: extracting at least one candidate interaction direction of the target object according to the canonical representation information of the target object; for one or more candidate interaction directions, generating semantic description information corresponding to the candidate interaction direction, and calculating a relevance score of the semantic description information and the subtask of the stage; and sorting the corresponding candidate interaction directions according to the relevance score to determine at least one interaction direction. The relevance score may, for example, be expressed in the form of a ten-point scale, a percentage, or a percentage, which is not limited in the above embodiments. In addition to the relevance score, a correlation level or the like can also be used for evaluation and sorting, which is not limited in the above embodiments.
[0077] In order to efficiently obtain semantic description information, in some embodiments, the generation of the semantic description information corresponding to the candidate interaction direction can include processing the candidate interaction direction using a visual language model to generate the corresponding semantic description information.
[0078] In some embodiments, the calculation of the relevance score of the semantic description information and the subtask of the stage can include processing the semantic description information and the subtask description information of the stage using an LLM large language model to obtain the relevance score of the semantic description information and the subtask of the stage. The subtask description information of the stage, for example, is information for describing the subtask of the stage, which can be obtained when the subtask of the stage is determined.
[0079] See also Figure 3 , Figure 3 This is a schematic diagram of an interactive direction extraction process provided by an embodiment of the present application. In the figure, Stage is a stage. Figure 3 As shown, for the teapot, the three main axes in the figure can be regarded as candidate interaction directions (at this time, the three main axes can be called candidate axes). Due to the limitations of the current VLM's ability to understand space, it is still challenging to evaluate the relevance of these directions to the task. To this end, the above embodiment proposes a mechanism that combines VLM to generate descriptions and LLM scoring. First, VLM is used to generate semantic description information for each candidate interaction direction. As an example, the semantic description information of the first candidate axis is, for example: the axis passes through the center of the teapot, extends outward from the side, and is perpendicular to the spout outlet; the semantic description information of the second candidate axis is, for example: the axis is arranged vertically through the teapot, from bottom to top, and is perpendicular to the spout outlet; the semantic description information of the third candidate axis is, for example: the axis is horizontally aligned, passes through the center of the teapot, and extends outward from the spout outlet. Then, LLM is used to infer and evaluate the relevance of these semantic description information to the task, and a relevance score for each semantic description information is obtained. Through this process, a set of candidate interaction directions sorted by task requirements can be generated. In Figure 3 In the example, the ranking results from high to low relevance scores are: the third candidate axis, the second candidate axis, and the first candidate axis. Based on the ranking results, at least one interaction direction can be determined, for example, the direction where the third candidate axis is located.
[0080] In some embodiments, the spatial constraint planning is performed based on the canonical interaction primitives of the subtasks of the stage and the target object to obtain spatial constraint information based on the canonical interaction primitives, which may include: performing spatial constraint planning based on the canonical interaction primitives of the subtasks of the stage and the target object to obtain at least one unverified constraint information, wherein the at least one unverified constraint information includes distance constraint information and / or angle constraint information; for one or more unverified constraint information, verifying the unverified constraint information to obtain corresponding verification results of the unverified constraint information; and if the corresponding verification results of the unverified constraint information are successful, including the unverified constraint information in the spatial constraint information.
[0081] In practical applications, assuming that verification results can be divided into three categories: success, failure, and optimization, targeted treatment can be carried out for the two types of verification results other than success.
[0082] In some embodiments, the method may further include: verifying the next unverified constraint information if the verification result corresponding to the unverified constraint information is failure.
[0083] In some embodiments, the method can further include, in a case where the verification result corresponding to the unverified constraint information is optimization, resampling based on the current canonical interaction primitive to achieve adjustment of the canonical interaction primitive, and re-planning of the spatial constraint based on the adjusted canonical interaction primitive to obtain at least one unverified constraint information. For the unverified constraint information with the verification result of optimization, resampling of the canonical interaction primitive and re-planning of the spatial constraint can be performed, and after obtaining new unverified constraint information, verification can be performed on the unverified constraint information.
[0084] In some embodiments, the verification of the unverified constraint information to obtain the verification result corresponding to the unverified constraint information can include rendering an interaction image corresponding to the unverified constraint information, and verifying the interaction image to obtain the verification result corresponding to the unverified constraint information. In some embodiments, the verification of the interaction image to obtain the verification result corresponding to the unverified constraint information includes verifying the interaction image using a visual language model to obtain the verification result corresponding to the unverified constraint information.
[0085] In the above embodiments, the spatial constraint information is defined using canonical interaction primitives. In each stage S i , a set of spatial constraints C i is used to regulate the spatial relationship between the primary object and the secondary object. As an example, the spatial constraints can include distance constraints and / or angle constraints. Distance constraints d i are used to adjust the distance between interaction points. Angle constraints θ i are used to ensure the correct alignment of interaction directions. These spatial constraints collectively define the geometric rules required for precise spatial alignment and task execution. The overall spatial constraints of each stage S i may be represented, for example, as formula (1).
[0086]
[0087] In the above embodiments, the spatial constraint information is defined using canonical interaction primitives. As an example, an ordered list of constraint interaction primitives can be generated for each stage S i , denoted as where N is a positive integer. The above embodiments extract canonical interaction primitives of the primary object and the secondary object, denoted as O active and O passive , and the spatial constraints C that define their spatial relationship. However, this process belongs to open-loop reasoning, which inherently limits the robustness and adaptability of the system. Its main limitations come from the hallucination effect of large models and the dynamic characteristics of real-world environments.
[0088] To overcome these challenges, Figure 1a As shown, the embodiments of the present application propose a dual closed-loop system. The dual closed-loop system includes, for example, closed-loop planning and closed-loop execution. Closed-loop planning corresponds to the process of constructing spatial constraint information, while closed-loop execution corresponds to the process of controlling the robot's operations. The following first describes the closed-loop planning process, followed by the closed-loop execution process.
[0089] The above embodiments can achieve closed-loop planning. Specifically, to improve the accuracy of standardized interaction primitives and alleviate the illusion problem of the VLM, the above embodiments propose a self-correction mechanism (RRC) based on resampling, rendering, and checking. This mechanism uses timely feedback from the visual language model (VLM) to detect and correct interaction errors, thereby ensuring accurate execution of the task. The RRC process can be divided into two phases: the initial phase and the optimization phase.
[0090] In the initial stage, the interaction constraints K defined above can be i is considered as unverified constraint information, which specifies the spatial relationship between the active object and the passive object. For each unverified constraint C (k) , can render interactive images according to the current configuration I i , and transform it (ie, the interaction image I i ) is submitted to the VLM for verification. The VLM returns one of the following three results: success, failure, or optimization. For example, if it returns success, the constraint is accepted and the next constraint is evaluated (until all unverified constraint information is verified); if it returns failure, the next constraint is evaluated (until all unverified constraint information is verified); if it returns optimization, the optimization phase is entered for further adjustment. Taking the interaction direction in the specification interaction primitive as an example, in the optimization phase, the predicted interaction direction v can be used. i Fine resampling is performed around the object to correct the misalignment between the functional axis and the geometric axis. For example, in v i Resample one or more optimization directions around v (j) Wherein, j is a positive integer, and the number of optimization directions can be 1, 2, 4, 6, etc., which is not limited in the above embodiment. As an example, when the number of optimization directions is greater than 1, a uniform sampling method can be used.
[0091] The above introduces the closed-loop planning process, and the following introduces the closed-loop execution process.
[0092] In some embodiments, the controlling the operation of the end effector of the robot according to the subtask and the space constraint information of the stage can include: calculating a loss value of a target loss function according to the subtask and the space constraint information of the stage to determine a target pose of the end effector satisfying a target optimization condition; the target loss function includes one or more loss terms of a constraint loss term, a collision loss term, and a path loss term, and the loss value of at least one loss term is calculated according to the pose of the end effector; performing trajectory planning on the end effector using the target pose to obtain trajectory information of the end effector; and the trajectory information is used to control the operation of the end effector.
[0093] In some embodiments, the constraint loss term can be calculated according to the pose of the target object of the stage, the pose of the main object can be determined according to the pose of the end effector, and the method can further include: performing pose tracking on the target object of the stage to update the pose of the target object; updating the loss value of at least one loss term in the target loss function based on the updated pose of the target object to realize the update of the loss value of the target loss function; and re-determining the target pose of the end effector satisfying the target optimization condition according to the updated loss value of the target loss function. Next, the trajectory information of the end effector can be updated by performing trajectory planning on the end effector using the re-determined target pose; and the trajectory information is used to control the operation of the end effector, which will not be described herein.
[0094] In some embodiments, the pose tracking on the target object of the stage to update the pose of the target object can include: using a 6D pose tracker to perform pose tracking on the target object of the stage to update the pose of the target object.
[0095] The above embodiments extract object-centered canonical interaction primitives as space constraints in a closed-loop manner in the planning process of at least one stage, optimize the trajectory of the end effector by constraints in the execution process of at least one stage, and update the target pose in the closed loop. As an example, in the closed-loop planning process of each stage, a VLM visual language model can be used to extract object-centered canonical interaction primitives as space constraints in a closed-loop manner. In the closed-loop execution process of each stage, the trajectory of the end effector is optimized by constraints, and the 6D pose tracker is used for updating in the closed loop.
[0096] The above embodiments combine the fine-grained geometric information of key points and the stability of rigid 6D poses. As an example, precise manipulation can be achieved by automatically extracting the interaction points and interaction directions in the canonical coordinate frame of an object (corresponding to the canonical space of the object) by a VLM. In the framework provided by the above embodiments, a complex robot task can be decomposed into multiple stages, each of which is defined by a canonical interaction primitive with spatial constraints. Since these canonical interaction primitives are based on the canonical space of an object, they can be referred to as canonical interaction primitives. This structured way can precisely define the task requirements and facilitate the execution of complex operation tasks. In the following, examples will be given to illustrate how canonical interaction primitives serve as the basis for spatial constraints to achieve robust manipulation.
[0097] Once the canonical interaction primitives and their corresponding spatial constraint information are defined for each stage, the task execution can be formalized as an optimization problem. As an example, the optimization objective can be to determine the target pose P ee* of the end-effector by minimizing an objective loss function L
[0098]
[0099] where arg is the English abbreviation of argument (i.e., independent variable), arg min is the value of the independent variable when the expression behind it is minimized, and j is the serial number of the loss term. It can be seen that the objective loss function L includes, for example, a constraint loss term (i.e., constraint loss L C ), a collision loss term (i.e., collision loss L collision ), and a path loss term (i.e., path loss L path ), and N in formula (2) is, for example, 3. The loss value of at least one loss term can be calculated according to the pose P ee of the end-effector.
[0100] where the constraint loss L C ensures that the action complies with the spatial constraints C of the task, the definition of which can be as shown in formula (3).
[0101]
[0102] where ρ(·) measures the deviation between the current spatial relationship of the primary object P active and the secondary object P passive and the desired constraint C, Φ(·) maps the pose of the end-effector to the pose of the primary object, and the subscript t is the time.
[0103] The collision loss L collision reduces the collision of the end-effector with obstacles in the environment, the definition of which can be as shown in formula (4).
[0104]
[0105] where d(p ee , O j ) denotes the distance between the end-effector and the obstacle O j , and d min is the minimum safe distance allowed. d min may be selected or set according to the needs in practical applications, for example.
[0106] The path loss L p ensures the smoothness of motion, which can be defined as shown in equation (5).
[0107]
[0108] where d trans (·) and d rot (·) denote the translational and rotational displacement of the end-effector, respectively, and λ1 and λ2 are weight factors to balance the impact of translation and rotation.
[0109] By minimizing the objective loss function, the target pose P ee* of the end-effector can be dynamically adjusted to ensure successful task execution while reducing collisions and maintaining smooth motion. Although equation (3) outlines how to optimize the target pose P ee* of the end-effector using canonical interaction primitives and their corresponding spatial constraint information, real-world task execution often involves significant dynamic factors. For example, in a grasping task, deviations in the grasping pose can cause the object to move unexpectedly during grasping. In addition, in some dynamic environments, the target object may be displaced. These challenges highlight the critical role of closed-loop execution in dealing with uncertainties.
[0110] To address uncertainties in the real world, the above embodiments not only use object-centric canonical interaction primitives, but also employ pose tracking algorithms (e.g., 6D pose tracking algorithms) to update the pose P active of the primary object and the pose P passive of the secondary object in a timely manner. This timely feedback allows for dynamic adjustment of the target pose of the end-effector, enabling robust and precise closed-loop execution.
[0111] In one specific application scenario, experiments related to the above embodiments are conducted. First, the experimental setup is introduced as follows.
[0112] As an example, the hardware is configured as follows. The experimental platform is based on a Franka Emika Panda robot arm, with the parallel gripper finger portion replaced by a UMI finger. In terms of perception, two Intel RealSense D415 depth cameras are used. One camera is mounted on the gripper to provide a first perspective of the manipulation region; the other camera is placed opposite the robot to provide a third perspective of the workspace.
[0113] As an example, 12 tasks are designed to evaluate the model's manipulation capabilities in real-world scenarios. Six of these tasks involve manipulation of rigid objects (e.g., pouring tea), and the other six tasks focus on manipulation of articulated objects (e.g., opening a drawer). These tasks cover a diverse set of objects and aim to evaluate the model's ability to generalize and adapt in complex environments. For each task, 10 experiments are performed for each method, and the success rate is recorded. After each experiment, the object layout can be reconfigured to ensure the robustness of the evaluation.
[0114] As an example, the method provided by the embodiments of the present application (i.e., the OmniManip method) is compared with the following three baseline (track) methods. The first method is the VoxPoser method, which uses a large language model (LLM) and a visual language model (VLM) to generate a three-dimensional value map to synthesize robot trajectories, and is good at zero-shot learning and closed-loop control. The second method is the CoPa method, which introduces spatial constraints of object components and combines VLM to realize open-vocabulary manipulation. The third method is the ReKep method, which uses relationship key point constraints and hierarchical optimization to generate real-time actions from natural language instructions.
[0115] As an example, GPT-4o of the OpenAI API can be used as a visual language model, using a set of interaction examples as prompts to guide the model's reasoning in manipulation tasks. As an example, a relevant model is used to implement 6-DOF general grasping, and GenPose++ (as an example of a pose estimation model) is used to implement general 6D pose estimation.
[0116] During the above experiment process, the OmniManip is comprehensively evaluated in 12 open-vocabulary manipulation tasks, which cover from simple grasp and place operations to more complex tasks, such as object-object interaction with direction constraints and articulated object manipulation. As shown in Table 1, the OmniManip method exhibits strong zero-shot generalization ability and excellent overall performance without task-specific training. This generalization ability benefits from the common sense knowledge embedded in the VLM, and in addition, the efficient, object-centric canonical interaction primitive facilitates accurate three-dimensional perception and execution. OmniManip exhibits significant performance advantages over the baseline method, which is mainly due to the following two factors. The first factor is the efficiency and stability of the object-centric canonical interaction primitive, which is verified by a large number of experiments. The second factor is the advanced double closed-loop system for planning and execution. By introducing a new self-correction mechanism based on RRC, the double closed-loop system effectively alleviates the hallucination problem of large models.
[0117] Table 1
[0118]
[0119]
[0120] Table 1 shows the quantitative results of 12 real-world operation tasks. The first six tasks focus on the operation of rigid objects, and the last six involve the operation of articulated objects. In the table, “-” indicates that the method is difficult to handle the task due to its inherent principles. OmniManip (Ours) is the method provided by the embodiments of the present application, where Closed-loop refers to disabling closed-loop planning, and Open-loop refers to enabling closed-loop planning. As shown in Table 1, this closed-loop planning achieves more than 15% performance improvement in both rigid and articulated object manipulation tasks.
[0121] The reliability of OmniManip is explained as follows. Referring to Figure 4 and Figure 5 , Figure 4 is a stability analysis diagram of an interaction primitive provided by the embodiments of the present application (visualization of planning and corresponding execution results based on different methods, taking the “pour tea” task as an example), Figure 5 is a qualitative analysis diagram of the impact of perspective on performance provided by the embodiments of the present application (taking the “recycle battery” task as an example). In Figure 4 , Planning refers to planning, and Execution refers to execution.
[0122] To effectively connect visual language models (VLMs) with low-level manipulation, reliable interaction primitives are very important. In the experimental process, the reliability of OmniManip is evaluated from two dimensions: stability and viewpoint consistency. Stability, for example, refers to the ability to reliably extract task-relevant canonical interaction primitives. As shown in Figure 4 ReKep extracts keypoint proposals through semantic clustering, but it lacks sensitivity to spatial geometry and tasks, making it difficult to generate enough task-relevant keypoints. CoPa extracts parts through explicit pixel segmentation, which is highly sensitive to image texture and part shape. In contrast, the object-centric canonical interaction primitives of OmniManip sample interaction points in canonical space aligned with the object's functional axis, ensuring robustness and task-specific precision. Viewpoint consistency, for example, refers to the consistency of canonical interaction primitive extraction under different viewpoints, which is crucial to ensuring manipulation stability. Since ReKep and CoPa directly sample points from the object's surface, they both perform poorly in terms of viewpoint consistency. Taking ReKep as an example, Figure 5 shows the planning results of ReKep and OmniManip for different viewpoints in the task "recycle battery". As shown in Figure 5 ReKep successfully identifies interaction points at a 90° overhead viewpoint, but fails at a 0° frontal viewpoint, where the ideal target point floats in the air. In contrast, OmniManip represents object-centric canonical interaction primitives in canonical space, ensuring viewpoint invariance.
[0123] Table 2
[0124]
[0125] Table 2 is a quantitative analysis of the impact of viewpoint (or angle) on performance, using the task "recycle battery" as a case study. The quantitative comparison results provided in Table 2 show that the performance of OmniManip is almost invariant under different viewpoints, while the performance of ReKep is significantly affected by changes in viewpoint.
[0126] The efficiency of OmniManip is explained below. The interaction direction planning in OmniManip is driven by a goal-oriented sampling strategy. Compared with uniform sampling of the rotation space SO(3), OmniManip samples along the principal axes of the object's canonical space. Since the canonical space of the object is aligned with the functional axis, it can ensure the efficiency and effectiveness of sampling.
[0127] To evaluate this efficiency, the experimental process compares the sampling strategy of OmniManip with the uniform sampling strategy of SO(3) through two key indicators: the number of iterations and the corresponding task success rate.
[0128] Table 3
[0129]
[0130] Table 3 shows the quantitative analysis of the sampling efficiency for canonical interaction primitives. As shown in Table 3, OmniManip not only requires fewer iterations, but also achieves higher task performance, indicating that aligning the sampling process with the object functional axis can reduce the sampling overhead while improving overall performance.
[0131] Referring to Figure 6 , Figure 6 is a closed-loop planning schematic provided by an embodiment of the application, which realizes a self-correction mechanism through RRC.
[0132] The above embodiment can realize closed-loop planning. In the related method, the planning component of VLM operates in an open-loop manner, i.e., it is difficult to verify the correctness of the planning before execution. Although ReKep realizes closed-loop control through point tracking, this is limited to the execution phase and it is difficult to provide feedback on the planning results generated by VLM. In contrast, OmniManip introduces a self-correction mechanism through RRC to realize closed-loop planning, greatly reducing the planning failure caused by VLM hallucination, thereby providing more reliable planning. As an example, the results of disabling closed-loop planning (corresponding to Closed-loop) are provided in Table 1, in which the success rates of both rigid and articulated object manipulation tasks decrease by more than 15%, indicating the effectiveness of the closed-loop planning method. In Figure 6 In the embodiment, taking the task of "inserting a pen into a pen holder" as an example, the closed-loop planning result is qualitatively demonstrated. Obviously, OmniManip can effectively pre-render the planning result and perform self-correction through the RRC mechanism, thereby realizing closed-loop planning.
[0133] Even if the planning is perfect, open-loop execution can still cause task failure. Referring to Figure 7 , Figure 7 are two typical failure cases without closed-loop execution provided by an embodiment of the application.
[0134] Figure 7 Two typical examples are shown. As an example, the relative pose between the gripper and the object changes during the interaction; or the pose of the target object (for example, a cup) changes, such as the case where the object moves during task execution. In these examples, although the planning is successful, open-loop execution causes failure. To address these challenges, OmniManip uses pose tracking to achieve real-time closed-loop execution. ReKep can use point tracking for closed-loop control, but its failure rate is high due to occlusion problems. In contrast, OmniManip exhibits higher robustness to occlusion caused by object movement, which benefits from object-centered pose tracking, which can continue to track the interaction primitive in canonical space according to the object pose even if the canonical interaction primitive is no longer visible.
[0135] In some embodiments, OmniManip can be used to generate automatic demonstration data. Unlike related approaches that rely on task-specific information, OmniManip is able to collect demonstration trajectories for new tasks in a zero-shot manner without task-specific details or prior object knowledge. To validate the effectiveness of the data generated by OmniManip, 150 trajectories can be collected for each task for training a behavior cloning policy.
[0136] Table 4
[0137]
[0138] Table 4 shows the results of behavior cloning using examples generated by OmniManip. It can be seen that OmniManip achieves a high success rate.
[0139] To develop a more efficient and more general representation to connect the high-level reasoning capability of VLMs with precise low-level robot operations, the above embodiments provide a novel object-centric intermediate representation (e.g., canonical interaction primitives) that integrates interaction points and interaction directions in the canonical space of objects. This representation bridges the gap between the high-level common-sense reasoning of VLMs and the precise 3D spatial understanding. The canonical space of an object can be defined based on its functional affordances. Thus, the functionality of an object can be described in a more structured and semantically meaningful way in the canonical space. Meanwhile, technical progress in general object pose estimation makes it possible to canonicalize a wide range of objects.
[0140] By way of illustration, the above embodiments can employ a general 6D pose estimation model to canonicalize objects and describe the rigid body transformation of objects during interaction. In addition, a single-view 3D generative network can be used to generate detailed object meshes (e.g., 3D models). In the canonical space, initial sampling of interaction directions along the principal axes of an object provides a set of coarse interaction possibilities (e.g., candidate interaction directions). Then, VLMs are used to predict interaction points. Subsequently, VLMs are used to identify task-related canonical interaction primitives and estimate the spatial constraints among them. To improve the hallucination problem in VLM reasoning, a self-correcting mechanism can be introduced by interaction rendering and primitive resampling, thereby enabling closed-loop reasoning. Once the final policy is determined, actions can be computed by constraint optimization and robust and timely control can be ensured by pose tracking in the closed-loop execution phase.
[0141] The above embodiments can achieve efficient and effective canonical interaction primitive sampling. By utilizing the canonical space of an object, efficient and effective canonical interaction primitive sampling is achieved, enhancing the reasoning capability of the robot operating system. In addition, based on the object-centered intermediate representation, a double-loop, open-vocabulary robot operating system is provided. Furthermore, the rendering and resampling process drive the decision-making reasoning loop, while the pose tracking ensures the motion execution loop.
[0142] In summary, the embodiments of the present application propose a novel object-centered interaction representation (e.g., canonical interaction primitive), which bridges the gap between the high-level common-sense reasoning of VLM and the low-level robot operation, and realizes a double-loop, open-vocabulary operating system without fine-tuning the VLM. Extensive experiments show that the method provided by the above embodiments has strong zero-shot generalization capability in diversified operation tasks, and also demonstrates its potential in automatically generating robot operation data. Specifically, the above embodiments propose a novel object-centered intermediate representation method, which effectively bridges the gap between the visual language model (VLM) and the precise spatial reasoning required for robot manipulation. Secondly, the above embodiments construct interaction primitives in the canonical space of an object, converting high-level semantic reasoning into operational three-dimensional spatial constraints. In addition, the double-loop system ensures the robustness of decision-making and execution without fine-tuning the VLM. The method provided by the above embodiments shows strong zero-shot generalization capability in various manipulation tasks, highlighting its potential in automating robot data generation and improving the efficiency of robot systems in unstructured environments. The above embodiments provide a promising foundation for future research on scalable, open-vocabulary robot manipulation.
[0143] The embodiments of the present application also provide a robot operating system, which includes a task planning module, a constraint planning module, and an operation control module. The task planning module is configured to perform task planning by using target task information and visual input information to determine at least one stage of sub-tasks and target objects, the target objects of the stage including a main object and a target object. The constraint planning module is configured to perform spatial constraint planning for at least one stage according to the canonical interaction primitives of the target objects and the sub-tasks of the stage to obtain spatial constraint information based on the canonical interaction primitives. The operation control module is configured to control the operation of the end effector of the robot according to the sub-tasks and the spatial constraint information of the stage.
[0144] Referring to Figure 8 , Figure 8 is a structural block diagram of a robot provided by the embodiments of the present application.
[0145] The embodiments of the present application also provide a robot, which comprises a robot operating system and an end effector, and the robot operating system is used for executing any of the above methods to control the operation of the end effector.
[0146] In some embodiments, the robot is a humanoid robot (i.e. a human-like robot, a biped robot), a quadruped robot, a wheeled robot, a multi-joint robot arm or other automated equipment.
[0147] In some embodiments, the robot can further comprise one or more of an odometer, an IMU (Inertial Measurement Unit), an image sensor, a laser sensor, an angle encoder, a torque sensor, a PIR sensor (Passive Infrared Sensor).
[0148] In some embodiments, the image sensor can comprise a camera (or a camera module).
[0149] In some embodiments, the laser sensor can comprise a 2D laser sensor and / or a 3D laser sensor.
[0150] In some embodiments, the robot can be a multi-joint robot.
[0151] The embodiments of the present application also provide a computer readable storage medium, which stores a computer program, and the computer program is executed by a processor to implement any of the above methods.
[0152] The embodiments of the present application also provide a computer program product, which comprises a computer program, and the computer program is executed by a processor to implement any of the above methods.
[0153] The computer program product can adopt a portable compact disc read-only memory (CD-ROM) and comprise program codes, and can run on a terminal device, for example, a personal computer. However, the computer program product of the present application is not limited to this, and the computer program product can adopt any combination of one or more computer readable media.
[0154] The embodiments of the present application also provide a computer device, which comprises a memory and a processor, the memory stores a computer program, and the processor implements any of the above methods when executing the computer program.
[0155] Referring to Figure 9 , Figure 9 is a structural block diagram of a computer device provided by the embodiments of the present application.
[0156] The computer device is not limited in the embodiments of the present application, and for example, can be a local computer device, a cloud computer device, a distributed computer device, etc.
[0157] The computer device can include a memory 110, a processor 120, and a communication interface 130. The memory 110, the processor 120, and the communication interface 130 are connected through an internal connection path.
[0158] The memory 110 is configured to store a computer program. In some implementations, the computer program can include codes for implementing the method of the embodiments of the present application.
[0159] The processor 120 is configured to execute the computer program stored in the memory 110 to control the communication interface 130 to receive input data and information, output operation results, etc. In some implementations, when the scheme of the embodiments of the present application is implemented by software or firmware, the computer program for implementing the scheme of the embodiments of the present application can be stored in the processor 120 and executed by the processor 120.
[0160] The memory 110 can be a volatile memory or a non-volatile memory, or can include both volatile and non-volatile memories. The non-volatile memory can be a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), or a flash memory. The volatile memory can be a random access memory (RAM). It should be noted that the memory 110 described herein is intended to include, but is not limited to, these and other suitable types of any memory. As an example, the memory 110 includes a random access memory (RAM), a cache memory, and a read-only memory (ROM). The memory 110 stores a computer program, which can be executed by the processor 120, so that the processor 120 implements the steps of any of the above methods.
[0161] The processor 120 can be a central processing unit (CPU), and can also be other general-purpose processors, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field programmable gate array (FPGA) or other programmable logic device, discrete gate or transistor logic, discrete hardware components, etc. The general-purpose processor can be a microprocessor, or the processor 120 can also be any conventional processor.
[0162] In the implementation process, each step of the above method can be completed by the integrated logic circuit of hardware in the processor 120 or the instruction in the form of software. The method disclosed in combination with the embodiments of the present application can be directly embodied as hardware processor execution completion, or combined with hardware and software modules in the processor 120 to complete execution. The software module can be located in a storage medium mature in the art such as random access memory, flash memory, read-only memory, programmable read-only memory, or electrically erasable programmable memory, register, etc. The storage medium is located in the memory 110, and the processor 120 reads the information in the memory 110, and combines the hardware to complete the steps of the above method. To avoid repetition, it will not be described in detail here.
[0163] In some implementations, in addition to the hardware units introduced above, the computer device can also include software modules, where the software modules can be, for example, an operating system, a basic input and output system (BIOS), application software, etc.
[0164] The operating system is used to manage hardware and / or software resources of the computer device, and is the kernel and cornerstone of the computer device. The operating system needs to handle basic transactions such as managing and configuring memory, determining the priority of system resource supply and demand, controlling input and output devices, operating network and managing file system, etc. In order to facilitate user operation, most operating systems will provide an operation interface for user to interact with the system.
[0165] The BIOS is used to run hardware initialization in the power-on boot stage, and provides runtime services for the operating system and application programs. In some implementations, the BIOS can also monitor the processor temperature and perform temperature protection strategies, etc.
[0166] Application software, also called application program, can be understood as software written for a specific application purpose of a user, and is one of the main classifications of computer software. For example, application software can be a program for realizing power control, temperature management, etc.
[0167] It can be understood that the specific examples in the present application are only to help those skilled in the art better understand the embodiments of the present application, and do not limit the protection scope of the present application.
[0168] It can be understood that in various embodiments of the present application, the size of the serial number of each process does not mean the order of execution, and the execution order of each process should be determined according to its function and inherent logic, and should not constitute any limitation on the implementation process of the present application.
[0169] It can be understood that the various embodiments described in the present application can be implemented alone or in combination, and the present application does not limit this.
[0170] Unless otherwise specified, all technical and scientific terms used in the present application have the same meaning as understood by those skilled in the art of the present application. The terms used in the present application are only for the purpose of describing the specific embodiments and are not intended to limit the scope of the present application. The term "and / or" used in the present application includes any and all combinations of one or more related listed items. The singular forms "a", "an" and "the" used in the present application and the appended claims are also intended to include the plural forms, unless the context clearly indicates otherwise.
[0171] Those of ordinary skill in the art can realize that the units and algorithm steps of the examples described in connection with the embodiments disclosed herein can be realized in electronic hardware, or a combination of computer software and electronic hardware. Whether the functions are realized in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of the present application.
[0172] Those skilled in the art can clearly understand that, for the convenience and brevity of description, the specific working processes of the above-described embodiments can refer to the corresponding processes in other embodiments, which will not be described here.
[0173] In several embodiments provided in the present application, it should be understood that the disclosed system, device and method can be implemented by other manners. For example, the above-described device embodiments are merely illustrative, for example, the division of units is merely a logical function division, and actual implementation can have another division manner, for example, a plurality of units or components can be combined or integrated into another system, or some features can be ignored or not executed. In addition, the displayed or discussed mutual coupling or direct coupling or communication connection can be indirect coupling or communication connection through some interfaces, devices or units, and can be electrical, mechanical or other forms.
[0174] The units described as separate components can or can not be physically separate, and the components shown as units can or can not be physical units, i.e., can be located in one place or can be distributed to a plurality of network units. Part or all of the units can be selected according to actual needs to achieve the purpose of the technical solutions of the present application.
[0175] In addition, the functional units in each embodiment of the present application can be integrated into one processing unit, or each unit can be physically present separately, or two or more units can be integrated into one unit.
[0176] If the functions are realized in the form of software function units and sold or used as independent products, they can be stored in a computer readable storage medium. Based on this understanding, the technical solutions of the present application or the parts that essentially contribute to the related art or the parts of the technical solutions can be embodied in the form of software products. The computer software product is stored in a storage medium, including a plurality of instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in the embodiments of the present application. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM), random access memory (RAM), magnetic disk or optical disk, and various program code storage media.
[0177] The above is merely a specific embodiment of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art can easily think of changes or replacements within the technical scope disclosed in the present application, which should be covered within the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.
Claims
1. A robot operation method, characterized in that: The method comprises: Performing task planning using target task information and visual input information to determine subtasks and target objects for at least one stage, wherein the target objects for the stage include active objects and passive objects, wherein the active objects and passive objects corresponding to the same stage meet the following conditions: the passive object is one of the at least one task objects; the active object is an end effector, or the active object is one of the at least one task objects and is different from the passive object; For at least one stage, performing spatial constraint planning based on subtasks of the stage and canonical interaction primitives of a target object to obtain spatial constraint information based on the canonical interaction primitives; and controlling an operation of an end effector of the robot based on the subtasks of the stage and the spatial constraint information; The process of acquiring the canonical interaction primitives of the target object includes: extracting the canonical interaction primitives using the subtasks of the stage and the canonical representation information of the target object to obtain the canonical interaction primitives of the target object; the canonical representation information includes a three-dimensional model and a pose; The standard interaction primitives of the target object include interaction points and / or interaction directions.
2. The robot operation method according to claim 1, characterized in that: The process of acquiring the at least one task object includes: processing the visual input information to mark at least one input object; At least one task object corresponding to the target task information is filtered out from the at least one input object.
3. The robot operation method according to claim 1, characterized in that: The three-dimensional model of the target object is obtained by reconstructing the three-dimensional structure of the target object using the visual input information, and / or the position and posture of the target object is obtained by estimating the position and posture of the target object using the visual input information.
4. The robot operation method according to claim 1, characterized in that: The interaction point includes at least one of the following: a first interaction point that is visible and touchable; and a second interaction point that is invisible or untouchable.
5. The robot operation method according to claim 4, characterized in that: The process of extracting the interaction points includes: superimposing a Cartesian network on a target image containing the target object to obtain a superimposed image; the target image is determined based on the visual input information; According to the subtasks of the stage, the interaction point is located using the superimposed image.
6. The robot operation method according to claim 5, characterized in that: The target image corresponding to the first interaction point includes an input image, which is obtained by processing the visual input information; and / or, the target image corresponding to the first interaction point includes at least one view of the three-dimensional model of the target object, and the at least one view includes at least one of the six views.
7. The robot operation method according to claim 5, characterized in that: The method of locating the interaction point using the superimposed image according to the subtask of the stage includes: The interaction point is located by using the superimposed image according to the task type of the subtask of the stage and the object type of the target object; the object type includes active and passive.
8. The robot operation method according to claim 1, characterized in that: The process of extracting the interaction points includes: Extracting standardized interaction primitives based on the standardized representation information of the target object to obtain multiple candidate interaction points; Processing a plurality of candidate interaction points according to the subtasks of the stage to generate an interaction point heat map; At least one interaction point is determined from a plurality of candidate interaction points using the interaction point heat map.
9. The robot operation method according to claim 1, characterized in that: The process of extracting the interaction direction includes: extracting at least one candidate interaction direction of the target object according to the normalized representation information of the target object; For one or more candidate interaction directions, generate semantic description information corresponding to the candidate interaction directions, and calculate a relevance score between the semantic description information and the subtasks of the stage; The corresponding candidate interaction directions are sorted according to the relevance scores to determine at least one interaction direction.
10. The robot operation method according to claim 1, characterized in that: The performing of spatial constraint planning according to the subtasks of the stage and the canonical interaction primitives of the target object to obtain spatial constraint information based on the canonical interaction primitives includes: Performing spatial constraint planning based on the subtasks of the stage and the canonical interaction primitives of the target object to obtain at least one unverified constraint information, wherein the at least one unverified constraint information includes distance constraint information and / or angle constraint information; For one or more unverified constraint information, the unverified constraint information is verified to obtain a verification result corresponding to the unverified constraint information; if the verification result corresponding to the unverified constraint information is successful, the unverified constraint information is included in the spatial constraint information.
11. The robot operation method according to claim 10, characterized in that: The verifying the unverified constraint information to obtain a verification result corresponding to the unverified constraint information includes: Rendering an interactive image corresponding to the unverified constraint information; The interactive image is verified to obtain a verification result corresponding to the unverified constraint information.
12. The robot operation method according to claim 10, characterized in that: The method further comprises: If the verification result of the unverified constraint information is failure, verify the next unverified constraint information; or When the verification result corresponding to the unverified constraint information is optimized, resampling is performed based on the current standard interaction primitive to achieve adjustment of the standard interaction primitive, and spatial constraint planning is re-performed based on the adjusted standard interaction primitive to obtain at least one unverified constraint information.
13. The robot operation method according to claim 1, characterized in that: The controlling the operation of the end effector of the robot according to the subtasks and spatial constraint information of the stage includes: Calculating a loss value of a target loss function based on the subtasks and spatial constraint information of the stage to determine a target pose of the end effector that satisfies a target optimization condition; the target loss function includes one or more loss terms selected from a constraint loss term, a collision loss term, and a path loss term, and a loss value of at least one loss term is calculated based on the pose of the end effector; The target posture is used to perform trajectory planning on the end effector to obtain trajectory information of the end effector; the trajectory information is used to control the operation of the end effector.
14. A robot, characterized in that: The robot comprises a robot operating system and an end effector, wherein the robot operating system is configured to execute the method according to any one of claims 1 to 13 to control the operation of the end effector.
15. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the method according to any one of claims 1 to 13 is implemented.
16. A computer device, characterized in that: The computer device includes a memory and a processor, the memory stores a computer program, and the processor implements the method according to any one of claims 1 to 13 when executing the computer program.
17. A computer program product, characterized in that The computer program product comprises a computer program, and when the computer program is executed by a processor, the method according to any one of claims 1 to 13 is implemented.
Citation Information
Patent Citations
Neural task planner for autonomous vehicles
CN113139652A
Hierarchical robot skill expression method, terminal and computer readable storage medium
CN114131598A