Robot Manipulation Learning Method and Device Through Prediction of Interaction

Through the robot manipulation learning method of predicting interaction, the transition frame and interactive object position are predicted using visual and language embedding sequences, the problem of neglecting interaction behavior in the prior art is solved, and the expressiveness of robot manipulation and the performance of downstream applications are improved.

CN118081770BActive Publication Date: 2025-06-17SHANGHAI ARTIFICIAL INTELLIGENCE INNOVATION CENT
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202410439723.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-04-12
Publication Date
2025-06-17
Estimated Expiration
2044-04-12

AI Technical Summary

Technical Problem

The prior art ignores the behavioral patterns and physical interactions between the robot and the environment in robot manipulation, making it difficult for the model to identify and capture dynamic interactions in manipulation scenarios.

Method used

Through the robot manipulation learning method of predicting interaction, the keyframes of the initial state and the termination state in the interaction process are determined, and the visual encoding module and the language encoding module are used to obtain the visual embedding sequence and the language embedding sequence. Based on these sequences, pre-constructed predictors and detectors are used to predict the position of unseen transition frames and interactive objects, and information exchange is realized through bidirectional attention.

Benefits of technology

This method can more effectively capture and understand the interactive behavior in robot manipulation, improve the performance of downstream robot applications, realize a centralized learning process, and enhance the model's ability to represent specific knowledge of manipulation tasks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118081770B_ABST
    Figure CN118081770B_ABST
Patent Text Reader

Abstract

The present application relates to the field of computer vision technology, and discloses a robot manipulation learning method and device through predicted interaction. In terms of the encoder, the key frames of the initial state and the termination state in the interaction process are used as the input of the visual coding module to obtain a visual embedding sequence, and the language description in the interaction process is segmented according to the positions in the predefined vocabulary set to obtain a language embedding sequence. In the decoder, a predictor is used to predict unseen transition frames, representing the interaction between the initial state and the termination state, enabling the model to understand "how to interact"; a detector is used to infer the positions of the interacting objects, enabling the model to know "where to interact". Among them, information exchange between the transition frames and the positions of the interacting objects is achieved through bidirectional attention, promoting mutual support in the training process, which can improve the performance of downstream robot applications. And information irrelevant to the interaction is filtered out, emphasizing the key states, thereby realizing a concentrated learning process.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computer vision technology, and particularly to a robot control learning method and device through predictive interaction. Background Art

[0002] The visual motion control of a robot system involves perceiving and interpreting the surrounding environment from visual inputs, making informed decisions, and performing appropriate actions. This ability is crucial for a wide range of robot applications, including object manipulation, grasping, navigation, etc. Driven by the achievements in large-scale pre-trained vision and natural language processing, the robotics field attempts to utilize large-scale datasets to build generalizable representations. However, for robot manipulation, collecting demonstration data has proven to be both time-consuming and expensive. Therefore, researchers have started to explore representation learning methods that bypass the dependence on limited-domain robot data, which has become an important and popular research focus.

[0003] In the field of robotics, recent research efforts have utilized large-scale subjective perspective human video datasets to build a foundation for representation learning in robot manipulation. As Figure 2 (a)-(b) shown, previous methods widely adopted contrastive learning and masked signal modeling. Specifically, Figure 2 (a) Random image frames are obtained from the initial frame, and contrastive learning is performed between the random image frames and the language description "Remove top of the blender". During the contrastive learning process, the image and the language are aligned. Figure 2 (b) Random image frames are obtained from the initial frame, and masked signal reconstruction is performed between the random image frames and the language description "Remove top of the blender", such as filling in the masked part "Remove___blender". Although these methods provide valuable insights for improving the performance of the robot system in downstream tasks, their main focus often lies in distinguishing high-level semantic cues or capturing fine-grained pixel information, neglecting the crucial interaction dynamics, i.e., the behavior patterns and physical interactions that occur between the robot and the environment.

[0004] In most environments where robot systems are deployed, their functions are not limited to passive perception capabilities but also include active interaction with the environment. This motivates us to explore how to connect the interactions that shape the world and the visual representations that express the world. In this regard, a feasible method is video prediction pre-training, as Figure 2As shown in (c), a random image frame is obtained from the initial frame, and the random image frame and the language description "Remove top of the blender" are used for video prediction, and the (k + 1)-th frame is predicted based on the k-th frame. By learning to predict future frames, the model inherently acquires the ability to represent the temporal evolution of the scene. However, in the specific context of a robotic manipulation scenario, objects usually do not exhibit autonomous movement, and scene changes mainly come from interaction-based motion. Traditional methods only model the temporal relationship of consecutive frames, and this task setting is relatively simple, which is likely to introduce noise or redundant information, thus hindering the model from identifying interaction-related patterns or effectively capturing dynamic interactions in the manipulation scenario. Summary of the Invention

[0005] Embodiments of the present application provide a robot control learning method by predicting interactions, so as to solve the problem in the prior art that the behavioral patterns and physical interactions occurring between the robot and the environment are ignored.

[0006] Correspondingly, embodiments of the present application also provide a robot control learning device by predicting interactions, an electronic device, and a computer-readable storage medium, which are used to ensure the implementation and application of the above method.

[0007] To solve the above technical problems, embodiments of the present application disclose a robot control learning method by predicting interactions, and the method includes:

[0008] Determine the key frames of the initial state and the termination state during the interaction process, input them into the visual encoding module, and obtain a visual embedding sequence;

[0009] Segment the language description during the interaction process according to the positions in the predefined vocabulary set to obtain a language embedding sequence; the language embedding sequence includes a series of numerical identifiers;

[0010] Based on the visual embedding sequence and the language embedding sequence, use the pre-constructed predictor to predict unseen transition frames;

[0011] Use the pre-constructed detector to infer the positions of the interacting objects in the transition frames;

[0012] Among them, information exchange between the transition frames and the positions of the interacting objects is achieved through bidirectional attention.

[0013] Embodiments of the present application also disclose a robot control learning device by predicting interactions, and the device includes:

[0014] A visual encoding module, which is used to determine the key frames of the initial state and the termination state during the interaction process, input them into the visual encoding module, and obtain a visual embedding sequence;

[0015] A language encoding module for segmenting a language description in an interaction process according to positions in a predefined vocabulary set to obtain a language embedding sequence; the language embedding sequence includes a series of numerical identifiers;

[0016] An interaction mode prediction module for predicting unseen transition frames based on the visual embedding sequence and the language embedding sequence by using a pre-constructed predictor;

[0017] An interaction position prediction module for inferring the position of an interaction object in a transition frame by using a pre-constructed detector;

[0018] Wherein, information exchange between the transition frame and the position of the interaction object is realized through bidirectional attention.

[0019] An embodiment of the present application also discloses an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, one or more methods described in the embodiments of the present application are implemented.

[0020] An embodiment of the present application also discloses a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, one or more methods described in the embodiments of the present application are implemented.

[0021] In the embodiment of the present application, in terms of the encoder, the key frames of the initial state and the termination state in the interaction process are used as the input of the visual encoding module to obtain a visual embedding sequence. The language description in the interaction process is segmented according to the positions in the predefined vocabulary set to obtain a series of numerical identifiers as the language embedding sequence. In the decoder, based on the visual embedding sequence and the language embedding sequence, a pre-constructed predictor is used to predict unseen transition frames, representing the interaction between the initial state and the termination state, so that the model understands "how to interact"; a pre-constructed detector is used to infer the position of the interaction object in the transition frame, so that the model can obtain the knowledge of "where to interact". Wherein, information exchange between the transition frame and the position of the interaction object is realized through bidirectional attention, promoting mutual support in the training process. In the embodiment of the present application, interaction-oriented learning is introduced into representation learning to embed representations with manipulation task-specific knowledge, which can improve the performance of downstream robot applications. And the embodiment of the present application filters out information irrelevant to the interaction and emphasizes key states, thereby realizing a centralized learning process.

[0022] Additional aspects and advantages of the embodiments of the present application will be given in the following description part, which will become obvious from the following description, or can be understood through the practice of the present application. Description of the Drawings

[0023] The above and / or additional aspects and advantages of the present application will become apparent and be readily understood from the following description of embodiments in conjunction with the accompanying drawings, where:

[0024] Figure 1 is a flowchart of a robot manipulation learning method through predictive interaction provided for an embodiment of the present application;

[0025] Figure 2 is a schematic diagram of the network structures of an existing video prediction method and the MPI in an embodiment of the present application;

[0026] Figure 3 is a schematic diagram of the network structure of a robot manipulation learning method through predictive interaction provided for an embodiment of the present application;

[0027] Figure 4 is a diagram of experimental results provided for an embodiment of the present application;

[0028] Figure 5 is a schematic diagram of the structure of a robot manipulation learning device through predictive interaction provided for an embodiment of the present application;

[0029] Figure 6 is a schematic diagram of the structure of an electronic device provided for an embodiment of the present application. Detailed Embodiments

[0030] The embodiments of the present application are described in detail below. The examples of the embodiments are shown in the accompanying drawings, where the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described by referring to the accompanying drawings are exemplary and are only used to explain the present application and should not be construed as a limitation of the present application.

[0031] Those skilled in the art of the present technology can understand that unless specifically stated otherwise, the singular forms "a", "an", "the", and "said" used herein may also include the plural forms. It should be further understood that the term "comprising" used in the specification of the present application means the presence of features, integers, steps, operations, elements, and / or components, but does not exclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or combinations thereof. It should be understood that when we say an element is "connected" or "coupled" to another element, it can be directly connected or coupled to other elements, or there may also be intermediate elements. In addition, the "connection" or "coupling" used herein may include wireless connection or wireless coupling. The phrase "and / or" used herein includes all or any unit and all combinations of one or more related listed items.

[0032] Those skilled in the art can understand that, unless otherwise defined, all terms used herein (including technical terms and scientific terms) have the same meaning as the general understanding of those of ordinary skill in the art to which the present invention belongs. It should also be understood that terms such as those defined in a general dictionary should be understood to have a meaning consistent with the meaning in the context of the prior art, and will not be interpreted in an idealized or overly formal sense unless specifically defined as here.

[0033] The solution provided by the embodiments of the present application can be executed by any electronic device, such as a terminal device or a server. Among them, the server can be an independent physical server, a server cluster or a distributed system composed of multiple physical servers, or a cloud server providing cloud computing services. The terminal can be a smart phone, a tablet computer, a laptop computer, a desktop computer, a smart speaker, a smart watch, etc., but is not limited thereto. The terminal and the server can be directly or indirectly connected through wired or wireless communication methods, and the present application does not make any restrictions here. For the technical problems existing in the prior art, the robot manipulation learning method and device provided by the present application through predicting interaction are intended to solve at least one of the technical problems in the prior art.

[0034] The technical solution of the present application and how the technical solution of the present application solves the above technical problems will be described in detail below with specific embodiments. These several specific embodiments below can be combined with each other, and the same or similar concepts or processes may not be repeated in some embodiments. The embodiments of the present application will be described below with reference to the accompanying drawings.

[0035] The embodiments of the present application provide a possible implementation manner, such as Figure 1 shown, a flowchart of a robot manipulation learning method through predicting interaction is provided. This solution can be executed by any electronic device, and optionally, can be executed on the server side or the terminal device.

[0036] This method learns to manipulate (Manipulation by Predicting the Interaction, MPI) and enhances the visual representation through predicting interaction. As Figure 2As shown in (d), an interaction-oriented video prediction pre-training approach is adopted, and two main objectives are formulated, which are also the key elements constituting the interaction, summarized as "how to interact" and "where to interact". In terms of video input, a set of key frames (initial frame, transition frame, termination frame); in terms of the representation learning framework, the key image frames (key frames of the initial state and the termination state) are obtained and input into the MPI together with the language description (such as "Remove top of the blender"), a predictor is introduced to predict the frames of the transition state (transition frames), and a detector is introduced to detect the positions of the interacting objects in the transition frames. In terms of performance verification, real machine experiments, simulation environment experiments and reference index positioning are carried out to verify the performance of the method in the embodiments of the present application.

[0037] Specifically, as Figure 1 shown in, the method may include the following steps:

[0038] Step 101, determine the key frames of the initial state and the termination state during the interaction process, input them into the visual encoding module, and obtain the visual embedding sequence.

[0039] Step 102, segment the language description during the interaction process according to the positions in the predefined vocabulary set to obtain the language embedding sequence. Wherein, the language embedding sequence includes a series of numerical identifiers.

[0040] Based on the Transformer-based encoding-decoding architecture, in terms of the encoder, starting from a set of key frames, these key frames represent the initial state, the transition state and the termination state during the interaction process. Among them, the key frames of the initial state and the termination state (which can be called the initial frame and the termination frame) are used as the input of the visual encoder, while the key frames of the transition state are used as the prediction target, and the visual embedding sequence is obtained through encoding by the visual encoding module.

[0041] Step 103, based on the visual embedding sequence and the language embedding sequence, use the pre-constructed predictor to predict the unseen transition frames.

[0042] Step 104, use the pre-constructed detector to infer the positions of the interacting objects in the transition frames.

[0043] Among them, information exchange between the transition frames and the positions of the interacting objects is achieved through bidirectional attention.

[0044] In the embodiments of the present application, in terms of the decoder, a predictor and a detector for interaction frame prediction and interaction object detection are constructed. Through the information exchange between the predictor and the detector, the model can establish the potential relationship between "how to interact" and "where to interact", so as to optimize the two objectives.

[0045] In the embodiments of the present application, in terms of the encoder, the key frames of the initial state and the termination state in the interaction process are used as the input of the visual coding module to obtain a visual embedding sequence. The language description in the interaction process is segmented according to the positions in the predefined vocabulary set to obtain a series of numerical identifiers as the language embedding sequence. In the decoder, based on the visual embedding sequence and the language embedding sequence, a pre-constructed predictor is used to predict unseen transition frames, representing the interaction between the initial state and the termination state, enabling the model to understand "how to interact"; a pre-constructed detector is used to infer the positions of the interacting objects in the transition frames, enabling the model to obtain the knowledge of "where to interact". Among them, information exchange between the transition frames and the positions of the interacting objects is achieved through bidirectional attention, promoting mutual support in the training process. In the embodiments of the present application, interaction-oriented learning is introduced into representation learning to embed representations with manipulation task-specific knowledge, which can improve the performance of downstream robot applications. And the embodiments of the present application filter out information irrelevant to the interaction and emphasize the key states, thus realizing a centralized learning process.

[0046] In an alternative embodiment, determining the key frames of the initial state and the termination state in the interaction process and inputting them into the visual coding module to obtain a visual embedding sequence includes:

[0047] Using a visual encoder to divide the key frames of the initial state and the key frames of the termination state in the interaction process, respectively obtaining a plurality of non-overlapping block windows;

[0048] Converting the plurality of block windows corresponding to the key frames of the initial state into initial state visual embeddings, and converting the plurality of block windows corresponding to the key frames of the termination state into termination state visual embeddings;

[0049] In the embodiments of the present application, a decoupled multi-modal encoder is introduced. The multi-modal encoder plays a core role in learning representations, as Figure 3 shown in the light blue and orange parts. For the visual encoder design (such as Figure 3 the light blue part), the established Vision Transformer (ViT) framework is followed. The input frames (the key frames of the initial state and the termination state) F i are divided into non-overlapping windows of size s and then converted into a visual embedding sequence suitable for ViT processing Here, L vis represents the length of the visual embedding, and d represents the feature dimension. The value of L vis is equal to the number of divided windows HW / s 2 , where H and W respectively represent the height and width of the input frame.

[0050] Among them, "ViT" is a method proposed by Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, Jakob Uszkoreit, and Neil Houlsby in "An image is worth 16x16 words: Transformers for image recognition at scale" published at the International Conference on Learning Representations (ICLR).

[0051] Using the initial state visual embedding as the query, and the termination state visual embedding as both the key and the value simultaneously, capture the relationships between patch windows through the deformable attention layer, and fuse to obtain the visual embedding sequence.

[0052] To establish the state transition relationship between the initial state and the termination state of the input, introduce a causal reasoning module to promote the dynamic attention between the two states. The causal reasoning module takes the initial state visual embedding and the termination state visual embedding {v0, v1} ∈ RL vis × d encoded in parallel as the input. The initial state visual embedding is used as the query Q, while the final visual state embedding is used as both the key K and the value V. These encoded frames are processed through a deformable attention layer to capture the relationships between patches. This process can be interpreted as:

[0053] v′0 = DeformAttn(Q = v0, K = V = v1)

[0054] v = Norm(v0 + v′0)

[0055] The fused visual representation v ∈ RL vis × d encodes the causal relationship between the input states and is utilized in the decoder. In addition, this module reduces the computational burden in subsequent stages by merging the embeddings of the two frames into a unified representation. For downstream adaptation tasks with single-image input, simply copy the image features from ViT to generate the input of this module.

[0056] In an optional embodiment, tokenize the language description during the interaction process according to the positions in the predefined vocabulary set to obtain the language embedding sequence, including:

[0057] The language encoder is used to segment the language description in the interaction process according to the position in the predefined vocabulary set to obtain a series of numerical identifiers, which are filled to a preset maximum length to obtain a language embedding sequence.

[0058] The present application embodiment uses DistillBERT as a language encoder (such as Figure 3 The linguistic description of the interaction process (e.g., “Remove top of the blender”) is first tokenized into a series of numerical identifiers according to their positions in the predefined vocabulary set V and padded to a maximum length L. lang Among them, the maximum length L lang It is a preset constant. The segmented words obtained after segmentation can be filled with zeros to form a series of numerical identifiers and filled to the maximum length L lang The language encoder remains frozen (no parameter updates are performed). In addition, the visual encoder and language encoder are decoupled, so the language input can be omitted in downstream tasks.

[0059] Among them, "DistillBERT" is the method proposed by Victor Sanh, Lysandre Debut, Julien Chaumond, and Thomas Wolf in "DistilBERT, a distilled version of BERT: smaller, faster, cheaper and lighter" published on the arXiv preprint (arXiv preprint arXiv).

[0060] In an optional embodiment, before predicting an unseen transition frame using a pre-built predictor based on the visual embedding sequence and the language embedding sequence, the method further comprises:

[0061] Input the visual embedding sequence and the language embedding sequence into the feature clusterer, and use the feature clusterer to extract aggregation tokens;

[0062] Aggregate visual embedding and aggregate language embedding are generated by minimizing the distance between aggregated tokens.

[0063] In this embodiment of the present application, a token aggregator is introduced, which is inspired by two main considerations: 1) In order to address the challenge of inconsistent lengths between visual and language embeddings v,l, this embodiment of the present application aims to minimize the aggregation of tokens. The distance between them enables the visual encoder to generate language-aligned and semantically rich features; 2) by using aggregate tokens as a concise representation for downstream tasks, it facilitates fair comparison with existing methods, which naturally rely on aggregate features.

[0064] Multi-Head Attention Pooling (MAP) has proven to be more effective than other common methods (such as [cls] tokens or global pooling) in extracting compact representations from a series of feature embeddings. Based on this view, the embodiments of this application also employ a MAP block, shared between visual and language representations, for extracting aggregated tokens. The aggregation process starts with a latent embedding vector initialized with zeros, denoted as visual and language Taking the visual part as an example, this process can be formally expressed as

[0065]

[0066]

[0067] In an alternative embodiment, based on the visual embedding sequence and the language embedding sequence, a pre-constructed predictor is utilized to predict unseen transition frames, including:

[0068] Randomly initialize a set of prediction queries, with each prediction query corresponding to each patch window;

[0069] Perform cross-attention between the visual embedding sequence and the language embedding sequence, and use the prediction queries to estimate the transition frames.

[0070] In the embodiments of this application, the detailed structure of the Prediction Transformer is as Figure 3 shown in the dark blue part on the right. Given the visual embedding sequence and the language embedding sequence from the encoder, the Prediction Transformer is responsible for predicting the pixels of the frames representing unseen interaction states. Different from using queries to represent the masked regions in the input image in the original MAE, the prediction process in the embodiments of this application starts from a set of randomly initialized prediction queries beginning, and these queries correspond to each patch window of the target frame. To obtain the information required for accurate prediction, cross-attention is performed between the visual embedding sequence and the language embedding sequence {v, l} (achieved through the information interaction between the text cross-attention and the visual cross-attention in Figure 3 ), and q pred is used as the query.

[0071] Among them, "MAE" is the method proposed by Kaiming He, Xinlei Chen, Saining Xie, Yanghao Li, Piotr Dollar, and Ross Girshick in "Masked autoencoders are scalable vision learners" published at the IEEE International Conference on Computer Vision and Pattern Recognition (CVPR).

[0072] In an alternative embodiment, after performing cross-attention between the visual embedding sequence and the language embedding sequence and using the prediction query to estimate the transition frame, the method further includes:

[0073] Using a bi-directional attention mechanism to encapsulate semantic information related to the interaction object in the detection query, enhancing the prediction query.

[0074] After iterative processing within the multi-layer decoder, the prediction query generates a pixel-level prediction result through linear projection (using F), denoted as

[0075] In an alternative embodiment, using a pre-constructed detector to infer the position of the interaction object in the transition frame includes:

[0076] Initializing the detection query using the aggregated visual embedding sequence;

[0077] Using the information inferred from the prediction query by the detection query to regress the position of the interaction object.

[0078] Considering the object-centered nature of the interaction process, in the embodiments of the present application, a detection query q det ∈R d is introduced in the DetectionTransformer part to locate the interaction object, as shown in the purple part of Figure 3 . To accelerate the training convergence speed, the aggregated visual embedding is used to initialize the detection query. It should be noted that the detection query q det mainly regresses the object position based on the information inferred from the prediction query q pred . In the embodiments of the present application, it is expected that the prediction query contains comprehensive information about the target state. Information exchange is carried out through the bi-directional attention mechanism between the two sets of queries. Subsequently, a two-layer multi-layer perceptron (MLP) is applied to obtain the regression box:

[0079] Empirical analysis shows that although a slower convergence speed is observed in the early training stage due to insufficient information representation at both ends, the transmission of query information helps to eventually converge both tasks. The integrated modeling and optimization of these two tasks mutually enhance the model's ability to obtain more effective representations for robot operation.

[0080] Using the method in the embodiments of the present application for experiments, Figure 4(a) Schematic diagram for real machine experiments. The experimental contents include "removing the spatula from the cabinet shelf", "putting the pot into the sink", "putting the banana into the drawer", "lifting the pot lid", and "closing the drawer".

[0081] The success rate in a clean background is as Figure 4 (b) shown. The experimental contents include "Put the orange into basket", "Pick up bread", "Close laptop", "Scan code", "Push block", "Stack block", "Water roses", "Put croissant on the plate", "Pick up ice cream", "Put pepper on the plate". From left to right in the figure are the success rates obtained using DINO, R3M, MVP, Voltron, and MPI respectively. Among them, MPI is proposed in this application embodiment and Figure 4 is denoted as MPI (Ours) in the figure.

[0082] Among them, "DINO" is the method proposed by Mathilde Caron, Hugo Touvron, Ishan Misra, Herve Jegou, Julien Mairal, Piotr Bojanowski, and Armand Joulin in "Emerging properties in self-supervised vision transformers" published at the International Conference on Computer Vision (ICCV);

[0083] "R3M" is the method proposed by Suraj Nair, Aravind Rajeswaran, Vikash Kumar, Chelsea Finn, and Abhinav Gupta in "R3M: A universal visual representation for robot manipulation" published at CoRL;

[0084] "MVP" is the method proposed by Ilija Radosavovic, Tete Xiao, Stephen James, Pieter Abbeel, Jitendra Malik, and Trevor Darrell in "Real-world robot learning with masked visual pre-training" published on CoRL;

[0085] "Voltron" is the method proposed by Siddharth Karamcheti, Suraj Nair, Annie S Chen, Thomas Kollar, Chelsea Finn, Dorsa Sadigh, and Percy Liang in "Language-driven representation learning for robotics" published on RSS.

[0086] According to Figure 4 (b), although some models show excellent performance in specific tasks (e.g., Voltron performs well in the "Scan QR code" task), they cannot consistently maintain high performance in all tasks. In contrast, MPI proposed in the embodiments of this application focuses on interactive objects and obtains a more widely applicable representation, thus consistently maintaining improved performance in various tasks. In a more challenging kitchen environment, as Figure 4 (c) shows, including "Takespatula off the shelf", "Putpot into sink", "Putbanana into drawer", "Lift up the lid", "Close drawer", according to Figure 4 (c), MPI performs excellently in all manipulation tasks. Notably, it shows particular proficiency in tasks that require precise object perception to interact with small objects, such as "Lift up the lid" and "Take off the spatula". Figure 4 (d) shows the average success rate across tasks. Overall, as Figure 4 (d) shows, MPI outperforms previous methods (DINO, R3M, MVP, Voltron), with an average 26.3% increase in Success Rate in a wide range of 15 tasks.

[0087] Table 1. Success rate of generalization evaluation

[0088]

[0089]

[0090] Generalization Evaluation of Real Scenarios: To further evaluate the generalization ability and robustness of visual representations, two different evaluation scenarios are designed in a complex kitchen environment in the embodiments of this application. In the first scenario, the task is "put the banana into the drawer". Background interference is introduced by replacing the daisy with a rose. In addition, to test the generalization ability of various models to changes in the manipulated object, another detailed task is introduced in the embodiments of this application: "lift the pot lid", and the wooden pot is used instead of the white pot. The results are shown in Table 1. Obviously, when encountering out-of-distribution situations, the performance of all existing methods will decline. However, MPI shows significant robustness to this kind of interference, with only a slight decrease of 8.3% in the presence of background interference and a minimum decrease of 37.5% in terms of object changes. These results verify the generalization ability of the method in the embodiments of this application in unseen environments and manipulated objects.

[0091] Table 2 shows the results of single-task visual-motor control in the Franka Kitchen simulation environment. The success rate (%) on 50 randomly sampled trajectories is reported in the embodiments of this application. The best results of models with similar parameters are marked in bold, and the second-best results are underlined. "INSUP." represents classification supervised learning based on ImageNet. MPI always shows superior performance in multiple tasks.

[0092] Table 2. Franka Kitchen

[0093]

[0094] As shown in Table 2. It is worth noting that the representation learning framework customized for robot operation shows obvious advantages in the field of computer vision, exceeding two widely adopted visual pre-training methods: ImageNet classification and CLIP contrastive pre-training. Although the previous methods performed quite similarly, the embodiments of this application are implemented based on smaller and larger visual backbones, showing significant advantages compared with the prior art. In particular, MPI combined with the ViT-Small architecture has a 6% higher average success rate compared to the leading precedent Voltron. This improvement is further increased by 7.2% by using the ViT-Base backbone. In addition, the method in the embodiments of this application always shows the best or near-best performance in all tasks, highlighting its continuous advantages in dealing with complex scenarios encountered in Franka Kitchen.

[0095] Among them, "ImageNet" is a method proposed by Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Li Fei-Fei in ImageNet: A large-scale hierarchical image database, which was published at the IEEE Conference on Computer Vision and Pattern Recognition (CVPR).

[0096] "CLIP" is a method proposed by Alec Radford, JongWook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, Gretchen Krueger, and Ilya Sutskever in "Learning transferable visual models from natural language supervision", which was published at the International Conference on Machine Learning (ICML).

[0097] Table 3 shows the results of single-task visual-motor control in the Meta-World simulation environment. In the embodiments of the present application, the success rate (%) on 50 randomly sampled trajectories is reported. The best results are marked in bold, and the second-highest results are underlined. MPI demonstrated excellent performance in three tasks, with a higher average success rate compared to previous methods.

[0098] Table 3. Meta-World

[0099]

[0100] The experimental results of visual-motor control on Meta-World are presented in Table 3. Notably, a smaller variant of MPI achieved the highest success rate, exceeding MVP, even though the latter's visual backbone is approximately four times that of the embodiments of the present application. Compared with Voltron, which also uses ViT-Small, the leading advantage of the method in the embodiments of the present application is further expanded to 17%. The excellent performance demonstrated in Meta-World highlights the effective generalization ability of position object movement achieved by MPI.

[0101] Table 4 shows the results of Referring Expression Grounding. In the embodiments of this application, the average precision (%) of smaller variants is reported (R3M uses ResNet50, and MVP, Voltron, and the method proposed in the embodiments of this application use ViT-Small) at three IoU thresholds. In the embodiments of this application, the aggregated visual embeddings generated by the encoder are utilized MPI can provide the best detection results regardless of whether it uses full-length visual embeddings or aggregated embeddings.

[0102] Table 4. Referring Expression Grounding

[0103]

[0104] The experimental results of the Referring Expression Grounding task are listed in Table 4. Although R3M has certain effects in visual motion control, it can be seen from its average precision of only 42.66% at 0.75 IoU that it has difficulties in accurately locating objects according to language descriptions. This limitation may stem from its pre-training objective, which only relies on contrastive loss to obtain high-level semantic features and loses low-level localization features after global average pooling. In contrast, the method proposed in the embodiments of this application achieves superior average precision with embeddings of the same dimension, compared with the previous leading method MVP. In the embodiments of this application, it is observed that at three different IoUs, the precision is increased by +3.22%, +6.78%, and +11.5% respectively. In addition, the token aggregator learned from pre-training enables the model to achieve comparable or even better performance with fewer embeddings, demonstrating the effectiveness of the interaction-oriented learning objective in the embodiments of this application.

[0105] Based on the same principle as the method provided in the embodiments of this application, the embodiments of this application also provide a robot manipulation learning device for predicting interactions, such as Figure 5 shown, the device includes:

[0106] A visual encoding module 501, configured to determine key frames of the initial state and the termination state during the interaction, input them into the visual encoding module, and obtain a sequence of visual embeddings;

[0107] A language encoding module 502, configured to segment the language description during the interaction according to positions in a predefined vocabulary set to obtain a sequence of language embeddings; wherein, the sequence of language embeddings includes a series of numerical identifiers.

[0108] The interaction mode prediction module 503 is configured to predict unseen transition frames based on the visual embedding sequence and the language embedding sequence by using a pre-constructed predictor;

[0109] The interaction position prediction module 504 is configured to infer the positions of interacting objects in the transition frames by using a pre-constructed detector;

[0110] Among them, information exchange between the transition frames and the positions of the interacting objects is achieved through bidirectional attention.

[0111] In the embodiments of the present application, in terms of the encoder, the key frames of the initial state and the termination state in the interaction process are used as the input of the visual encoding module to obtain a visual embedding sequence. The language descriptions in the interaction process are segmented according to the positions in the predefined vocabulary set to obtain a series of numerical identifiers as the language embedding sequence. In the decoder, based on the visual embedding sequence and the language embedding sequence, a pre-constructed predictor is used to predict unseen transition frames, representing the interaction between the initial state and the termination state, so that the model understands "how to interact"; a pre-constructed detector is used to infer the positions of the interacting objects in the transition frames, so that the model can obtain the knowledge of "where to interact". Among them, information exchange between the transition frames and the positions of the interacting objects is achieved through bidirectional attention, promoting mutual support in the training process. In the embodiments of the present application, interaction-oriented learning is introduced into representation learning to embed representations with manipulation task-specific knowledge, which can improve the performance of downstream robot applications. And the embodiments of the present application filter out information irrelevant to the interaction and emphasize the key states, thereby realizing a concentrated learning process.

[0112] The robot manipulation learning device for predicting interaction provided by the embodiments of the present application can implement Figures 1 to 4 each process implemented in the method embodiments. To avoid repetition, it will not be elaborated here.

[0113] The robot manipulation learning device for predicting interaction in the embodiments of the present application can execute the robot manipulation learning method for predicting interaction provided by the embodiments of the present application, and its implementation principle is similar. The actions performed by each module and unit in the robot manipulation learning device for predicting interaction in the embodiments of the present application correspond to the steps in the robot manipulation learning method for predicting interaction in the embodiments of the present application. For the detailed function descriptions of each module of the robot manipulation learning device for predicting interaction, reference can be specifically made to the descriptions in the corresponding robot manipulation learning method for predicting interaction shown above, which will not be elaborated here.

[0114] Based on the same principle as the method shown in the embodiments of the present application, the embodiments of the present application also provide an electronic device, which may include but is not limited to: a processor and a memory; the memory is used to store a computer program; the processor is used to execute the robot manipulation learning method through predictive interaction shown in any optional embodiment of the present application by calling the computer program. Compared with the prior art, in the encoder of the robot manipulation learning method through predictive interaction provided by the present application, the key frames of the initial state and the termination state in the interaction process are used as the input of the visual coding module to obtain a visual embedding sequence, and the language description in the interaction process is segmented according to the positions in the predefined vocabulary set to obtain a series of numerical identifiers as the language embedding sequence. In the decoder, based on the visual embedding sequence and the language embedding sequence, a pre-constructed predictor is used to predict unseen transition frames, representing the interaction between the initial state and the termination state, so that the model can understand "how to interact"; a pre-constructed detector is used to infer the positions of the interacting objects in the transition frames, so that the model can obtain the knowledge of "where to interact". Among them, information exchange between the transition frames and the positions of the interacting objects is realized through bidirectional attention, promoting mutual support in the training process. In the embodiments of the present application, interaction-oriented learning is introduced into representation learning to embed representations with task-specific knowledge of manipulation, which can improve the performance of downstream robot applications. And the embodiments of the present application filter out information irrelevant to the interaction and emphasize the key states, thereby realizing a centralized learning process.

[0115] In an optional embodiment, an electronic device is also provided, as Figure 6 shown Figure 6 The electronic device 600 shown may be a server, including: a processor 601 and a memory 603. Among them, the processor 601 and the memory 603 are connected, such as through a bus 602. Optionally, the electronic device 600 may further include a transceiver 604. It should be noted that in practical applications, the transceiver 604 is not limited to one, and the structure of the electronic device 600 does not constitute a limitation to the embodiments of the present application.

[0116] The processor 601 may be a CPU (Central Processing Unit), a general-purpose processor, a DSP (Digital Signal Processor), an ASIC (Application Specific Integrated Circuit), an FPGA (Field Programmable Gate Array), or other programmable logic devices, transistor logic devices, hardware components, or any combination thereof. It can implement or execute various exemplary logical blocks, modules, and circuits described in connection with the disclosure of this application. The processor 601 may also be a combination that implements computing functions, such as a combination of one or more microprocessors, a combination of a DSP and a microprocessor, etc.

[0117] The bus 602 may include a path for transmitting information between the above components. The bus 602 may be a PCI (Peripheral Component Interconnect) bus or an EISA (Extended Industry Standard Architecture) bus, etc. The bus 602 may be divided into an address bus, a data bus, a control bus, etc. For the sake of representation, Figure 6 only a thick line is used to represent it in the figure, but it does not mean that there is only one bus or one type of bus.

[0118] The memory 603 may be a ROM (Read Only Memory) or other types of static storage devices that can store static information and instructions, a RAM (Random Access Memory) or other types of dynamic storage devices that can store information and instructions, or it may also be an EEPROM (Electrically Erasable Programmable Read Only Memory), a CD-ROM (Compact Disc Read Only Memory), or other optical disc storage, optical disc storage (including compact discs, laser discs, optical discs, digital versatile discs, Blu-ray discs, etc.), magnetic disk storage media, or any other medium that can be used to carry or store the desired program code in the form of instructions or data structures and can be accessed by a computer, but is not limited thereto.

[0119] The memory 603 is used to store the application program code for executing the solution of this application, and is controlled by the processor 601 for execution. The processor 601 is used to execute the application program code stored in the memory 603 to implement the content shown in the foregoing method embodiments.

[0120] Among them, the electronic device includes but is not limited to: mobile terminals such as mobile phones, laptop computers, digital broadcast receivers, PDAs (Personal Digital Assistants), PADs (Tablet Computers), PMPs (Portable Multimedia Players), in-vehicle terminals (such as in-vehicle navigation terminals), etc., and fixed terminals such as digital TVs, desktop computers, etc. Figure 6 The electronic device shown is only an example and should not impose any limitations on the functions and usage scope of the embodiments of this application.

[0121] The server provided by this application can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers. It can also be a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN, as well as big data and artificial intelligence platforms. The terminal can be a smart phone, a tablet computer, a laptop computer, a desktop computer, a smart speaker, a smart watch, etc., but is not limited thereto. The terminal and the server can be directly or indirectly connected through wired or wireless communication methods, and this application does not make any restrictions here.

[0122] The embodiments of this application provide a computer-readable storage medium, on which a computer program is stored. When it runs on a computer, it enables the computer to execute the corresponding content in the foregoing method embodiments.

[0123] It should be understood that although the steps in the flowchart of the accompanying drawings are shown in sequence according to the indication of the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless there is a clear indication in this article, the execution of these steps has no strict order limit, and they can be executed in other orders. Moreover, at least a part of the steps in the flowchart of the accompanying drawings may include multiple sub-steps or multiple stages. These sub-steps or stages are not necessarily executed at the same time, but can be executed at different times. Their execution order is not necessarily sequential, but can be executed alternately or alternately with at least a part of other steps or sub-steps or stages of other steps.

[0124] It should be noted that the above computer-readable storage medium in the present application can also be a computer-readable signal medium or a combination of a computer-readable storage medium and a computer-readable storage medium. The computer-readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples of the computer-readable storage medium can include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present application, the computer-readable storage medium can be any tangible medium that contains or stores a program, and this program can be used by or in combination with an instruction execution system, apparatus, or device. In the present application, the computer-readable signal medium can include a data signal propagated in a baseband or as part of a carrier wave, which carries computer-readable program code. Such a propagated data signal can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. The computer-readable signal medium can also be any computer-readable medium other than the computer-readable storage medium, and this computer-readable signal medium can send, propagate, or transmit a program for use by or in combination with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted by any appropriate medium, including but not limited to: wires, optical cables, RF (radio frequency), etc., or any suitable combination of the above.

[0125] The above computer-readable medium can be included in the above electronic device; or it can exist separately without being assembled into the electronic device.

[0126] The above computer-readable medium carries one or more programs, and when the above one or more programs are executed by the electronic device, the electronic device is caused to execute the method shown in the above embodiments.

[0127] According to one aspect of the present application, there is provided a computer program product or a computer program, which includes computer instructions stored in a computer-readable storage medium. The processor of the computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, causing the computer device to execute the method and apparatus for robot control learning through predictive interaction provided in the above various alternative implementation manners.

[0128] Computer program code for performing the operations of this application can be written in one or more programming languages or combinations thereof. The above-mentioned programming languages include object-oriented programming languages such as Java, Smalltalk, C++, and also include conventional procedural programming languages such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, executed as an independent software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer can be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or it can be connected to an external computer (for example, by using an Internet service provider to connect through the Internet).

[0129] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of this application. In this regard, each block in the flowchart or block diagram can represent a module, a program segment, or a part of the code, and this module, program segment, or part of the code contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks can occur in a different order than that marked in the accompanying drawings. For example, two consecutively represented blocks can actually be executed substantially in parallel, and they can sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, and the combination of blocks in the block diagram and / or flowchart, can be implemented by a dedicated hardware-based system for performing the specified functions or operations, or can be implemented by a combination of dedicated hardware and computer instructions.

[0130] The modules described in the embodiments of this application can be implemented in software or in hardware. Among them, the name of the module does not constitute a limitation to the module itself in some cases. For example, the visual coding module can also be described as "the visual coding module for determining the key frames of the initial state and the termination state during the interaction process, inputting the visual coding module, and obtaining the visual embedding sequence".

[0131] The above description is only a preferred embodiment of the present application and an explanation of the applied technical principles. Those skilled in the art should understand that the scope of disclosure involved in the present application is not limited to the technical solutions formed by the specific combination of the above technical features, but should also cover other technical solutions formed by any combination of the above technical features or their equivalent features without departing from the above disclosure concept. For example, the technical solutions formed by mutually replacing the above features with the technical features (but not limited to) disclosed in the present application that have similar functions.

Claims

1. A robot manipulation learning method through predictive interaction, characterized in that: The method comprises: Determine the key frames of the initial state and the terminal state during the interaction process, input them into the visual encoding module, and obtain the visual embedding sequence; Segmenting the language description in the interaction process according to the position in the predefined vocabulary set to obtain a language embedding sequence; the language embedding sequence includes a series of numerical identifiers; Predicting unseen transition frames using a pre-built predictor based on the visual embedding sequence and the language embedding sequence; Inferring the position of the interactive object in the transition frame using a pre-built detector; wherein information exchange between the transition frame and the position of the interactive object is achieved through bidirectional attention; The key frames of the initial state and the terminal state in the interaction process are determined and input into the visual encoding module to obtain the visual embedding sequence, including: Using a visual encoder, the key frames of the initial state and the key frames of the terminal state in the interaction process are divided to obtain a plurality of non-overlapping block windows respectively; Converting a plurality of block windows corresponding to the key frame of the initial state into an initial state visual embedding, and converting a plurality of block windows corresponding to the key frame of the terminal state into a terminal state visual embedding; The initial state visual embedding is used as a query, and the terminal state visual embedding is used as both a key and a value. The relationship between the block windows is captured through a deformable attention layer, and the visual embedding sequence is obtained by fusion.

2. The robot manipulation learning method through predictive interaction according to claim 1, characterized in that: The step of segmenting the language description in the interaction process according to the position in the predefined vocabulary set to obtain the language embedding sequence includes: The language description in the interaction process is segmented by a language encoder according to the position in a predefined vocabulary set to obtain a series of numerical identifiers, which are filled to a preset maximum length to obtain the language embedding sequence.

3. The robot manipulation learning method through predictive interaction according to claim 2, characterized in that: Before predicting an unseen transition frame using a pre-built predictor based on the visual embedding sequence and the language embedding sequence, the method further comprises: Inputting the visual embedding sequence and the language embedding sequence into a feature clusterer, and extracting aggregate tokens using the feature clusterer; By minimizing the distance between the aggregated tokens, an aggregated visual embedding and an aggregated language embedding are generated.

4. The robot manipulation learning method through predictive interaction according to claim 3, characterized in that: The predicting of unseen transition frames using a pre-built predictor based on the visual embedding sequence and the language embedding sequence comprises: Randomly initialize a group of prediction queries, each of the prediction queries corresponds to each of the block windows; Cross-attention is performed between the visual embedding sequence and the language embedding sequence, and the transition frame is estimated using the predicted query.

5. The robot manipulation learning method through predictive interaction according to claim 4, characterized in that: The inferring the position of the interactive object in the transition frame by using a pre-built detector comprises: Initializing a detection query using the aggregated visual embedding sequence; The position of the interactive object is regressed using information inferred from the prediction query by the detection query.

6. The robot manipulation learning method through predictive interaction according to claim 4, characterized in that: After performing cross attention between the visual embedding sequence and the language embedding sequence and estimating the transition frame using the predicted query, the method further includes: A bidirectional attention mechanism is used to encapsulate semantic information related to the interacting objects in the detection query.

7. A robot control learning device through predictive interaction, characterized in that: The device comprises: The visual encoding module is used to determine the key frames of the initial state and the terminal state during the interaction process, and input them into the visual encoding module to obtain the visual embedding sequence; A language encoding module, used to segment the language description in the interaction process according to the position in the predefined vocabulary set to obtain a language embedding sequence; the language embedding sequence includes a series of numerical identifiers; An interaction mode prediction module, configured to predict unseen transition frames using a pre-built predictor based on the visual embedding sequence and the language embedding sequence; An interactive position prediction module, used for inferring the position of the interactive object in the transition frame using a pre-built detector; wherein information exchange between the transition frame and the position of the interactive object is achieved through bidirectional attention; The key frames of the initial state and the terminal state in the interaction process are determined and input into the visual encoding module to obtain the visual embedding sequence, including: Using a visual encoder, the key frames of the initial state and the key frames of the terminal state in the interaction process are divided to obtain a plurality of non-overlapping block windows respectively; Converting a plurality of block windows corresponding to the key frame of the initial state into an initial state visual embedding, and converting a plurality of block windows corresponding to the key frame of the terminal state into a terminal state visual embedding; The initial state visual embedding is used as a query, and the terminal state visual embedding is used as both a key and a value. The relationship between the block windows is captured through a deformable attention layer, and the visual embedding sequence is obtained by fusion.

8. An electronic device, characterized in that: The method comprises a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the method according to any one of claims 1 to 6 when executing the program.

9. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the method according to any one of claims 1 to 6 is implemented.

Citation Information

Patent Citations

  • Video thumbnail recommendation method fusing visual semantic information

    CN111680190A

  • Robot fixed-point welding auxiliary device

    CN114643447A