Generalized robot operation method and system based on segmentation mask representation
By adopting a method based on segmentation mask representation in robot operation technology, using a multimodal large model to generate segmentation masks and guide the training of operation strategy networks, the problem of insufficient generalization ability in the existing technology is solved, and efficient and generalizable robot operations are achieved.
Patent Information
- Application Number
- CN202510170495.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-17
- Publication Date
- 2025-06-06
- Estimated Expiration
- 2045-02-17
AI Technical Summary
When facing new scenarios or tasks, existing robot operation technologies lack generalization capabilities and need to rely on large-scale data sets and additional model fine-tuning, resulting in high costs and limited scalability of the model in practical applications.
The generalized robot operation method based on segmentation mask representation is adopted. By constructing a three-dimensional object library and a desktop scene library, the desktop scene layout is generated in a virtual environment. Combining the robot's perspective image and text instructions, a pre-trained multi-modal big model is used to generate the segmentation mask of the target object and region, and fuse it with the robot's perspective image features to guide the training and execution of the robot's operation strategy network.
It significantly improves the robot's generalization ability in multi-tasking and multi-scenarios, improves operation accuracy and scenario adaptability, reduces dependence on large-scale data sets and model fine-tuning, reduces costs and improves the scalability of the model.
Smart Images

Figure CN120107583A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of robot operation in desktop scenes, and in particular to a generalizable robot operation method and system based on segmentation mask representation. Background Art
[0002] Currently, the field of robot manipulation has become an important research direction in artificial intelligence and robotics, especially in complex desktop scenarios. This field aims to give robots efficient object manipulation capabilities by combining visual, language and motion data, covering applications such as object grasping, precise placement, and task execution in complex scenarios. Realizing universal operation of robots in diverse tasks and scenarios is the key to promoting the further development of robotics.
[0003] Existing robotic manipulation techniques can be roughly divided into the following three categories: 1) imitation learning models, 2) vision-goal-based policy models, and 3) Vision-Language-Action (VLA) models.
[0004] Imitation learning models collect a large amount of artificial demonstration data for training, focus on specific tasks or scenarios, and can achieve good performance under specific conditions, but are limited by the diversity of data and the limitations of the model, and their generalization ability is poor; visual target-based policy models usually use target images or predefined rules to provide target guidance for robots. Such models perform well in spatial positioning and operation accuracy, but rely on high-quality data and specific task settings, and are difficult to adapt to more complex dynamic scenarios; VLA models combine large-scale robot training data and pre-trained vision-language models (Vision-Language Models, VLMs) to demonstrate strong task reasoning capabilities and a wide range of operational adaptability, and can support a variety of robot operation tasks. However, these methods still face the challenge of insufficient generalization ability when facing new scenarios or tasks, and usually need to rely on large-scale data sets and additional model fine-tuning, which is not only costly, but also limits the scalability of the model in practical applications. Summary of the invention
[0005] In order to solve the problem of insufficient generalization ability in existing robot operation technology, the present invention provides a generalizable robot operation method and system based on segmentation mask representation, aiming to achieve the generalization ability of robot operation strategies in diverse scenarios and tasks.
[0006] The specific technical solution adopted by the present invention is:
[0007] In a first aspect, the present invention provides a generalizable robot operation method based on segmentation mask representation, wherein the robot is composed of a gripper and a mechanical arm, and comprises the following steps:
[0008] (1) Build a 3D object library and a desktop scene library, and randomly generate a large number of desktop scene layouts in virtual environments based on the data in the library;
[0009] (2) Generate robot operation trajectory data for the target object in the desktop scene by giving the target object and the target area; for each robot operation trajectory data, generate diversified text instructions by combining the appearance, spatial position relationship and common sense knowledge of all objects in the desktop scene;
[0010] (3) For each robot operation trajectory data, collect the robot view image, robot state data and a text instruction at each operation step as a training sample to construct a training sample set; the robot state data includes the robot's gripper switch and the joint angle of the robot arm;
[0011] (4) Using the pre-trained multimodal large model, the target object and target area indicated by the text instruction in each training sample are located to obtain the target object mask and the target area mask;
[0012] (5) inputting the target object mask, target area mask, and training samples corresponding to several historical operation steps into the robot operation strategy network, extracting the robot view image features, robot state features, and text instruction features, fusing the target object mask and target area mask with the robot view image features, predicting the action instructions of the robot's next operation step based on the fusion results and the robot state features, text instruction features, and action features corresponding to the learnable action token, and calculating the loss according to the actual state of the robot at the next operation step and the predicted action instructions to train the robot operation strategy network;
[0013] (6) Use the pre-trained multimodal large model and the trained robot operation strategy network to complete given instructions in actual desktop scenarios.
[0014] Furthermore, the desktop scene layout in the virtual environment includes a desktop scene, a target object and a plurality of interference objects.
[0015] Furthermore, the objects in the three-dimensional object library contain attribute information, and the attribute information includes category, color, shape and material.
[0016] Furthermore, in step (2), each robot operation trajectory data corresponds to a number of text instructions.
[0017] Furthermore, the text instructions are implemented by a large language model, and the input of the large language model includes the target object, the target area, the robot's initial perspective image, all object attribute information in the desktop scene, and related prompt words.
[0018] Furthermore, the pre-trained multimodal large model described in step (4) includes an image encoder, a multi-layer perceptron, a large language model and a positioning module;
[0019] Step (4) specifically includes:
[0020] (4-1) Use the image encoder to obtain the robot's initial view image x in the training sample v,0 Encode features and project them into the embedding feature space of the large language model through a multi-layer perceptron;
[0021] (4-2) Initializing a positioning module by a pre-trained SAM model, wherein the positioning module includes a pre-trained image encoder and a pre-trained image decoder, and the positioning module is connected after the large language model;
[0022] (4-3) Add special tokens to the vocabulary of the large language model <seg>, and given a large language model prompt word, the large language model is required to locate the target object and target area referred to in the text instruction according to the input image encoding features and text instructions, and generate text descriptions respectively Where CLIP(·) represents the image encoder, f v (·) represents a multi-layer perceptron, x t Represents a text instruction, represents a large language model, y t,1 Indicates the text description of the target object referred to in the text instruction. t,2 Indicates the text description of locating the target area referred to in the text instruction;
[0023] (4-4) When the large language model text output contains tokens <seg>When the last hidden layer feature vector before the text output is input into the positioning module, the corresponding segmentation mask is obtained. Among them, M o is the target object mask, M p is the target area mask, and both segmentation masks are 0-1 mask matrices with the same size as the robot view image; ε(·) represents the pre-trained image encoder in the localization module, represents the pre-trained image decoder in the localization module, Represents y t,1 Output the last hidden layer feature vector corresponding to the previous one, Represents y t,2 Output the corresponding last hidden layer feature vector.
[0024] Furthermore, the robot operation strategy network includes a pre-trained image encoder, a pre-trained text encoder, a multi-layer perceptron, a positioning perceptron, and a Transformer decoder, wherein the positioning perceptron is composed of several attention layers;
[0025] Step (5) specifically includes:
[0026] (5-1) For each operation step, the robot view image is imaged using the pre-trained image encoder in the robot operation strategy network to obtain the robot view image features. Includes a global image feature and a local image feature Among them, D v The feature dimension encoded for the image;
[0027] (5-2) The positioning sensor initializes a global query vector A target object query vector and a target region query vector Where D p is the initial feature dimension;
[0028] (5-3) In the first attention layer of the localization perceptron, the three vectors are concatenated together and projected into the hidden layer space to obtain the query vector Where d is the hidden layer feature dimension; robot view image feature After being projected by different projection matrices, they are concatenated with the query vector Q to obtain the key vectors Sum value vector Calculate the attention matrix based on the query vector Q and the key vector K
[0029] (5-4) Mask the target object M o and the target region mask M p They are mapped to a feature map size of 14×14 and converted into a one-dimensional vector of size 1×196, and then M o The corresponding one-dimensional vector is applied to A [1,:196] , M p The corresponding one-dimensional vector is applied to A [2,:196] , so that the attention value of the mask area is replaced by the maximum value of the current matrix, and the updated attention matrix A is obtained ′ ; Next, calculate A ′ The softmax of A is multiplied by the value vector V and passed through the feedforward network FFN(·) to get the output of the attention layer O = FFN(softmax(A ′ )×V);
[0030] (5-5) Return to step (5-3) and use the output O of the previous attention layer as the query vector calculated by the next attention layer Until the output of the last attention layer is obtained, the image features of the final output fused with mask information are recorded as
[0031] (5-6) Use the pre-trained text encoder and multi-layer perceptron to extract the robot state feature Z t , text instruction feature Z s , and initialize a learnable action token <act>And extract the action feature Z a ; Get the input sequence corresponding to each operation step
[0032] (5-7) The above input sequence corresponding to the N historical operation step data is input into the Transformer decoder to predict the next action instruction, where the action instruction includes the gripper action and the robot arm action.
[0033] Furthermore, step (6) specifically includes:
[0034] (6-1) In an actual desktop scenario, a text instruction is given to require the robot to take the target object to the specified target area; the robot's initial view image and the given text instruction are input into the pre-trained multimodal large model to generate a target object mask and a target area mask;
[0035] (6-2) Maintain an input sequence consisting of N historical operation step data When the historical data is less than N operation steps, the latest historical operation step is copied as a supplement; the input sequence corresponding to each operation step is obtained in the same way as the training stage, and the next action instruction is generated by the trained robot operation strategy network. The process is repeated until the robot completes the task of the given text instruction.
[0036] Furthermore, the pre-trained image encoder in the robot operation strategy network adopts a ViTMAE image encoder.
[0037] In a second aspect, the present invention proposes a generalizable robot operating system based on segmentation mask representation, which is used to implement the above-mentioned generalizable robot operating method based on segmentation mask representation.
[0038] Compared with the prior art, the present invention has the following beneficial effects:
[0039] The present invention is a generalizable robot operation method and system based on segmentation mask representation. During implementation, the present invention utilizes a multimodal large model to generate a segmentation mask, and uses this as an intermediate representation to guide the training and execution of the robot operation strategy.
[0040] (1) By using a multimodal large model, the present invention can jointly extract semantic information and visual features from the robot's view image and text instructions to generate segmentation masks of the target object and the placement area. As a pre-trained general knowledge system, the multimodal large model can effectively adapt to different tasks and scenarios, provide high-precision object positioning and area annotation, and provide reliable input for subsequent operation strategies.
[0041] (2) By using segmentation mask representation, the present invention combines the target object and scene masks generated by the multimodal large model with the robot view image and robot state information to provide accurate spatial and semantic guidance for the operation strategy. The segmentation mask not only clarifies the position and shape of the target object, but also significantly enhances the robot's operation accuracy and scene adaptability in complex tasks.
[0042] (3) The present invention can automatically generate large-scale training sets containing diverse objects, complex scenes, and rich task instructions to improve the generalization ability of the model. The segmentation masks generated by the multimodal large model are combined with these diverse data to effectively enhance the adaptability of the operation strategy to unknown tasks and scenes.
[0043] In summary, the present invention significantly improves the robot's generalization ability in multiple tasks and scenarios by using a multimodal large model to generate segmentation masks and using them as intermediate representations to guide the learning and execution of robot operation strategies. While improving operation accuracy, this method can adapt to diverse task instructions and complex scenarios, providing a new technical solution for achieving efficient and generalizable robot operations. BRIEF DESCRIPTION OF THE DRAWINGS
[0044] Figure 1 is a flow chart of a generalizable robot operation method based on segmentation mask representation of the present invention;
[0045] Figure 2 It is a schematic diagram of a multimodal large model;
[0046] Figure 3 It is a schematic diagram of the robot operation strategy network. DETAILED DESCRIPTION
[0047] The present invention will be further described and illustrated below in conjunction with the accompanying drawings and specific implementation methods.
[0048] like Figure 1 As shown, the present invention provides a generalizable robot operation method based on segmentation mask representation, comprising the following steps:
[0049] Step 1: Obtain a variety of three-dimensional objects and their desktop scene layouts, randomly select target objects and interference objects from the three-dimensional objects and place them in the target desktop scene, and record the robot's initial view image; here, the robot's initial view image contains the target object;
[0050] The robot operation trajectory data for the target object in the target desktop scene is generated by using an automated trajectory generation algorithm, and a variety of text instructions are generated by combining the appearance, spatial position relationship and common sense knowledge of all objects in the target desktop scene; here, each robot operation trajectory data corresponds to a number of text instructions;
[0051] Step 2: For each robot operation trajectory data, collect the robot view image, robot state data and a text instruction at each operation step as a training sample to construct a training set;
[0052] Step 3: Use the pre-trained multimodal large model to locate the target object and target area indicated by the text instruction in each training sample, and obtain the target object mask and target area mask;
[0053] Step 4: Input the target object mask, target region mask, and training samples into the downstream robot operation strategy network, use the pre-trained image encoder to extract the robot view image features, use the attention mechanism to fuse the target object mask, target region mask, and robot view image features, and use the Transformer decoder to predict the robot's next action instruction, which includes the gripper action and the robotic arm action;
[0054] Step 5: For the predicted action instructions, calculate the L1 loss of the gripper action and the binary cross entropy loss of the robotic arm action respectively, and use this loss to jointly train the robot operation strategy network to improve its adaptability in multiple tasks and multiple scenarios.
[0055] The robot operation strategy network of the present invention performs cross-modal joint training on automatically generated diverse data sets, utilizes the generalization ability of multimodal large models and the precise spatial guidance of segmentation masks, and realizes a generalizable robot operation method that can maintain efficient performance in complex scenarios and diverse tasks.
[0056] Step 6: Combine the pre-trained multimodal large model and the trained robot operation strategy network to complete the given instructions in the actual desktop scenario.
[0057] In the above step 1, the automated robot operation data generation process is implemented as follows:
[0058] 1.1) Collect high-quality 3D object assets: In this embodiment, for the 760K 3D objects in the public dataset Objaverse, their multi-view images are rendered, and the GPT-4o large model is used to generate detailed description information of the objects, including attribute information such as category, color, shape, and material. 3526 objects suitable for desktop scene robot operation are screened out through rules.
[0059] 1.2) Use the open source robot operation simulator RoboCasa to generate a variety of desktop scene layouts, and randomly select the target object A and 3 to 10 interference objects B from the three-dimensional object assets and place them in the scene. Here, RoboCasa will automatically mark some reference tracks, that is, the robot's gripper action and mechanical arm joint action when an object is placed in a certain area. The reference track here is used to provide a reference for the downstream MimicGen. It should be noted that the reference track does not need to be for the target object A. If sampling other solutions to generate desktop scene layouts, you can manually mark some reference tracks.
[0060] 1.3) Using the open source automated trajectory generation solution MimicGen, the robot operation trajectory data of the target object in the target desktop scene is generated based on the reference trajectory. In the automated trajectory generation process, the target area can be given; by updating the desktop scene layout, target object and interference object in step 1.2), a large amount of robot operation trajectory data can be generated.
[0061] Here, generating robot operation trajectory data may also be achieved by other automated methods.
[0062] 1.4) For each robot operation trajectory data, the GPT-4o large model is used to generate text instructions for appearance, spatial position relationship and common sense knowledge based on the target object, target area, robot initial view image, and all object attribute information in the desktop scene; diversified text instructions help improve the model's generalization ability for different instructions.
[0063] For example, for a direct instruction: "Bring the water cup to the seat", a text instruction targeting appearance can be generated: "Bring the transparent object with a handle to the seat", a text instruction targeting spatial position relationship can be generated: "Bring the object on the left of the kettle to the seat", and a text instruction targeting common sense knowledge can be generated: "I want to drink water".
[0064] In the above step 2, under each robot operation trajectory data, the state of the robot in each unit time can be obtained. The state refers to the robot's gripper switch and the joint angle of the robot arm. Taking the robot used in this embodiment as an example, the robot state includes a gripper switch with 1 degree of freedom and a joint angle of the robot arm with 7 degrees of freedom. In addition, the robot view image can be obtained in each unit time. The robot view image and robot state during the operation process are recorded, and the robot view image, robot state data and a text instruction under each operation step are collected as a training sample to construct a training set.
[0065] In step 3 above, if Figure 2 As shown in FIG. 1 , the steps for obtaining the segmentation mask corresponding to the target object and the target area are specifically as follows:
[0066] 3.1) The robot's initial view image x in the training data v and the text instruction x t Input a pre-trained multimodal large model, wherein the multimodal large model includes a CLIP image encoder, a multi-layer perceptron, a large language model, and a positioning module;
[0067] 3.2) Use CLIP image encoder to obtain image encoding features F v = CLIP(x v ), using multi-layer perceptron f v (·) Project it to the embedding feature space of the large language model and get Where D t is the embedding feature dimension;
[0068] 3.3) In order to locate the target object or area, an additional localization module is added after the large language model, which consists of a pre-trained image encoder ε(·) and a pre-trained image decoder In this embodiment, the pre-trained image encoder ε(·) and the pre-trained image decoder Initialized by the pre-trained SAM model;
[0069] 3.4) Add special tokens to the vocabulary of the large language model <seg>, and given a large language model prompt word, the large language model input content is required to locate the target object and target area referred to in the text instruction, and generate text descriptions respectively; under the prompt word, the large language model According to the image encoding features and text instructions x t Input, get text output When the text output of the large language model contains tokens <seg>When the corresponding last layer feature vector F seg Input positioning module; since the target object and target area models will output two texts respectively, the two texts need to be judged separately whether they contain tokens <seg>, and two different eigenvectors Input the positioning module separately to get the corresponding segmentation mask Among them, M o is the target object mask, M p is the target area mask, and both segmentation masks are 0-1 mask matrices with the same size as the robot's view image.
[0070] In step 4 above, if Figure 3 As shown, the robot operation strategy network includes a pre-trained image encoder, a pre-trained text encoder, a multi-layer perceptron, a positioning perceptron, and a Transformer decoder. The positioning perceptron is composed of several attention layers. An optional implementation process of the robot operation strategy network is as follows:
[0071] 4.1) For each input operation step, the robot view image x v , use the pre-trained image encoder ViTMAE to obtain image encoding features Where D v is the feature dimension of image encoding, which contains the global CLS image feature vector and the local 14×14 feature map embedding vector
[0073] 4.2) The localization sensor initializes a global query vector A target object query vector and a target region query vector Where D p is the initial feature dimension;
[0074] 4.3) In order to parallelize the processing, in the first attention layer, the three vectors are concatenated together and projected into the hidden layer space to obtain the query vector Where d is the hidden layer feature dimension, and the feature map embedding vector of the robot's view image The key vector is obtained by concatenating the projection with the query vector Q Sum value vector Calculate the attention matrix based on the query vector Q and the key vector K
[0075] 4.4) In order to introduce the segmentation mask information into the attention matrix, the segmentation masks of the target object and the region are mapped to the feature map size of 14×14 respectively, and converted into a one-dimensional vector of size 1×196, and then M o The corresponding one-dimensional vector is applied to A [1,:196] , M p The corresponding one-dimensional vector is applied to A [2,:196] , so that the attention value of the mask area is replaced by the maximum value of the current matrix, and the updated attention matrix A is obtained ′ ; Next, calculate A ′ The softmax of A is multiplied by the value vector V and passed through the feedforward network FFN(·) to get the output of the attention layer O = FFN(softmax(A ′ )×V);
[0076] 4.5) The localization perceptron consists of L layers of attention layers, returns to step 4.3), and uses the output O of the previous attention layer as the query vector calculated by the next attention layer Until the output of the last attention layer is obtained, the image features of the final output fused with mask information are recorded as
[0077] 4.6) For text instruction x t , use the pre-trained text encoder CLIP to extract the text feature vector Z t ;
[0078] For robot status information x s , use the multi-layer perceptron MLP projection to get the state feature vector Z s ;
[0079] For action instructions, initialize a learnable action token <act>, and its corresponding action feature vector is Z a ;
[0080] Therefore, the input sequence at the current moment is recorded as
[0081] 4.7) The complete input, i.e., the input sequence containing N historical operation step data, is used to predict the next action instruction using the Transformer decoder, wherein the action instruction includes a 1-degree-of-freedom gripper action and a 7-degree-of-freedom robot arm action.
[0082] In the above step 5, the loss function for jointly training the robot operation strategy network includes the gripper action loss and the robot arm action loss.
[0083] In a specific implementation of the present invention, taking the robot used as an example, the loss function includes the gripper action loss of 1 degree of freedom, which uses the binary cross entropy function to calculate the loss And, the 7-DOF robot motion loss, which uses the L1 function to calculate the loss The total loss is
[0084] Using this loss function, the gradient descent learning method is adopted to train the learnable parameters involved in the model; since the training data contains a variety of operation tasks, joint training on multiple different operation tasks can be achieved, so that the model can complete multiple robot operation tasks at the same time.
[0085] In the above step 6, the pre-trained multimodal large model and the trained robot operation strategy network are combined to complete the given instructions in the actual desktop scenario.
[0086] In an actual desktop scenario, given a text instruction, the robot is required to take the target object A to the specified target area; the robot's initial view image and the given text instruction are input into the pre-trained multimodal large model to generate the target object mask and target area mask;
[0087] Maintain an input sequence consisting of N historical operation step data When the historical data is less than N operation steps, the latest historical operation step data is copied as a supplement; the input sequence corresponding to each operation step data is obtained in the same way as the training stage, and the trained robot operation strategy network generates the next action instruction, and the process is repeated until the robot completes the task of the given text instruction.
[0088] The above method is applied to the following embodiments to demonstrate the technical effects of the present invention, and the specific steps in the embodiments are not repeated here.
[0089] The present invention conducts simulation experiments on the Pick and Place operation task based on the RoboCasa simulator. In order to objectively evaluate the performance of the present invention, the present invention randomly generates 400 test data, uses the success rate of task completion for evaluation, and compares it with the following prior art models:
[0090] Compared with 1.ACT model, it is a Transformer-based policy network, trained by conditional variational autoencoder, and uses action block mechanism for action prediction. Although the prediction efficiency is improved, the block method limits its performance in refined operation tasks and is easily affected by the mode collapse problem, resulting in insufficient action generation diversity.
[0091] Comparison 2. The BC-Transformer model is a multimodal fusion model based on behavioral cloning, which encodes language input through CLIP and combines visual features. Although the fusion ability is strong, the model is highly dependent on high-quality demonstration data, has poor adaptability to unseen scenes, and its generalization is limited in complex tasks.
[0092] Comparison 3.GR-1 model is a GPT-style robot operation model that achieves action prediction by combining visual and language information. Although its design can predict future actions and scenes, the model relies heavily on large-scale pre-training, resulting in high computational overhead and data requirements.
[0093] According to the steps described in the specific implementation manner, the experimental results obtained are shown in Tables 1 and 2.
[0094] Table 1: Success rate of robot operation test in RoboCasa simulation scenario according to the present invention
[0095]
[0096] Table 2: Test success rates of robot grasping and placing operations for novel object scenes of the present invention
[0097]
[0098] As can be seen from Table 1, the success rate of the present invention in robot operation tasks under various types of text instructions is significantly better than that of existing advanced methods. For grasping and placing tasks, the success rate of the present invention method in the difficult task settings of appearance, spatial relationship and common sense reasoning is increased from 13.8 to 30.5 (an increase of 121%), from 16.3 to 33.5 (an increase of 105%), and from 11.5 to 30.0 (an increase of 160%), respectively, showing significant performance improvements. These improvements are due to the fact that the present invention uses a pre-trained multimodal large model to generate segmentation masks for target objects and regions, making full use of the powerful reasoning and positioning capabilities of the multimodal large model, thereby significantly enhancing the overall performance of robot operation.
[0099] As can be seen from Table 2, when the segmentation mask information is not used, the success rate of the present invention in testing new object categories is relatively limited, but when the segmentation mask information is added, the success rate is significantly improved. This shows that the segmentation mask is an effective intermediate representation that can greatly improve its operation performance and task generalization ability for unseen objects, thereby realizing an efficient and generalizable robot operation strategy.
[0100] In this embodiment, a generalizable robot operating system based on segmentation mask representation is also provided, which is used to implement the above embodiment. The terms "module", "unit", etc. used below can implement a combination of software and / or hardware for a predetermined function. Although the system described in the following embodiments is preferably implemented in software, it is also possible to implement hardware, or a combination of software and hardware.
[0101] A generalizable robot operating system based on segmentation mask representation, comprising:
[0102] A training sample generation module is used to construct a three-dimensional object library and a desktop scene library, and randomly generate a large number of desktop scene layouts in a virtual environment based on the data in the library; by giving a target object and a target area, the robot operation trajectory data for the target object in the desktop scene is generated; for each robot operation trajectory data, a variety of text instructions are generated in combination with the appearance, spatial position relationship and common sense knowledge of all objects in the desktop scene; for each robot operation trajectory data, the robot view image, robot state data and a text instruction under each operation step are collected as a training sample to construct a training sample set; the robot state data includes the robot's gripper switch and the joint angle of the robot arm;
[0103] A multimodal large model module, which uses a pre-trained multimodal large model to locate the target object and target area indicated by the text instruction in each training sample, and obtains a target object mask and a target area mask;
[0104] An operation strategy network module, which takes the target object mask, target area mask and training samples corresponding to several historical operation steps as input, extracts the robot view image features, robot state features and text instruction features, fuses the target object mask and target area mask with the robot view image features, and predicts the action instructions of the robot's next operation step based on the fused results and the robot state features, text instruction features and action features corresponding to the learnable action token;
[0105] An operation strategy network training module is used to calculate the loss based on the actual state of the robot at the next operation step and the predicted action instructions to train the robot operation strategy network;
[0106] The robot module is used to use the pre-trained multimodal large model and the trained robot operation strategy network to complete given instructions in actual desktop scenarios.
[0107] For the system embodiment, since it basically corresponds to the method embodiment, the relevant parts can refer to the partial description of the method embodiment, and the implementation methods of the remaining modules will not be repeated here. The system embodiment described above is only schematic, wherein the units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they may be located in one place, or they may be distributed on multiple network units. Some or all of the modules may be selected according to actual needs to achieve the purpose of the scheme of the present invention. Ordinary technicians in this field can understand and implement it without paying creative work.
[0108] The embodiments of the system of the present invention can be applied to any device with data processing capabilities, and the device with data processing capabilities can be a device or apparatus such as a computer. The system embodiments can be implemented by software, or by hardware or a combination of software and hardware. Taking software implementation as an example, as a device in a logical sense, the corresponding computer program instructions in the non-volatile memory are read into the memory by the processor of any device with data processing capabilities and run.
[0109] The above examples are only specific embodiments of the present invention. Obviously, the present invention is not limited to the above examples, and many variations are possible. All variations that can be directly derived or associated with the contents disclosed by a person skilled in the art should be considered as the protection scope of the present invention.< / act> < / seg> < / seg> < / seg> < / act> < / seg> < / seg>
Claims
1. A generalizable robot operation method based on segmentation mask representation, wherein the robot is composed of a gripper and a mechanical arm, characterized in that: The following steps are involved: (1) Build a 3D object library and a desktop scene library, and randomly generate a large number of desktop scene layouts in virtual environments based on the data in the library; (2) Generate robot operation trajectory data for the target object in the desktop scene by giving the target object and the target area; for each robot operation trajectory data, generate diversified text instructions by combining the appearance, spatial position relationship and common sense knowledge of all objects in the desktop scene; (3) For each robot operation trajectory data, collect the robot view image, robot state data and a text instruction at each operation step as a training sample to construct a training sample set; the robot state data includes the robot's gripper switch and the joint angle of the robot arm; (4) Using the pre-trained multimodal large model, the target object and target area indicated by the text instruction in each training sample are located to obtain the target object mask and the target area mask; (5) inputting the target object mask, target area mask, and training samples corresponding to several historical operation steps into the robot operation strategy network, extracting the robot view image features, robot state features, and text instruction features, fusing the target object mask and target area mask with the robot view image features, and predicting the action instructions of the robot's next operation step based on the fusion results and the robot state features, text instruction features, and action features corresponding to the learnable action token, and calculating the loss according to the actual state of the robot at the next operation step and the predicted action instructions to train the robot operation strategy network; (6) Use the pre-trained multimodal large model and the trained robot operation strategy network to complete given instructions in actual desktop scenarios.
2. The generalizable robot operation method based on segmentation mask representation according to claim 1, characterized in that: In step (1), the desktop scene layout in the virtual environment includes a desktop scene, a target object and a plurality of interference objects.
3. The generalizable robot operation method based on segmentation mask representation according to claim 1 or 2, characterized in that: The objects in the three-dimensional object library contain attribute information, including category, color, shape and material.
4. The generalizable robot operation method based on segmentation mask representation according to claim 1, characterized in that: In step (2), each robot operation trajectory data corresponds to a number of text instructions.
5. The generalizable robot operation method based on segmentation mask representation according to claim 4, characterized in that: The text instructions are implemented by a large language model, and the input of the large language model includes the target object, the target area, the robot's initial viewing angle image, all object attribute information in the desktop scene, and related prompt words.
6. The generalizable robot operation method based on segmentation mask representation according to claim 1, characterized in that: The pre-trained multimodal large model described in step (4) includes an image encoder, a multi-layer perceptron, a large language model and a positioning module; Step (4) specifically includes: (4-1) Use the image encoder to obtain the robot's initial view image x in the training sample v,0 Encode features and project them into the embedding feature space of the large language model through a multi-layer perceptron; (4-2) Initializing a positioning module by a pre-trained SAM model, wherein the positioning module includes a pre-trained image encoder and a pre-trained image decoder, and the positioning module is connected after the large language model; (4-3) Add special tokens to the vocabulary of the large language model <seg>, and given a large language model prompt word, the large language model is required to locate the target object and target area referred to in the text instruction according to the input image encoding features and text instructions, and generate text descriptions respectively Where CLIP(·) represents the image encoder, f v (·) represents a multi-layer perceptron, x t Represents a text instruction, represents a large language model, y t,1 Indicates the text description of the target object referred to in the text instruction. t,2 Indicates the text description of locating the target area referred to in the text instruction;< / seg> (4-4) When the large language model text output contains tokens <seg>When the last hidden layer feature vector before the text output is input into the positioning module, the corresponding segmentation mask is obtained. Among them, M o is the target object mask, M p is the target area mask, and both segmentation masks are 0-1 mask matrices with the same size as the robot view image; ε(·) represents the pre-trained image encoder in the localization module, represents the pre-trained image decoder in the localization module, Represents y t,1 Output the corresponding last hidden layer feature vector before, Represents y t,2 Output the corresponding last hidden layer feature vector.< / seg> 7. The generalizable robot operation method based on segmentation mask representation according to claim 1, characterized in that: The robot operation strategy network includes a pre-trained image encoder, a pre-trained text encoder, a multi-layer perceptron, a positioning perceptron, and a Transformer decoder, wherein the positioning perceptron is composed of several attention layers; Step (5) specifically includes: (5-1) For each operation step, the robot view image is imaged using the pre-trained image encoder in the robot operation strategy network to obtain the robot view image features. Includes a global image feature and a local image feature Among them, D v The feature dimension encoded for the image; (5-2) The positioning sensor initializes a global query vector A target object query vector and a target region query vector Where D p is the initial feature dimension; (5-3) In the first attention layer of the localization perceptron, the three vectors are concatenated together and projected into the hidden layer space to obtain the query vector Where d is the hidden layer feature dimension; robot view image feature After being projected by different projection matrices, they are concatenated with the query vector Q to obtain the key vectors Sum value vector Calculate the attention matrix based on the query vector Q and the key vector K (5-4) Mask the target object M o and the target region mask M p They are mapped to a feature map size of 14×14 and converted into a one-dimensional vector of size 1×196, and then M o The corresponding one-dimensional vector is applied to A [1,:196] , M p The corresponding one-dimensional vector is applied to A [2,:196] , so that the attention value of the mask area is replaced by the maximum value of the current matrix, and the updated attention matrix A is obtained ′ ; Next, calculate A ′ The softmax of A is multiplied by the value vector V and passed through the feedforward network FFN(·) to get the output of the attention layer O = FFN(softmax(A ′ )×V); (5-5) Return to step (5-3) and use the output O of the previous attention layer as the query vector calculated by the next attention layer Until the output of the last attention layer is obtained, the image features of the final output fused with mask information are recorded as (5-6) Use the pre-trained text encoder and multi-layer perceptron to extract the robot state feature Z t , text instruction feature Z s , and initialize a learnable action token <act>And extract the action feature Z a ; Get the input sequence corresponding to each operation step < / act> (5-7) The above input sequence corresponding to the N historical operation step data is input into the Transformer decoder to predict the next action instruction, where the action instruction includes the gripper action and the robot arm action.
8. The generalizable robot operation method based on segmentation mask representation according to claim 7, characterized in that: Step (6) specifically includes: (6-1) In an actual desktop scenario, a text instruction is given to require the robot to take the target object to the specified target area; the robot's initial view image and the given text instruction are input into the pre-trained multimodal large model to generate a target object mask and a target area mask; (6-2) Maintain an input sequence consisting of N historical operation step data When the historical data is less than N operation steps, the latest historical operation step is copied as a supplement; the input sequence corresponding to each operation step is obtained in the same way as the training stage, and the next action instruction is generated by the trained robot operation strategy network. The process is repeated until the robot completes the task of the given text instruction.
9. The generalizable robot operation method based on segmentation mask representation according to claim 7, characterized in that: The pre-trained image encoder in the robot operation strategy network adopts the ViTMAE image encoder.
10. A generalizable robot operating system based on segmentation mask representation, used to implement the method of claim 1; characterized in that: The system comprises: A training sample generation module is used to construct a three-dimensional object library and a desktop scene library, and randomly generate a large number of desktop scene layouts in a virtual environment based on the data in the library; by giving a target object and a target area, the robot operation trajectory data for the target object in the desktop scene is generated; for each robot operation trajectory data, a variety of text instructions are generated in combination with the appearance, spatial position relationship and common sense knowledge of all objects in the desktop scene; for each robot operation trajectory data, the robot view image, robot state data and a text instruction under each operation step are collected as a training sample to construct a training sample set; the robot state data includes the robot's gripper switch and the joint angle of the robot arm; A multimodal large model module, which uses a pre-trained multimodal large model to locate the target object and target area indicated by the text instruction in each training sample, and obtains a target object mask and a target area mask; An operation strategy network module, which takes the target object mask, target area mask and training samples corresponding to several historical operation steps as input, extracts the robot view image features, robot state features and text instruction features, fuses the target object mask and target area mask with the robot view image features, and predicts the action instructions of the robot's next operation step based on the fused results and the robot state features, text instruction features and action features corresponding to the learnable action token; An operation strategy network training module is used to calculate the loss based on the actual state of the robot at the next operation step and the predicted action instructions to train the robot operation strategy network; The robot module is used to use the pre-trained multimodal large model and the trained robot operation strategy network to complete given instructions in actual desktop scenarios.
Citation Information
Patent Citations
Robot sequence task learning method based on visual simulation
CN111203878A
Dynamic interactive representation-based dexterous manipulator grabbing method
CN117798919A
Target object grabbing method and system based on multi-source knowledge driving
CN118038221A
Open-vocabulary robotic control using multi-modal language models
WO2024178241A1
Cited By
Method and system for editing abnormal region of railway freight train image
CN120472051A
A method and system for editing abnormal areas in railway freight train images
CN120472051B
Air-ground robot action prediction model training and application method, equipment and medium
CN122045828A