Robot manipulation method based on visual language large model
Patent Information
- Application Number
- CN202410784930.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-06-18
- Publication Date
- 2026-09-29
- Estimated Expiration
- 2044-06-18
AI Technical Summary
[0004]本发明的目的是为了解决现有机器人理解指令及视觉环境后执行的操纵任务完成准确率低的问题,而提出基于视觉语言大模型的机器人操纵方法
[0037]本发明机器人根据操作者所提的语言指令和深度相机获取的视觉环境包括物体颜色、位置等信息,对视觉模态和语言模态进行特征提取,并通过Perceiver Transformer作为算法进行视觉语言多模态特征融合。根据所提取到的特征,完成相关操纵任务。同时,基于先验知识提升模型泛化能力,使机器人即使面对完全未知的任务和物体时,也可以获取操纵物体的基本信息和对应物体部件位置方位,以方便完成操纵任务。
Smart Images

Figure CN118559711B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the fields of artificial intelligence and embodied intelligence, specifically to a method for supporting the manipulation of multi-task robots using large language model algorithms based on multimodal information such as visual language. Background Technology
[0002] With the continuous development of society and economy and the booming rise of artificial intelligence in China, the concept of embodied intelligence has recently attracted widespread attention from researchers in the field of artificial intelligence. Several major global AI companies have recently launched large-scale models, such as OpenAI's GPT-4 and MetaAI's SAM. These large models possess extremely strong semantic understanding capabilities, bringing the previous "Internet AI" to near-peak performance under current computing power conditions. Embodied intelligence aims to enable robots to develop an understanding of the objective world through their own learning after interacting with their environment, thereby achieving true intelligence. In recent years, sub-tasks under embodied intelligence have been continuously developed, such as visual language navigation and visual language question answering. Building on this research, attention has begun to turn to sub-tasks involving manipulating objects through human commands. These sub-tasks are called visual language-based robot manipulation tasks, where the operator provides verbal commands, the robot understands the commands and the visual environment, and completes the relevant manipulation tasks. These tasks can be widely applied in various scenarios such as industrial production and daily life.
[0003] Although robot manipulation tasks based on visual language information are closely related to human production and daily life, research in this field is still in its early stages, and there are many difficulties in completing such tasks. First, understanding abstract instructions is a major challenge; enabling the agent to comprehend abstract instructions and break them down into concrete sub-instructions is a significant hurdle. Second, agents struggle to judge the progress of long-term tasks, often getting stuck at a certain stage and unable to assess progress, thus hindering further progress. Finally, some current algorithms exhibit poor generalization performance, making it difficult to achieve high success rates in new scenarios and with new instructions. Summary of the Invention
[0004] The purpose of this invention is to solve the problem of low accuracy in performing manipulation tasks by existing robots after understanding instructions and the visual environment, and to propose a robot manipulation method based on a large visual language model.
[0005] The specific process of the robot manipulation method based on the large visual language model is as follows:
[0006] Step 1: Input the language instruction text and the RGBD image captured by the depth camera into the visual language large model;
[0007] The PC output of the large visual language model includes three-dimensional position coordinates, three-dimensional rotation pose, and the opening and closing state of the mechanical gripper.
[0008] Step 2: The Jetson Nano end of the large-scale visual language model robotic arm receives 3D position coordinates, 3D rotation pose, and the opening and closing state of the robotic gripper via ROS.
[0009] Step 3: Using the KDL library, the Jetson Nano end of the visual language large model robotic arm performs inverse kinematics calculations on the received 3D position coordinates, 3D rotation pose, and the opening and closing state information of the robotic gripper. The calculated joint angles are then input into the servo motor, and PID control is applied to the servo motor to complete the robotic arm's movements.
[0010] Preferably, the large visual language model includes a PC and a depth camera;
[0011] The PC is connected to the Dofbot six-degree-of-freedom robotic arm and a depth camera, respectively.
[0012] The Dofbot six-degree-of-freedom robotic arm has six servo motors, each with a rotation angle of 0° to 180°;
[0013] The six servos are connected in series, with the output shaft of one bus servo connected to the input shaft of the next servo.
[0014] The Dofbot six-degree-of-freedom robotic arm encapsulates servo control functions. The main control board runs a Python program to input the angles of each servo to control each servo, and the main control board runs a Python program to read the position information of each servo.
[0015] The main control board of the Dofbot six-degree-of-freedom robotic arm is a Jetson Nano development board, and the Dofbot six-degree-of-freedom robotic arm is equipped with multiple USB interfaces to connect to a depth camera.
[0016] The PC communicates with the Dofbot six-degree-of-freedom robotic arm using the ROS system, with the PC acting as the host computer for ROS.
[0017] Preferably, the depth camera is an Astra Pro camera manufactured by Orbbec. Before using the depth camera, the Zhang calibration method is first used to calibrate the depth camera to obtain the relationship between the depth camera coordinate system and the world coordinate system.
[0018] Preferably, the PC is a PC running Ubuntu 18.04.
[0019] Preferably, in step one, the language instruction text and the RGBD image captured by the depth camera are input into the visual language large model;
[0020] The PC output of the large visual language model includes three-dimensional position coordinates, three-dimensional rotation pose, and the opening and closing state of the mechanical gripper.
[0021] The specific process is as follows:
[0022] Step 11: Obtain additional specific component information;
[0023] Steps 1 and 2: Obtain the location information of specific components;
[0024] Step 13: Input the language instruction text and additional specific component information into the CLIP language encoder to obtain the language feature vector;
[0025] Voxel block features are extracted from RGBD images and the location information of specific components using a 3D convolutional visual encoder.
[0026] The obtained language feature vectors and voxel block features are input into the Perceiver Transformer algorithm model for feature fusion to obtain the fusion result.
[0027] The fusion results are input into three multilayer perceptrons, which respectively obtain three-dimensional position coordinate information, rotational pose information, and opening and closing state information of the manipulator.
[0028] The three-dimensional position coordinate information, rotational pose information, and opening and closing state information of the robotic gripper are input into the three-dimensional voxel encoder for encoding, and the encoded three-dimensional position coordinate information of the robotic arm end, the three-dimensional rotational pose information of the robotic arm end, and the opening and closing state information of the robotic gripper at the robotic arm end are output.
[0029] The encoded 3D position coordinates of the robotic arm end effector, the 3D rotation pose information of the robotic arm end effector, and the opening and closing state information of the robotic arm end effector claw, output by the PC of the large language model, are transmitted to the Jetson Nano terminal on the large visual language model robotic arm via ROS communication, so as to control the robotic arm to complete the corresponding actions.
[0030] Preferably, additional specific component information is obtained in each of the steps; the specific process is as follows:
[0031] The language instruction text is input into the visual language large model. The NLKT library is used to extract object-related nouns from the language instruction text. The object-related nouns are then input into the pre-trained language model GPT-4 in the form of a template. The object part information output by the pre-trained language model GPT-4 is used as additional specific part information.
[0032] Preferably, the location information of a specific component is obtained in steps one and two; the specific process is as follows:
[0033] The SAM pre-trained semantic segmentation model cuts the RGBD image into N parts, and inputs the information of the N parts into the CLIP visual encoder one by one. The CLIP visual encoder outputs feature vectors.
[0034] Simultaneously, the first piece of information from the additional specific component information is input into the CLIP language encoder, which outputs a feature vector; N takes the value of a positive integer.
[0035] Calculate the similarity between the feature vectors output by the CLIP visual encoder and the feature vectors output by the CLIP language encoder, and take the region with the highest similarity as the location information of the specific component.
[0036] The beneficial effects of this invention are as follows:
[0037] This invention's robot extracts features from both visual and linguistic modalities based on verbal commands from the operator and visual environment information, including object color and position, obtained from a depth camera. It then uses a Perceiver Transformer algorithm to fuse these visual and linguistic multimodal features. Based on the extracted features, the robot completes relevant manipulation tasks. Simultaneously, prior knowledge enhances the model's generalization ability, enabling the robot to acquire basic information about the object being manipulated and the position and orientation of its components, even when faced with completely unknown tasks and objects, thus facilitating the completion of manipulation tasks. Attached Figure Description
[0038] Figure 1 This is a schematic diagram of a robot manipulation method based on a large visual language model.
[0039] Figure 2 Diagram of the control mechanism for a dual-arm robot;
[0040] Figure 3 A flowchart for a large-scale model of a robot visual language for unknown scenarios;
[0041] Figure 4 A flowchart for obtaining additional specific component information;
[0042] Figure 5 A flowchart for obtaining the location information of specific components;
[0043] Figure 6 This is a flowchart of robot manipulation based on a large visual language model. Detailed Implementation
[0044] Specific Implementation Method 1: The specific process of the robot manipulation method based on the large visual language model in this implementation method is as follows:
[0045] Step 1: Input multimodal information, such as language instruction text and RGBD images captured by depth cameras, into the visual language large model;
[0046] The PC output of the large visual language model includes three-dimensional position coordinates, three-dimensional rotation pose, and the opening and closing state of the mechanical gripper.
[0047] Three-dimensional position coordinates and three-dimensional rotation pose are used to describe the six-degree-of-freedom pose of a large model of visual language;
[0048] The opening and closing state of the mechanical gripper controls the grasping and placement actions of the large model using visual language.
[0049] Taking "open the drawer" as an example, the input language command text is:<Open top drawer> ;
[0050] Step 2: The Jetson Nano end of the large-scale visual language model robotic arm receives 3D position coordinates, 3D rotation pose, and the opening and closing state of the robotic gripper via ROS.
[0051] Step 3: Using the KDL library, the Jetson Nano end of the visual language large model robotic arm performs inverse kinematics calculations on the received 3D position coordinates, 3D rotation pose, and the opening and closing state information of the robotic gripper. The calculated joint angles are then input into the servo motor, and PID control is applied to the servo motor to complete the next robotic arm action corresponding to the command observed in the current visual scene.
[0052] Specific Implementation Method Two: This implementation method differs from Specific Implementation Method One in that the large visual language model includes a PC and a depth camera;
[0053] The PC is connected to the Dofbot six-degree-of-freedom robotic arm and a depth camera, respectively.
[0054] The Dofbot six-degree-of-freedom robotic arm has six servo motors, each with a rotation angle of 0° to 180° (greater than or equal to 0° and less than or equal to 180°).
[0055] The six servos are connected in series, with the output shaft of one bus servo connected to the input shaft of the next servo.
[0056] The Dofbot six-degree-of-freedom robotic arm encapsulates servo control functions. The main control board runs a Python program to input the angles of each servo to control each servo. The main control board also runs a Python program to read the relevant position information of each servo.
[0057] The main control board of the Dofbot six-degree-of-freedom robotic arm is a Jetson Nano development board, and the Dofbot six-degree-of-freedom robotic arm is equipped with multiple USB ports to connect to external devices such as depth cameras;
[0058] The PC communicates with the Dofbot six-degree-of-freedom robotic arm using the ROS system, with the PC acting as the host computer for ROS.
[0059] Visual language large model manipulation device such as Figure 2 As shown, the dual arms utilize the Dofbot six-DOF robotic arm manufactured by Yabo Intelligent Technology Co., Ltd., featuring six bus servos with rotation angles ranging from 0 to 180°. The servos are cascaded in series, with the output shaft of one servo connected to the input shaft of the next. The Dofbot robot incorporates built-in servo control functions, allowing the main control board to run Python programs to input servo angles for control and to read servo position information. The main control board on the robotic arm uses a Jetson Nano development board with multiple USB ports for connecting external devices such as cameras. The depth camera is an Astra Pro camera manufactured by Orbbec. Before use, the RGBD camera is calibrated using Zhang's calibration method to correct camera distortion and obtain the relationship between the camera coordinate system and the world coordinate system. Integrated in the center is a PC running Ubuntu 18.04. The PC communicates with the robotic arm via the ROS system, with the PC acting as the host computer for ROS.
[0060] The other steps and parameters are the same as in Specific Implementation Method 1.
[0061] Specific Implementation Method 3: This implementation method differs from Specific Implementation Method 1 or 2 in that the depth camera used is the Astra Pro camera manufactured by Orbbec. Before using the depth camera, the Zhang's calibration method is first used to calibrate the depth RGBD camera in order to correct camera distortion and obtain the relationship between the depth camera coordinate system and the world coordinate system.
[0062] Other steps and parameters are the same as in specific implementation method one or two.
[0063] Specific Implementation Method Four: This implementation method differs from one of the specific implementation methods one to three in that the PC is a PC with the Ubuntu 18.04 system.
[0064] The other steps and parameters are the same as those in one of the specific implementation methods one to three.
[0065] Specific Implementation Method 5: This implementation method differs from one of the specific implementation methods one to four in that, in step one, multimodal information such as language instruction text and RGBD images captured by a depth camera are input into the visual language large model;
[0066] The PC output of the large visual language model includes three-dimensional position coordinates, three-dimensional rotation pose, and the opening and closing state of the mechanical gripper.
[0067] Three-dimensional position coordinates and three-dimensional rotation pose are used to describe the six-degree-of-freedom pose of a large model of visual language;
[0068] The opening and closing state of the mechanical gripper controls the grasping and placement actions of the large model using visual language.
[0069] The specific process is as follows:
[0070] Step 11: Obtain additional specific component information;
[0071] The process for obtaining additional specific component information is as follows: Figure 4 As shown;
[0072] Steps 1 and 2: Obtain the location information of specific components;
[0073] Step 13: Input the language instruction text and additional specific component information into the CLIP language encoder to obtain the language feature vector;
[0074] Voxel block features are extracted from RGBD images and the location information of specific components using a 3D convolutional visual encoder.
[0075] The obtained language feature vectors and voxel block features are input into the PerceiverTransformer algorithm model in the order of extraction to perform feature fusion and obtain the fusion result.
[0076] The fusion results are input into three multilayer perceptrons, which respectively obtain three-dimensional position coordinate information, rotational pose information (rotational pose information of the end effector along the XYZ axes), and opening and closing state information of the manipulator.
[0077] The three-dimensional position coordinate information, rotational pose information, and opening and closing state information of the robotic gripper are input into the three-dimensional voxel encoder for encoding, and the encoded three-dimensional position coordinate information of the robotic arm end, the three-dimensional rotational pose information of the robotic arm end, and the opening and closing state information of the robotic gripper at the robotic arm end are output.
[0078] The encoded 3D position coordinates of the robotic arm end effector, the 3D rotation pose information of the robotic arm end effector, and the opening and closing state information of the robotic arm end effector claw, output by the PC of the large language model, are transmitted to the Jetson Nano terminal on the large visual language model robotic arm via ROS communication, so as to control the robotic arm to complete the corresponding actions.
[0079] The other steps and parameters are the same as those in one of the specific implementation methods one to four.
[0080] Specific Implementation Method Six: This implementation method differs from Specific Implementation Methods One to Five in that additional specific component information is obtained in step one; the specific process is as follows:
[0081] The language instruction text is input into the visual language large model. The NLKT library is used to extract object-related nouns from the language instruction text. The object-related nouns are then input into the pre-trained language model GPT-4 in the form of a template. The object part information output by the pre-trained language model GPT-4 is used as additional specific part information.
[0082] Object-related terms refer to the object names extracted from the NLKT library in language instructions;
[0083] The NLKT library is an open-source library in Python that breaks down an instruction into various word classes;
[0084] GPT-4: You are a robot, and you are asked to finish the task requested by people with natural language instructions. Now you know what the language instruction is and what are the object related to the task, you need to identify what part is most relevant to the given language instruction.
[0085] Example 1: Instruction:Open the door
[0086] Object:door
[0087] You need to answer:['handle','knob','doorframe','lock','hinge','window','pane']
[0088] Example 2: Instructions:screw open the wine bottle
[0089] Object: bottle
[0090] You need to answer:['cap','base','label','neck']
[0091] Example 3: Instructions:<INSTRUCTION_TO_PROCESS>
[0092] Object:<GIVEN_BY_NLTK>
[0093] You need to answer:
[0094] The templates are Instructions and Object;
[0095] Instructions are the given language instructions, and Object is extracted from the NLKT library. Instructions and Object are used as input to the language model GPT-4, and the language model GPT-4 outputs "You need to answer". After training, a pre-trained language model GPT-4 is obtained; "You need to answer" is additional specific part information.
[0096] The process for obtaining additional specific component information is as follows: Figure 4 As shown;
[0097] The other steps and parameters are the same as those in one of the specific implementation methods one to five.
[0098] Specific Implementation Method Seven: This implementation method differs from Specific Implementation Methods One through Six in that the location information of the specific component is obtained in steps one and two; the specific process is as follows:
[0099] The SAM pre-trained semantic segmentation model cuts the RGBD image into N parts, and inputs the information of the N parts into the CLIP visual encoder one by one. The CLIP visual encoder outputs feature vectors.
[0100] Simultaneously, the first piece of information from the additional specific component information (GPT infers a series of nouns that it thinks may be related to the instructions based on the language instructions, and the first of these nouns is the most relevant) is input into the CLIP language encoder, and the CLIP language encoder outputs a feature vector; N takes the value of a positive integer (the number of objects in the RGBD image determines the number of parts to be segmented; after setting some precision parameters for segmentation, the input image can be automatically segmented by the SAM pre-trained semantic segmentation model);
[0101] Calculate the similarity between the feature vector output by the CLIP visual encoder and the feature vector output by the CLIP language encoder, and take the region with the highest similarity as the location information of the specific component;
[0102] The process for obtaining the location information of specific components is as follows: Figure 5 As shown;
[0103] Taking "open the top drawer" as an example, follow the entered language command.<Open the top drawer> The model infers nouns such as "handle," "knob," "track," and "lock" from a GPT-4 pre-trained language model. The first noun, "handle," serves as additional guidance for "focusing on handle," and is compared with the output of the language visual encoder, selecting the position with the highest similarity as the component location of "handle." During image input, a one-dimensional dimension is added to the original RGB 3D channels to represent the component location. The fourth dimension of the voxel unit at the "handle" location is set to 1, while other units are set to 0, thus introducing visual component location matching information. Through these two types of information, the model can significantly reduce the operational range required for manipulation tasks when facing completely unknown objects, thereby greatly improving the success rate of manipulation tasks.
[0104] The other steps and parameters are the same as those in one of the specific implementation methods one to six.
[0105] This invention may have other embodiments. Without departing from the spirit and essence of this invention, those skilled in the art can make various corresponding changes and modifications according to this invention, but these corresponding changes and modifications should all fall within the protection scope of the appended claims.
Claims
1. A robot manipulation method based on a large visual language model, characterized by: The specific process of the method is as follows: Step 1: Input the language instruction text and the RGBD image captured by the depth camera into the visual language large model; The PC output of the large visual language model includes three-dimensional position coordinates, three-dimensional rotation pose, and the opening and closing state of the mechanical gripper. Step 2: The Jetson Nano end of the large-scale visual language model robotic arm receives 3D position coordinates, 3D rotation pose, and the opening and closing state of the robotic gripper via ROS. Step 3: Using the KDL library, the Jetson Nano end of the visual language large model robotic arm performs inverse kinematics calculations on the received 3D position coordinates, 3D rotation pose, and the opening and closing state of the robotic gripper. The calculated joint angles are then input into the servo motor, and PID control is applied to the servo motor to complete the robotic arm's movements. The large-scale visual language model includes a PC and a depth camera; The PC is connected to the Dofbot six-degree-of-freedom robotic arm and a depth camera, respectively. The Dofbot six-degree-of-freedom robotic arm has six servo motors, each with a rotation angle of 0˚ to 180˚; The six servos are connected in series, with the output shaft of one bus servo connected to the input shaft of the next servo. The Dofbot six-degree-of-freedom robotic arm encapsulates servo control functions. The main control board runs a Python program to input the angles of each servo to control each servo, and the main control board runs a Python program to read the position information of each servo. The main control board of the Dofbot six-degree-of-freedom robotic arm is a Jetson Nano development board, and the Dofbot six-degree-of-freedom robotic arm is equipped with multiple USB interfaces to connect to a depth camera. The PC communicates with the Dofbot six-degree-of-freedom robotic arm using the ROS system, with the PC acting as the host computer for ROS.
2. The robot manipulation method based on a large visual language model according to claim 1, characterized in that: The depth camera used is the Astra Pro camera manufactured by Orbbec. Before using the depth camera, the Zhang calibration method is first used to calibrate the depth camera to obtain the relationship between the depth camera coordinate system and the world coordinate system.
3. The robot manipulation method based on a large visual language model according to claim 2, characterized in that: The PC is a PC running Ubuntu 18.
04.
4. The robot manipulation method based on a large visual language model according to claim 3, characterized in that: In step one, the language instruction text and the RGBD image captured by the depth camera are input into the visual language large model; The PC output of the large visual language model includes three-dimensional position coordinates, three-dimensional rotation pose, and the opening and closing state of the mechanical gripper. The specific process is as follows: Step 11: Obtain additional specific component information; Steps 1 and 2: Obtain the location information of specific components; Step 13: Input the language instruction text and additional specific component information into the CLIP language encoder to obtain the language feature vector; Voxel block features are extracted from RGBD images and the location information of specific components using a 3D convolutional visual encoder. The obtained language feature vectors and voxel block features are input into the Perceiver Transformer algorithm model for feature fusion to obtain the fusion result. The fusion results are input into three multilayer perceptrons, which respectively obtain three-dimensional position coordinate information, rotational pose information, and opening and closing state information of the manipulator. The three-dimensional position coordinate information, rotational pose information, and opening and closing state information of the robotic gripper are input into the three-dimensional voxel encoder for encoding, and the encoded three-dimensional position coordinate information of the robotic arm end, the three-dimensional rotational pose information of the robotic arm end, and the opening and closing state information of the robotic gripper at the robotic arm end are output. The encoded 3D position coordinates of the robotic arm end effector, the 3D rotation pose information of the robotic arm end effector, and the opening and closing state information of the robotic arm end effector claw, output by the PC of the visual language large model, are transmitted to the Jetson Nano terminal on the visual language large model robotic arm via ROS communication, so as to control the robotic arm to complete the corresponding actions.
5. The robot manipulation method based on a large visual language model according to claim 4, characterized in that: The steps described above obtain additional specific component information; the specific process is as follows: The language instruction text is input into the visual language large model. The NLKT library is used to extract object-related nouns from the language instruction text. The object-related nouns are then input into the pre-trained language model GPT-4 in the form of a template. The object part information output by the pre-trained language model GPT-4 is used as additional specific part information.
6. The robot manipulation method based on a large visual language model according to claim 5, characterized in that: The location information of the specific component is obtained in steps one and two; The specific process is as follows: The SAM pre-trained semantic segmentation model cuts the RGBD image into N parts, and inputs the information of the N parts into the CLIP visual encoder one by one. The CLIP visual encoder outputs feature vectors. Simultaneously, the first piece of information from the additional specific component information is input into the CLIP language encoder, which outputs a feature vector; N takes the value of a positive integer. Calculate the similarity between the feature vectors output by the CLIP visual encoder and the feature vectors output by the CLIP language encoder, and take the region with the highest similarity as the location information of the specific component.
Citation Information
Patent Citations
Robot instruction operation method and system based on natural language and medium
CN116690616A
Mechanical arm grabbing method driven by natural language
CN117773920A