Mechanical arm autonomous grabbing system based on visual language model and control method thereof

Through the robotic arm autonomous grasping system based on visual language model, the problem of object recognition and grasping in complex environments is solved, and the accurate recognition and grasping of natural language designated objects is achieved, which is suitable for a variety of scenarios.

CN119952694APending Publication Date: 2025-05-09HANGZHOU DIANZI UNIV
View PDF 0 Cites 8 Cited by

Patent Information

Application Number
CN202510022599.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-07
Publication Date
2025-05-09

AI Technical Summary

Technical Problem

The existing robotic arm grasping technology lacks reasoning capabilities and human-machine interaction capabilities in complex environments, making it difficult to effectively identify and grasp objects in complex scenarios such as shelves.

Method used

A robotic arm autonomous grasping system based on visual language model is adopted. This system realizes the recognition and capture of natural language designated objects through image acquisition, visual language positioning, high-level object selection strategy, low-level action selection strategy and robotic arm motion control.

Benefits of technology

It realizes accurate identification and capture of human natural language designated objects in complex environments, and is suitable for various scenarios such as warehouses and pharmacies, improving the inference and operation capabilities of robotic arms in complex environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119952694A_ABST
    Figure CN119952694A_ABST
Patent Text Reader

Abstract

The invention discloses a mechanical arm autonomous grabbing system based on a visual language model and a control method of the mechanical arm autonomous grabbing system. The method comprises the steps that firstly, a mechanical arm is initialized, and a color image and a depth image of a working interval are collected; secondly, based on the color image and the natural language description, generating a target object bounding box and confidence, and judging whether a target object exists in the scene or not; and when the target object does not exist, instance segmentation and capture feasibility detection are carried out on the visible object in the current scene. And then the depth image and the object binary mask passing through grabbing feasibility detection are sent to an interaction object selection strategy and an action selection strategy to obtain an object to be interacted and a pushing action direction. And finally, center coordinates of the to-be-interacted object and center coordinates of the mechanical arm are obtained, the obstacle is moved in the pushing action direction until the target object appears and can be grabbed, and the mechanical arm grabs the target object and places the target object at a designated position. According to the method, the object specified by the human natural language can be accurately recognized and grabbed.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of robot arm grasping control, and in particular to a robot arm autonomous grasping system based on a visual language model and a control method thereof. Background Art

[0002] Robotic grasping technology has attracted much attention in fields such as smart homes and service robots. However, most existing robotic grasping technologies for shelf scenarios are only suitable for grasping objects in simple scenarios, and lack reasoning capabilities and human-computer interaction capabilities in complex environments. This is because traditional robotic grasping systems often rely on images or sensor information to infer the characteristics of objects and make grasping decisions. Systems based on visual language models can use the combination of images and text to enhance perception capabilities. For example, object description: using natural language descriptions to help the system understand the semantic information of objects. For example, "green round cup" provides more background knowledge than simple visual information (color and shape), which helps to improve the understanding of objects. In complex scenes such as shelves in reality, it is of great significance to guide robots to reason and perform various target-level grasping tasks through natural language, such as taking out objects that are blocked deep in the shelf and sorting out messy objects. Summary of the invention

[0003] In view of the deficiencies of the prior art, the present invention proposes a method and system for autonomous grasping of a robot arm based on a visual language model. The system completes the pushing and grasping operations of the corresponding object through image acquisition, visual language positioning, high-level object selection strategy, low-level action selection strategy and robot arm motion control. It can realize the recognition and grasping of objects specified by human natural language, and can be applied to various scenes such as warehouses and pharmacies. In order to solve the above technical problems, the technical solution of the present invention is as follows: The present invention provides an autonomous exploration method of a robot arm based on a visual language model. The method introduces a visual language model to determine whether the target object is visible, determines whether it is graspable through grasping detection, and explores the environment through a two-layer reinforcement learning algorithm when the target object is invisible or uncatchable, finds the target object, and grasps the target object.

[0004] A robotic arm autonomous grasping system based on a visual language model comprises an image data acquisition module, a visual language positioning module, a grasping feasibility judgment module, an interactive object selection module, an action selection module and a robotic arm motion module.

[0005] The image data acquisition module is used to collect the color image I of the working area. rgb and depth image I depth , and the depth map is normalized at the same time.

[0006] The visual language localization module transforms the scene color image Irgb and natural language description L as input, and use the visual language model to generate a bounding box and its confidence around the target object corresponding to the natural language description to determine whether the target object is visible and determine its position if it is visible.

[0007] The grasping feasibility judgment module judges whether the target object is graspable based on the color image data and the target object bounding box. If the target object is invisible or ungraspable, the object in the current scene is instance segmented for input to the subsequent reinforcement learning algorithm.

[0008] The interactive object selection module is used to input the normalized scene depth image and the scene instance segmentation mask of the depth image, and select the interactive object using the reinforcement learning algorithm.

[0009] The action selection module is used to input the local depth image and mask of the selected interactive object and use the reinforcement learning algorithm to determine the action to be performed by the robotic arm.

[0010] The robot arm motion module is used to receive the output of the reinforcement learning algorithm and plan the motion trajectory. The robot arm moves according to the motion trajectory to grab or push objects.

[0011] Preferably, the image data acquisition module further includes preprocessing of the acquired color images and depth images.

[0012] Preferably, the robot arm motion module includes a UR10 six-degree-of-freedom robot arm, an AG-95 two-finger gripper and a terminal force sensor, wherein the AG-95 two-finger gripper is connected to the end of the robot arm, and the force sensor is located at the end of the gripper. It is mainly used to receive the target position output by the algorithm and plan the robot arm motion trajectory, while cooperating with the AG-95 two-finger gripper to grasp or move the target object. The terminal force sensor ensures that the object will not be damaged due to excessive force applied, nor will it slip due to insufficient force.

[0013] The present invention also provides a control method for a robotic arm autonomous grasping system based on a visual language model, comprising the following steps:

[0014] Step 1: Initialize the robot arm and move it to the starting position; collect the color image I of the working area rgb and depth image I depth And pre-process and save.

[0015] Step 2: Transform the color image I rgbThe natural language description L is used as input, and the referring transformer network (a deep learning model for solving the problem of reference resolution) is used to perform the visual grounding task and generate the corresponding target object bounding box and target object confidence. If the confidence is higher than the set threshold, it is considered to be the target object. If the confidence is lower than the set threshold, it is considered not to be the target object, that is, there is no target object in the scene.

[0016] Step 3: When there is no target object, perform instance segmentation on all visible objects in the current scene based on the color image data, and perform grasping feasibility detection on the segmented instance masks.

[0017] Step 4: Send the scene depth image and the binary mask of the object detected by grasping feasibility into the interactive object selection strategy and action selection strategy to obtain the object to be interacted with and the direction of the pushing action.

[0018] Step 5: Obtain the center coordinate information of the object to be interacted with through camera coordinate transformation and hand-eye calibration according to the coordinates of the object's bounding box. The robotic arm performs the operation of moving the obstacle according to the received center coordinate information of the object to be interacted with and the pushing action direction.

[0019] Step 6: Repeat steps 1 to 5 until the target object appears and can be grasped. The robot arm performs a grasping action to grasp the target object and place it at the specified location.

[0020] Preferably, in step 2, the Referring Transformer network uses a vision transformer (VisionTransformer) and a pre-trained language model (BERT) as the main framework. The method for generating the bounding box of the target object corresponding to the natural language is:

[0021] Step 2.1: Input the color image I of the current scene containing multiple objects rgb , input the natural language describing the target object in the scene.

[0022] Step 2.2: The color image is encoded through the Vision Transformer network to extract image features; the natural language description is passed to a pre-trained language model (BERT). The natural language description will be segmented and embedded into a fixed-dimensional vector space to obtain text features.

[0023] Step 2.3: In the Referring Transformer, the image features and text features are associated with different regions of the image and the specific descriptions in the text through the self-attention mechanism in the transformer architecture, thereby achieving the fusion of image and natural language.

[0024] Step 2.4, Referring Transformer generates a region proposal through the fused visual-linguistic features. These regions are areas that may contain target objects, and each region has a corresponding confidence. The region with the highest confidence is selected. If the confidence of the region is higher than the set threshold, it is considered to be the target object. If the confidence is lower than the set threshold, it is considered not to be the target object, that is, there is no target object in the scene.

[0025] Step 2.5, Referring Transformer uses the convolutional layer to generate the bounding box coordinates of the target object.

[0026] Preferably, the grasping feasibility detection in step 3 is performed by checking the integrity of the instance mask and whether the blank area around it supports the gripper to extend into the instance mask.

[0027] Preferably, the method for performing instance segmentation based on depth image data in step 3 adopts Mask R-CNN, and the specific steps are as follows:

[0028] Step 3.1: Input a color image of the current scene containing multiple objects, usually of size H×W. The input image is first processed by a convolutional neural network (ResNet) to obtain a feature map FM (Feature Map). This process extracts high-level semantic features in the image, and the size of the feature map is usually H′×W′×C, where H′ and W′ are the spatial dimensions of the feature map, and C is the number of channels.

[0029] Step 3.2: Based on the feature map FM, use Region Proposal Network (RPN) to generate candidate object regions. RPN scans the feature map by sliding a window and outputs multiple bounding boxes and their scores.

[0030] Step 3.3: Use the RoIAlign operation to perform precise spatial alignment on each candidate object region (RoI) to obtain a fixed-size feature map.

[0031] Step 3.4: Use the fully connected layer to classify each RoI and regress its bounding box. At the same time, Mask R-CNN generates a mask branch on each RoI. This mask branch uses convolution operations to generate a binary mask of the object instance.

[0032] Step 3.5, Mask R-CNN outputs the mask of each detected object.

[0033] Preferably, both strategies in step 4 adopt the Double DQN network, wherein the interactive object selection strategy is described as follows:

[0034] When the target object is not visible or is severely occluded, the exploration reward r e Defined by:

[0035]

[0036] Where c represents the selected object, p e represents the area that needs to be explored. This formula shows that the strategy encourages the robot arm to select objects in the unexplored area for processing. As the exploration progresses, the explored area becomes larger and larger, and the probability of the strategy finding the target object becomes greater and greater.

[0037] When the target object is partially occluded but can be directly located by the visual language model, the target reward r t Defined by:

[0038]

[0039] where x t Represents the central horizontal coordinate of the target object instance segmentation mask, x o Represents the central horizontal coordinate of the instance segmentation mask of the selected operation object, x j represents the central horizontal coordinate of the instance segmentation mask of each graspable object in the scene. This formula indicates that the selection reward r t This will prompt the intelligent user to select the object closest to the target for interaction.

[0040] The action selection strategy is described as follows:

[0041] Based on the characteristics of the shelf scene, in order to simplify the grasping process, the pushing action is limited to two actions, left and right, and the moving distance is fixed at 10 cm each time. When the target object is invisible or severely obscured, the object selected based on the strategy should be moved to reduce the unexplored area. The unexplored area refers to the area where the target object may appear due to being obscured by the surface objects. To facilitate the calculation of rewards, the size of the unexplored area is quantified and defined by the vector of the horizontal axis. Since the unexplored area appears behind the surface objects of the current scene, the unexplored area can be represented by the sum of the one-dimensional distribution of the current scene instance mask on the X-axis of the image, and finally represented as a one-dimensional binary vector v s Since the unexplored area has the property of only decreasing but not increasing, the v before and after the move s 、v s 'The vector obtained by AND is used as the quantitative representation of the final unexplored area. Therefore, the exploration reward r e ' is defined by the following formula:

[0042]

[0043] Where X represents the length of the vector, which is also the length of the X-axis of the image, i represents the index of the vector element, and · represents the AND operation. This formula shows that the low-level action selection strategy encourages the robot arm to select actions that can reduce the unexplored area.

[0044] When the target object is partially visible but cannot be grasped directly due to the presence of other objects, the space around it needs to be expanded to facilitate grasping. The target reward is r t ' is defined by the following formula:

[0045]

[0046] where v t The vector representing the visible part of the instance segmentation mask of the target object, which indicates that the target reward r t 'Encourage the robot arm to expand the space around the target object so that the robot can grasp it.

[0047] Preferably, in step 6, the movement of the object by the robot arm is performed by grasping movement, and the grasping point is based on the center coordinates of the object. The grasping of the target object by the robot arm is based on the center coordinates of the target object.

[0048] Preferably, the center coordinate information of the object to be interacted with in step 6 is calculated as follows:

[0049] The plane pixel coordinates of the object center point are expressed as (x center ,y center ), using the camera internal parameters (including focal length f x f y and light center c x c y ) to calculate the plane coordinates of the object in the camera coordinate system.

[0050]

[0051] In this way, the plane coordinates p of the object in the camera coordinate system are obtained. camera =[x camera ,y camera ]. Next, the coordinates are transformed from the camera coordinate system to the robot end coordinate system (hand-eye calibration), and the transformation matrix T camera_to_arm It can be expressed as:

[0052]

[0053] Where R camera_to_arm is the rotation matrix, T' camera_to_arm is the translation vector. The conversion process is:

[0054] P arm =R camera_to_arm ·P camera +T camera_to_arm

[0055] This will transform the plane coordinates in the camera coordinate system into the plane coordinates P in the robot end coordinate system. arm .

[0056] Finally, the plane coordinate P in the robot end coordinate system arm Convert to world coordinate system,

[0057] The transformation matrix between the robot end coordinate system and the world coordinate system can be expressed as:

[0058]

[0059] R arm_to_word is the rotation matrix, T' arm_to_word is the translation vector, and the conversion process is:

[0060] P world =R arm_to_world ·P arm_to_world +T arm_to_world

[0061] In this way, we get the plane coordinates P of the center point of the object in the world coordinate system. world ,Then, through the algorithm of the MoveIt motion library in ROS, the motion information of the robot arm in the joint space is obtained and planned.

[0062] The present invention has the following characteristics and beneficial effects:

[0063] The above technical solution is adopted to use the visual language model to parse the natural language commands to realize the positioning of the target object, and use deep reinforcement learning to explore the scene when the target object is invisible or partially visible to find the target object, and finally use the robot arm path planning to complete these actions and realize the grasping of the target object. The system adopts modular design, with clear structural hierarchy and important theoretical value and practical significance. The present invention can accurately recognize and grasp objects specified by human natural language, and can be applied to warehouses, pharmacies and other scenes. BRIEF DESCRIPTION OF THE DRAWINGS

[0064] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative labor.

[0065] Figure 1 The figure is a schematic diagram of a method flow of an embodiment of the present invention.

[0066] Figure 2(a) is the scene depth map at the beginning;

[0067] Figure 2(b) shows the bounding box of the target object obtained after the visual language model;

[0068] Figure 2(c) shows the mapping of the obtained bounding box onto the updated scene depth map;

[0069] Figure 2(d) shows that the bounding box confidence is higher than the threshold and the object is identified as the target object;

[0070] Figure 2(e) shows the target object being captured successfully. DETAILED DESCRIPTION

[0071] It should be noted that, in the absence of conflict, the embodiments of the present invention and the features in the embodiments may be combined with each other.

[0072] In order to make the objectives, technical solutions and advantages of the present invention more clearly understood, the present invention is further described in detail below with reference to the accompanying drawings.

[0073] On the contrary, the present invention covers any substitution, modification, equivalent method and scheme made on the essence and scope of the present invention as defined by the claims. Further, in order to make the public have a better understanding of the present invention, some specific details are described in detail in the detailed description of the present invention below. Those skilled in the art can fully understand the present invention without the description of these details.

[0074] Example 1

[0075] This embodiment provides a robotic arm autonomous grasping system based on a visual language model, including the following modules:

[0076] The image data acquisition module, Realsense D435 camera is used to collect color images of the working area. rgb and depth image I depth The resolution of the collected data is 640×480. At the same time, the depth map is normalized and its value range is scaled to between 0 and 255. This module also includes preprocessing of the collected image data.

[0077] The visual language localization module transforms the scene color image I rgb and natural language description L as input, and use the visual language model to generate a bounding box and its confidence around the target object corresponding to the natural language description to determine whether the target object is visible and determine its position.

[0078] The grasping feasibility judgment module judges whether the target object is graspable based on the color image data and the target object bounding box. If the target object is invisible or ungraspable, the object in the current scene is instance segmented for input to the subsequent reinforcement learning algorithm.

[0079] An interactive object selection module, which is used to input a normalized scene depth image and a scene instance segmentation mask of the depth image, and select an object to be moved using a reinforcement learning algorithm;

[0080] The action selection module is used to input the local depth image and mask of the selected interactive object and use the reinforcement learning algorithm to determine the action to be performed by the robot arm;

[0081] The robot arm motion module is used to receive the output of the reinforcement learning algorithm and plan the motion trajectory. The robot arm moves according to the motion trajectory to grab or push objects.

[0082] The robot arm motion module includes the UR10 six-degree-of-freedom robot arm, the AG-95 two-finger gripper and the end force sensor. The AG-95 two-finger gripper is connected to the end of the robot arm, and the force sensor is located at the end of the gripper. It is mainly used to receive the target position output by the algorithm and plan the robot arm motion trajectory, while cooperating with the two-finger gripper to grab or move the target object. The end force sensor ensures that the object will not be damaged due to excessive force, nor will it slip due to insufficient force.

[0083] Example 2

[0084] This embodiment provides a control method for a robotic arm autonomous grasping system based on a visual language model. Figure 1 As shown, the following steps are included:

[0085] Step 1: Initialize the robot arm and move it to the starting position; the image data acquisition module collects the color image I of the working area. rgb and depth image I depth And pre-process and save.

[0086] Step 2: Transform the color image I rgbThe natural language description L is used as input, and the referring transformer is used to perform the visual grounding task, and the language is generated to generate the corresponding target object bounding box and the target object confidence. If the confidence is higher than the set threshold, it is considered to be the target object. If the confidence is lower than the set threshold, it is considered not to be the target object, that is, there is no target object in the scene. The natural language corresponding target bounding box is generated by encoding the input image through the vision transformer network to extract image features; passing the input natural language description to a pre-trained language model (BERT). The text description will be segmented and embedded in a fixed-dimensional vector space to obtain the semantic representation of the text; in the referring transformer, the image features and text features are associated with different regions of the image and the specific description in the text through the self-attention mechanism in the transformer architecture, thereby realizing the fusion of image and natural language; through the fused visual-language features, the referring transformer generates a region proposal. These regions are regions that may contain target objects. At the same time, the Referring Transformer uses a convolutional layer to generate the bounding box coordinates of the target object.

[0087] Step 3: Perform instance segmentation on the objects in the current scene based on the depth image data, and perform grasping feasibility detection on the segmented instance masks. The grasping feasibility detection is performed by checking the integrity of the instance mask and whether the blank area around it supports the gripper to extend. The method for instance segmentation based on depth image data is to use Mask R-CNN, and input the current scene image containing multiple objects, which is usually of size H×W. The input image is first processed by a convolutional neural network (ResNet) to obtain a feature map. This process extracts high-level semantic features in the image. The size of the feature map is usually H′×W′×C, where H′ and W′ are the spatial dimensions of the feature map and C is the number of channels; Region Proposal Network (RPN) is used to generate candidate object regions. RPN scans the feature map in a sliding window manner and outputs multiple bounding boxes and their scores; RoIAlign operation is used for each candidate region (RoI) to perform precise spatial alignment to obtain a fixed-size feature map, and a fully connected layer is used to classify each RoI and regress its bounding box. At the same time, Mask R-CNN generates a mask branch on each RoI. This branch uses convolution operations to generate binary masks of object instances, and finally outputs the binary masks of the objects visible in the current scene.

[0088] Step 4: Send the scene depth map and the binary mask of the object detected by grasping feasibility into the interactive object selection strategy to obtain the object to be interacted with. This strategy uses a reinforcement learning algorithm, and the target reward r t Defined by:

[0089]

[0090] where x t Represents the central horizontal coordinate of the target object instance segmentation mask, x o Represents the central horizontal coordinate of the instance segmentation mask of the selected operation object, x j Represents the central abscissa of the instance segmentation mask of each graspable object in the scene. This formula forces the robot arm to choose the object closest to the target for interaction.

[0091] After selecting the object to be interacted with, the action selection strategy is called to determine the direction of moving the object. The action selection strategy also uses the reinforcement learning algorithm, and the target reward r t ' is defined by the following formula:

[0092]

[0093] where v t A vector representing the visible portion of the instance segmentation mask of the target object, which forces the robot to expand the space around the target object for subsequent grasping.

[0094] Step 5: The robot arm moves the obstacle according to the received center coordinate information of the object to be interacted with and the direction of the push action, and repeats the process until the target object appears and can be grasped. The grasping of the object to be interacted with and the target object is based on their center coordinates, which can be calculated from the center point of their bounding box coordinates and obtained through camera coordinate transformation, hand-eye calibration and world coordinate conversion.

[0095] Example 3

[0096] The difference between this embodiment and embodiment 2 is that when the target object in the scene is severely blocked or completely invisible, the object confidence output by the visual language model is lower than the set threshold, that is, the target object cannot be located. At this time, the robot arm needs to explore autonomously in a complex environment to find the target object. Therefore, an additional reinforcement learning process is added to train the robot arm to explore autonomously.

[0097] Specifically, it is to prompt the robot arm to select objects in the unexplored area for interaction and to select actions that can reduce the unexplored area. The definition of the unexplored area is explained in detail below. The reward function is as follows:

[0098] Explore Rewards e Defined by:

[0099]

[0100] This formula enables the robot arm to select objects in unexplored areas as interactive objects.

[0101] After selecting the object to be interacted with, the action selection strategy is called to determine the direction of the moving object. The unexplored area refers to the area where the target may appear. To facilitate the calculation of rewards, the size of the unexplored area is quantified by the vector of the horizontal axis. The unexplored area can be represented by the sum of the one-dimensional distribution of the instance mask on the X-axis of the image, and finally represented as a one-dimensional binary vector v s Since the unexplored area has the property of only decreasing but not increasing, the v before and after the move s 、v s 'And the obtained v f As a quantitative representation of the final unexplored area, the exploration reward r e ' is defined by the following formula:

[0102]

[0103] Where X represents the length of the vector, which is also the length of the X axis of the image, and i represents the index of the vector element.

[0104] As can be seen from the figure below, after reinforcement learning training, the robot arm can explore the environment, discover and grasp the target object. Input natural language, please help me grasp the detergent. Figure 2(a) is the scene depth map at the beginning. The bounding box of the target object obtained after the visual language model is mapped to the depth map, as shown in Figure 2(b). The confidence of the bounding box is less than the threshold, indicating that the target object cannot be found in the current scene. Therefore, the robot arm explores the current scene, inputs the updated scene into the visual language model, and maps the obtained bounding box to the updated scene depth map, as shown in Figure 2(c). The confidence of the bounding box is also lower than the threshold. Continue to explore the environment, and obtain Figure 2(d) through the above process. The confidence of the bounding box is higher than the threshold, so the object is identified as the target object. Through the feasibility detection of grasping, the target object can be grasped, and the robot arm performs the grasping action, as shown in Figure 2(e), and successfully grasps the target object.

[0105] The embodiments of the present invention are described in detail above with reference to the accompanying drawings, but the present invention is not limited to the described embodiments. For those skilled in the art, various changes, modifications, substitutions and variations of these embodiments including components are made without departing from the principles and spirit of the present invention, and still fall within the scope of protection of the present invention.

Claims

1. A robotic arm autonomous grasping system based on a visual language model, characterized in that: Includes the following modules; An image data acquisition module, used to acquire color images and depth images of the working area; The visual language localization module uses the visual language model to generate a bounding box and its confidence around the target object described in natural language based on the color image and natural language description, and determines whether the target object is visible. If visible, it determines its location; The grasping feasibility judgment module determines whether the target object is graspable based on the color image and the target object bounding box. When the target object is not visible or graspable, the object in the current scene is instance segmented. The interactive object selection module selects interactive objects using a reinforcement learning algorithm through the normalized depth image and instance segmentation mask; The action selection module is used to input the local depth image of the interactive object and its mask, and use the reinforcement learning algorithm to determine the action performed by the robotic arm; The robot arm motion module is used to receive the output of the reinforcement learning algorithm and plan the motion trajectory. The robot arm moves according to the motion trajectory to grab or push objects.

2. The robotic arm autonomous grasping system based on the visual language model according to claim 1 is characterized in that: The image data acquisition module also includes preprocessing of the acquired color images and depth images.

3. The robotic arm autonomous grasping system based on the visual language model according to claim 1 is characterized in that: The robotic arm motion module includes a UR10 six-degree-of-freedom robotic arm, an AG-95 two-finger gripper and an end force sensor, wherein the AG-95 two-finger gripper is connected to the end of the UR10 six-degree-of-freedom robotic arm, and the force sensor is located at the fingertip of the AG-95 two-finger gripper, which is used to receive the target position and plan the robotic arm motion trajectory, and cooperate with the two-finger gripper to grasp or move the target object.

4. A control method for a robotic arm autonomous grasping system based on a visual language model, used to control the robotic arm autonomous grasping system according to any one of claims 1 to 3, characterized in that: The steps include: Step 1: Initialize the robot arm and move it to the starting position; collect color images and depth images of the working area, pre-process and save them; Step 2: Take the color image and natural language description as input, generate the target object bounding box and confidence through the Referring Transformer network, and use the confidence to determine whether the target object exists in the scene; Step 3: When there is no target object, all visible objects in the current scene are segmented based on the color image, and the segmented instance masks are tested for grasping feasibility; Step 4: Send the scene depth image and the binary mask of the object detected by grasping feasibility into the interactive object selection strategy and the action selection strategy to obtain the object to be interacted with and the direction of the push action; Step 5: Based on the coordinates of the object bounding box, the center coordinate information of the object to be interacted is obtained through coordinate transformation, and the robotic arm performs the operation of moving the obstacle according to the center coordinate information of the object to be interacted and the pushing action direction; Step 6: Repeat steps 1 to 5 until the target object appears and can be grasped. The robot arm performs a grasping action to grasp the target object and place it at the specified location.

5. The control method of the robot arm autonomous grasping system based on the visual language model according to claim 4 is characterized in that: The specific implementation process of step 2 is as follows: Step 2.1: Input the color image I of the current scene containing multiple objects rgb , input the natural language describing the target object in the scene; Step 2.2: The color image is encoded through the visual transformer network to extract image features; The natural language description is passed to a pre-trained language model BERT; the natural language description is tokenized and embedded into a fixed-dimensional vector space to obtain text features; Step 2.3: In the Referring Transformer, the image features and text features are associated with different regions of the image and the descriptions in the text through the self-attention mechanism in the transformer architecture, thus realizing the fusion of image and natural language. Step 2.4: Using the fused visual-linguistic features, the Referring Transformer generates a region proposal. Each region has a corresponding confidence score. The region with the highest confidence score is selected. If the confidence score of the region is higher than the set threshold, it is considered to be the target object. If the confidence score is lower than the set threshold, it is considered not to be the target object, that is, there is no target object in the scene. Step 2.5, Referring Transformer uses the convolutional layer to generate the bounding box coordinates of the target object.

6. The control method of the robot arm autonomous grasping system based on the visual language model according to claim 5 is characterized in that: The grasping feasibility detection in step 3 is performed by checking the completeness of the instance mask and whether the blank area around it supports the gripper to extend into the instance mask.

7. The control method of the robot arm autonomous grasping system based on the visual language model according to claim 6 is characterized in that: The specific process of instance segmentation in step 3 is as follows: Step 3.1: Input a color image of the current scene containing multiple objects, with a size of H×W. The input image is first processed by a convolutional neural network to obtain a feature map FM, with a size of H′×W′×C, where H′ and W′ are the spatial sizes of the feature map and C is the number of channels. Step 3.2: Based on the feature map FM, use RPN to generate candidate object regions; Step 3.3: Use the RoIAlign operation to spatially align each candidate object region RoI. Step 3.4: Use the fully connected layer to classify each RoI and regress its bounding box. At the same time, on each RoI, MaskR-CNN generates a mask branch; the mask branch uses convolution operations to generate a binary mask of the object instance; Step 3.5, Mask R-CNN outputs the mask of each detected object.

8. The control method of the robot arm autonomous grasping system based on the visual language model according to claim 7 is characterized in that: In step 4, both strategies use the Double DQN network; Step 4.1, the interactive object selection strategy is described as follows: When the target object is not visible, the exploration reward r e Defined by: Where c represents the selected object, p e Indicates areas that need to be explored; When the target object is partially occluded but can be directly located by the visual language model, the target reward r t Defined by: where x t Represents the central horizontal coordinate of the target object instance segmentation mask, x o Represents the central horizontal coordinate of the instance segmentation mask of the selected operation object, x j The central abscissa of the instance segmentation mask representing each graspable object in the scene; Step 4.2, the action selection strategy is described as follows: Based on the scene, the pushing action is limited to two actions, left and right, and the moving distance is fixed at 10cm each time. When the target object is not visible, based on the selected object, the unexplored area should be reduced after moving it. The unexplored area refers to the area where the target object may appear due to being blocked by the surface object. The size of the unexplored area is quantified by the vector of the horizontal axis. The unexplored area is represented by the sum of the one-dimensional distribution of the current scene instance mask on the X-axis of the image, and is finally represented as a one-dimensional binary vector v s ; Move the v before and after s 、v s 'The vector obtained by AND is used as the quantitative representation of the final unexplored area; therefore, the exploration reward r e ' is defined by the following formula: Where X represents the length of the vector, which is also the length of the X axis of the image, i represents the index of the vector element, and · represents and operation; When the target object is partially visible but cannot be grasped directly due to the presence of other objects, the space around it is enlarged, and the target reward r t ' is defined by the following formula: where v t A vector representing the visible portion of the instance segmentation mask of the target object.

9. The control method of the robot arm autonomous grasping system based on the visual language model according to claim 8 is characterized in that: The specific process of obtaining the center coordinate information of the object to be interacted is as follows: The plane pixel coordinates of the object center point are expressed as (x center ,y center ), using the camera internal parameters, including the focal length f x 、f y and light center c x 、c y ; Calculate the plane coordinates of the object in the camera coordinate system: Get the plane coordinates p of the object in the camera coordinate system camera =[x camera ,y camera ]; The transformation matrix T is used to transform the coordinates from the camera coordinate system to the robot end coordinate system. camera_to_arm It is expressed as: Where R camera_to_arm is the rotation matrix, T' camera_to_arm is the translation vector; The conversion process is: P arm =R camera_to_arm ·P camera +T camera_to_arm Convert the plane coordinates in the camera coordinate system to the plane coordinates P in the robot end coordinate system arm ; Finally, the plane coordinate P in the robot end coordinate system arm Convert to the world coordinate system, where the transformation matrix between the robot end coordinate system and the world coordinate system is expressed as: R arm_to_word is the rotation matrix, T' arm_to_word is the translation vector, and the conversion process is: P world =R arm_to_world ·P arm +T arm_to_world Get the plane coordinates P of the center point of the object in the world coordinate system world ,Then, through the algorithm of the MoveIt motion library in ROS, the motion information of the robot arm in the joint space is obtained and planned.

Citation Information

Cited By

  • Lightweight mechanical arm grabbing detection method

    CN120735023A

  • Mechanical arm grabbing method based on self-adaptive updating of stimulation control

    CN121105001A

  • Robotic arm grasping method based on adaptive update of stimulation control

    CN121105001B

  • Smoke sensing test method based on depth camera perception and mechanical arm control

    CN121625118A

  • Smoke sensing test method based on depth camera perception and robot control

    CN121625118B