Object grabbing method and system based on image segmentation and object pose estimation

The object mask is generated through the SAM model and the DINOv2 model, combined with rotation translation sampling and pose refining network, the problem of insufficient pose estimation accuracy in the middle of multi-object scenes is solved, and high-precision and real-time object grabbing is achieved.

CN120269580AActive Publication Date: 2025-07-08SHIRUI (BEIJING) ROBOT TECHNOLOGY CO LTD

Patent Information

Application Number
CN202510769372.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-10
Publication Date
2025-07-08
Estimated Expiration
2045-06-10

AI Technical Summary

Technical Problem

In the prior art, when the object grasping method is partially blocked or there are multiple objects in the scene, the position estimation accuracy is insufficient, making it difficult to meet the high-precision grasping requirements, and the calculation speed is slow, so it cannot meet the real-time control requirements.

Method used

The object mask is generated by the SAM model and the DINOv2 model, combined with the visibility ratio weight adjustment mechanism, and the rotation translation sampling technology and the pose refining network are used to perform preliminary and refined pose estimation, and the robotic arm joint angle is calculated through the inverse kinematics algorithm for grabbing.

Benefits of technology

It improves the accuracy and calculation speed of position estimation, meets the needs of high-precision grabbing, realizes real-time control, and improves the efficiency and stability of robot operation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120269580A_ABST
    Figure CN120269580A_ABST
Patent Text Reader

Abstract

The invention provides an object grabbing method and system based on image segmentation and object pose estimation, and relates to the technical field of computer vision, and the method comprises the steps: obtaining an original image of a to-be-grabbed object; preprocessing the original image to obtain a target image; according to the target image, generating an object mask of the to-be-grabbed object; according to the object mask, determining a 2D bounding box of the object to be grabbed in the target image; the pose of the to-be-grabbed object is preliminarily estimated through the rotation and translation sampling technology, and the preliminary pose is obtained; performing refining estimation on the pose of the to-be-grabbed object through the pose refining network to obtain a refined pose; the refining pose is converted from a camera coordinate system to a mechanical arm coordinate system; on a mechanical arm coordinate system, based on predefined fixed transformation, a target grabbing pose is calculated; and through an inverse kinematics algorithm, the target grabbing pose is converted into a joint angle of the mechanical arm, and the mechanical arm is controlled to grab the to-be-grabbed object according to the joint angle.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer vision technology, and in particular to an object grasping method and system based on image segmentation and object pose estimation. Background Art

[0002] With the rapid development of robot technology, robots are increasingly widely used in fields such as industry, logistics, and healthcare. Especially in object grasping tasks, how to accurately estimate the pose of an object and quickly perform grasping has become one of the key technologies. The pose estimation accuracy and grasping speed directly affect the working efficiency and operation stability of the robot.

[0003] In the prior art, object grasping usually relies on relatively simple pose estimation methods, such as traditional algorithms based on feature matching or template matching. Although these methods can work properly in some simple environments, in complex environments, especially when the object is partially occluded or there are multiple objects in the scene, the estimation accuracy of these methods is often insufficient and difficult to meet the requirements of high-precision grasping.

[0004] In addition, the calculation speed of these traditional methods is slow and cannot meet the requirements of real-time control. Summary of the Invention

[0005] In order to solve the technical problems that in the prior art, the pose estimation methods are often insufficient in estimation accuracy when the object is partially occluded or there are multiple objects in the scene, difficult to meet the requirements of high-precision grasping, and the calculation speed is slow and cannot meet the requirements of real-time control, the present invention provides an object grasping method and system based on image segmentation and object pose estimation.

[0006] The technical solutions provided by the embodiments of the present invention are as follows:

[0007] First aspect:

[0008] An object grasping method based on image segmentation and object pose estimation provided by an embodiment of the present invention includes:

[0009] S1: Obtain the original image of the object to be grasped;

[0010] S2: Preprocess the original image to obtain a target image;

[0011] S3: Generate an object mask of the object to be grasped according to the target image through the SAM model and the DINOv2 model;

[0012] S4: Determine the 2D bounding box of the object to be grasped in the target image according to the object mask;

[0013] S5: Based on the median depth of the 2D bounding box, use the rotation and translation sampling technique to perform a preliminary estimation of the pose of the object to be grasped, and obtain multiple preliminary poses;

[0014] S6: According to each of the preliminary poses and the target image, use a pose refinement network to perform a refined estimation of the pose of the object to be grasped, and obtain a refined pose;

[0015] S7: Convert the refined pose from the camera coordinate system to the robotic arm coordinate system;

[0016] S8: On the robotic arm coordinate system, based on the predefined fixed transformation between the end effector and the object to be grasped, calculate the target grasping pose;

[0017] S9: Through the inverse kinematics algorithm, convert the target grasping pose into the joint angles of the robotic arm;

[0018] S10: According to the joint angles, control the robotic arm to grasp the object to be grasped.

[0019] Second aspect:

[0020] An object grasping system based on image segmentation and object pose estimation provided by an embodiment of the present invention includes:

[0021] A processor;

[0022] A memory, on which computer-readable instructions are stored, and when the computer-readable instructions are executed by the processor, the object grasping method based on image segmentation and object pose estimation as described in the first aspect is implemented.

[0023] Third aspect:

[0024] A computer-readable storage medium provided by an embodiment of the present invention, on which a computer program is stored, and when the program is executed by a processor, the object grasping method based on image segmentation and object pose estimation as described in the first aspect is implemented.

[0025] The beneficial effects brought by the technical solutions provided by the embodiments of the present invention at least include:

[0026] In the embodiments of the present invention, an object mask of the object to be grasped is generated through the SAM model and the DINOv2 model. When the object is partially occluded or there are multiple objects in the scene, the accuracy of pose estimation is improved, which can meet the requirements of high-precision grasping. According to the object mask, a 2D bounding box of the object to be grasped in the target image is determined. Based on the median depth of the 2D bounding box, the pose of the object to be grasped is initially estimated through the rotation and translation sampling technique to obtain multiple initial poses. According to each initial pose and the target image, the pose of the object to be grasped is refined through a pose refinement network to obtain a refined pose, which improves the calculation speed and can meet the requirements of real-time control. BRIEF DESCRIPTION OF THE DRAWINGS

[0027] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0028] Figure 1 It is a schematic flow chart of an object grasping method based on image segmentation and object pose estimation provided by an embodiment of the present invention;

[0029] Figure 2 It is a schematic structural diagram of an object grasping system based on image segmentation and object pose estimation provided by an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0030] The following will describe the technical solutions in the present invention with reference to the drawings.

[0031] In the embodiments of the present invention, words such as "exemplarily" and "for example" are used to indicate examples, illustrations or explanations. Any embodiment or design solution described as an "example" in the present invention should not be construed as being more preferred or more advantageous than other embodiments or design solutions. Exactly speaking, the use of the word "example" is intended to present concepts in a specific way. In addition, in the embodiments of the present invention, the meaning expressed by "and / or" can be both, or either one of the two.

[0032] In the embodiments of the present invention, "image" and "picture" can sometimes be used interchangeably. It should be noted that when the difference is not emphasized, the meaning they express is the same. "(of)", "corresponding", and "corresponding" can sometimes be used interchangeably. It should be noted that when the difference is not emphasized, the meaning they express is the same.

[0033] In the embodiments of the present invention, sometimes a subscript such as W1 may be written in a non-subscript form such as W1. When the difference is not emphasized, the meanings they express are the same.

[0034] To make the technical problems, technical solutions, and advantages to be solved by the present invention clearer, the following will be described in detail with reference to the accompanying drawings and specific embodiments.

[0035] Refer to the attached Figure 1 figures, which show a schematic flow diagram of an object grasping method based on image segmentation and object pose estimation provided by an embodiment of the present invention.

[0036] The embodiments of the present invention provide an object grasping method based on image segmentation and object pose estimation. This method can be implemented by an object grasping device based on image segmentation and object pose estimation. The object grasping device based on image segmentation and object pose estimation can be a terminal or a server. The processing flow of the object grasping method based on image segmentation and object pose estimation can include the following steps:

[0037] S1: Obtain the original image of the object to be grasped.

[0038] Specifically, an industrial control computer is used to control an RGBD camera to obtain the original image of the object to be grasped.

[0039] It should be noted that the original image includes the color image and the depth image of the object to be grasped.

[0040] It should be noted that an RGBD camera is a camera that combines a traditional RGB camera and a depth sensor. It can not only capture the color image (RGB image) of an object, but also obtain the depth information (D, i.e., Depth) of each pixel point, that is, the distance between the object and the camera. This enables the RGBD camera to provide depth information in three-dimensional space, thereby helping the computer vision system to perceive and understand the shape, position, and spatial relationship of objects in the scene. Common RGBD cameras include Microsoft's Kinect and Intel's RealSense series, which are widely used in fields such as robotics, virtual reality, augmented reality, and 3D modeling.

[0041] In the present invention, by using an RGBD camera to obtain the original image of the object to be grasped, the color image and depth information of the object can be obtained simultaneously, thereby providing the accurate position and shape of the object in three-dimensional space. This enables the system to better understand the spatial relationship of the object, provides an accurate data basis for subsequent object recognition, pose estimation, and grasping operations, and improves the grasping accuracy and stability, especially in complex environments or cases of partial occlusion.

[0042] S2: Preprocess the original image to obtain the target image.

[0043] Specifically, the industrial control computer performs an alignment operation on the acquired color image and depth image to obtain a target image.

[0044] Optionally, the color image and the depth image are aligned according to the following formula:

[0045]

[0046] where I RGB (x,y) represents the pixel value at the coordinate (x,y) in the original image, represents the pixel value at the coordinate in the depth image.

[0047] It should be noted that the alignment operation refers to making the pixels in the color image and the depth image spatially corresponding, so that the pixel positions of the two match one by one, thereby obtaining a target image that combines color and depth information. This operation usually uses the internal parameter matrix of the camera and the coordinate transformation method to map the three-dimensional spatial information in the depth image to the two-dimensional pixel coordinate system of the color image, so as to ensure that the depth value can accurately correspond to the color information and provide an accurate data basis for subsequent image analysis and processing.

[0048] In the present invention, by performing an alignment operation on the color image and the depth image, the pixel positions of the two are matched one by one, thereby obtaining a target image that combines color and depth information. This alignment operation uses the internal parameter matrix of the camera and the coordinate transformation method to map the three-dimensional spatial information in the depth image to the two-dimensional pixel coordinate system of the color image, so that the depth value can accurately correspond to the color information. In this way, the visual features and spatial features of the object can be obtained simultaneously, providing an accurate data basis for subsequent tasks such as object recognition, pose estimation, and grasping, reducing noise and errors, and improving the accuracy and robustness of subsequent processing. Especially in complex environments and cases of object occlusion, it ensures the precise positioning and grasping of the object.

[0049] S3: According to the target image, an object mask of the object to be grasped is generated through the SAM model and the DINOv2 model.

[0050] It should be noted that the SAM model (Segment Anything Model) is an image segmentation model designed to automatically generate object masks from images. Through deep learning techniques, the SAM model can accurately segment objects in images, and is applicable not only to known object types but also to unknown objects.

[0051] The DINOv2 model (Distilled Knowledge Network v2) is a self-supervised learning model that focuses on image feature extraction and representation learning. It is pre-trained with large-scale unlabeled data and can extract fine-grained visual features from images, including semantic, appearance, and geometric information. DINOv2 has excellent performance in tasks such as image segmentation and object detection. Especially in the zero-shot setting, it can effectively associate visual information with the context and enhance the model's ability to handle unseen objects.

[0052] Optionally, the SAM model adopts 42 perspective templates rendered by Blender.

[0053] It should be noted that Blender is an open-source 3D creation software widely used in multiple fields such as modeling, rendering, animation, sculpting, texturing, simulation, and video editing. It provides a complete set of tools suitable for artists and developers to create 3D content. Blender supports multiple 3D production workflows, including from modeling, sculpting, animation production to rendering, physical simulation, etc. Its rendering engines, such as Eevee and Cycles, can generate high-quality images and animations and are widely used in film, game development, and virtual reality production.

[0054] In the present invention, by combining the SAM model, the DINOv2 model, and 42 perspective templates rendered by Blender, the accuracy and robustness of object segmentation and pose estimation can be significantly improved. Through the 42 perspective templates rendered by Blender, the system can generate object images from multiple angles, increasing the diversity of visual information and helping the system accurately identify and segment objects even when the object is partially occluded or in a complex environment, thus achieving efficient object grasping and pose estimation.

[0055] In a possible implementation, S3 specifically includes sub-steps S301 to S303:

[0056] S301: Segment the target image through the SAM model to generate multiple candidate segmentation regions.

[0057] Optionally, segment the target image according to the following formula to generate multiple candidate segmentation regions.

[0058]

[0059] where M represents the candidate segmentation region, C represents the confidence score corresponding to the candidate segmentation region, represents the mask decoder of the SAM model, which is used to generate candidate segmentation regions and confidence scores, An image encoder representing the SAM model is used to extract image features. I represents the input target image. A prompt encoder is used to input prompt information. Pr represents the prompt information.

[0060] S302: Extract the image features of each candidate segmentation region through the DINOv2 model.

[0061] Optionally, according to the following formula, extract the image features of each candidate segmentation region:

[0062]

[0063] Where F m Represents the image features of the m-th candidate segmentation region extracted. Represents the image feature extraction process of the DINOv2 model. Preprocess( ) represents preprocessing the candidate segmentation region (such as resizing or normalizing). M m Represents the m th candidate segmentation region.

[0064] Optionally, the image features include: semantic features, appearance features, and geometric features.

[0065] S303: Generate an object mask according to the image features of each candidate segmentation region through a visibility ratio weight adjustment mechanism.

[0066] It should be noted that the visibility ratio weight adjustment mechanism is a strategy used in image processing or computer vision tasks to handle cases of partial occlusion or incomplete visibility of objects. In image segmentation or object detection, an object may be partially hidden due to the viewing angle or background occlusion. This mechanism analyzes the ratio of the visible area to the occluded area of the object in the image and dynamically adjusts the weights of each part to ensure the generation of the object mask as accurately as possible. When segmenting an object, the visibility ratio weight adjustment mechanism can give higher weights to the visible parts of the object, thereby improving the accuracy of the segmentation results, especially in complex or partially occluded scenarios.

[0067] In the present invention, the SAM model can accurately segment objects in images through deep learning technology, adapt to the detection of known and unknown objects, and perform excellently especially in the case of complex backgrounds and partial occlusions. As a self-supervised learning model, the DINOv2 model can extract fine-grained visual features from unlabeled data, enhance the model's ability to process new objects, and improve the accuracy of image segmentation and object detection. Combining with the visibility ratio weight adjustment mechanism, the system can adjust the weights according to the visible parts of the objects, especially in the case of object occlusion or complex backgrounds, to ensure the accuracy of the segmentation results.

[0068] In a possible implementation manner, S303 specifically includes sub-steps S3031 to S3035:

[0069] S3031: According to the semantic features, perform matching scoring on each candidate segmentation region to obtain the semantic matching scores of each candidate segmentation region.

[0070] Optionally, according to the following formula, perform matching scoring on each candidate segmentation region to obtain the semantic matching scores of each candidate segmentation region:

[0071]

[0072] where S sem represents the semantic matching score of the candidate segmentation region, topK represents selecting the top K candidate segmentation regions with the strongest similarity, represents the CLS embedding of the m-th candidate segmentation region and represents the CLS embedding of the k-th template image .

[0073] It should be noted that the CLS embedding is a feature representation in the DINOv2 model, which extracts high-level semantic features from images through self-supervised learning. Each candidate segmentation region and template image will generate a CLS embedding to represent its semantic features. Suppose we have an image containing multiple objects, and the goal is to identify a specific cup from it. Multiple candidate segmentation regions are generated through the SAM model, and then the DINOv2 model is used to extract the CLS embedding of each region. Calculate the cosine similarity between the CLS embedding of each candidate region and the CLS embedding of the cup template, and finally obtain the semantic matching score of each candidate region. The candidate region with the highest score will be selected as the object mask of the cup.

[0074] S3032: According to the appearance features, perform matching scoring on each candidate segmentation region to obtain the appearance matching scores of each candidate segmentation region.

[0075] Optionally, according to the following formula, match scores are calculated for each candidate segmentation region to obtain the appearance match scores of each candidate segmentation region:

[0076]

[0077] where S appe represents the appearance match score of the candidate segmentation region, represents the number of block embeddings of the m-th candidate segmentation region max represents taking the maximum value, represents the j-th block embedding of the m-th candidate segmentation region, represents the i-th block embedding of the best-matching template image of.

[0078] It should be noted that the appearance match score is used to evaluate the matching degree of the candidate segmentation region with the target object by analyzing the appearance features of the candidate segmentation region. Suppose we have an image containing multiple objects, and the goal is to identify a specific cup from it. Multiple candidate segmentation regions are generated through the SAM model, and then the block embeddings of each region are extracted using the DINOv2 model. The cosine similarity between the block embeddings of each candidate region and the block embeddings of the cup template is calculated, and finally the appearance match score of each candidate region is obtained. The candidate region with the highest score will be selected as the object mask of the cup.

[0079] S3033: According to the geometric features, match scores are calculated for each candidate segmentation region to obtain the geometric match scores of each candidate segmentation region.

[0080] Optionally, according to the following formula, match scores are calculated for each candidate segmentation region to obtain the geometric match scores of each candidate segmentation region:

[0081]

[0082] where S geo represents the geometric match score of the candidate segmentation region, represents the 2D bounding box of the m-th candidate segmentation region, represents the 2D bounding box of the object to be grasped after projection of the initial pose.

[0083] It should be noted that in the object grasping method based on image segmentation and object pose estimation, the geometric matching score is used to evaluate the matching degree of a candidate segmentation region with the target object by analyzing its geometric features. Geometric features usually include information such as the bounding box, shape, and position of the object. The geometric matching degree is evaluated by calculating the intersection and union between the bounding box of the candidate segmentation region and the bounding box of the target object. Suppose we have an image containing multiple objects, and the goal is to identify a specific cup from it. Multiple candidate segmentation regions are generated through the SAM model, and then the IoU between the bounding box of each candidate region and the bounding box of the cup template is calculated to finally obtain the geometric matching score of each candidate region. The candidate region with the highest score will be selected as the object mask of the cup.

[0084] S3034: Through the visibility ratio weight adjustment mechanism, the semantic matching score, appearance matching score, and geometric matching score of each candidate segmentation region are weighted and summed to obtain the comprehensive matching score of each candidate segmentation region.

[0085] Optionally, according to the following formula, through the visibility ratio weight adjustment mechanism, the semantic matching score, appearance matching score, and geometric matching score of each candidate segmentation region are weighted and summed to obtain the comprehensive matching score of each candidate segmentation region:

[0086]

[0087] where s m represents the comprehensive matching score of the candidate segmentation region, and r vis represents the visibility ratio weight.

[0088] S3035: Select the candidate segmentation region with the highest comprehensive matching score as the object mask.

[0089] In the present invention, by weighting the matching scores of semantic features, appearance features, and geometric features and combining the visibility ratio weight adjustment mechanism, the accuracy and robustness of object segmentation can be effectively improved. Semantic features, appearance features, and geometric features evaluate the segmentation regions of the object from different perspectives, making the score of each candidate segmentation region more comprehensive and accurate. The visibility ratio weight adjustment mechanism can better handle the situation of object occlusion and partial visibility by adjusting the weight of each candidate segmentation region. By giving priority to the scores of the visible parts, the missegmentation caused by occlusion or incomplete object information can be reduced, thereby improving the accuracy of the mask.

[0090] Specifically, the industrial control computer sends the target image to the server, and the server generates the object mask.

[0091] Furthermore, the server performs pose estimation based on the generated object mask to obtain the refined pose.

[0092] S4: Determine the 2D bounding box of the object to be grasped in the target image according to the object mask.

[0093] In the present invention, generating a 2D bounding box through the object mask can clearly determine the position and size of the object in the image, providing precise spatial information for subsequent pose estimation. The bounding box helps the algorithm focus on the effective area of the object, reducing the interference of the background and noise on pose estimation, thereby improving the accuracy of the estimation.

[0094] S5: Based on the median depth of the 2D bounding box, preliminarily estimate the pose of the object to be grasped through the rotation and translation sampling technique to obtain multiple preliminary poses.

[0095] It should be noted that the median depth refers to the depth value at the middle position in the depth data of an object or scene in three-dimensional space. In a depth image, each pixel represents the distance from the object surface to the camera. The median depth is represented by sorting all depth values and selecting the value at the middle.

[0096] It should be noted that the rotation and translation sampling technique is a method for object pose estimation. By sampling the pose of the object in three-dimensional space, the position and orientation of the object are determined. Specifically, the rotation and translation sampling technique uniformly generates multiple sampling points around the object, which include the rotation angle and position offset of the object, thereby generating multiple possible object poses. In the rotation stage, the sampling rotates around the object at different angles. In the translation stage, the sampling involves different position offsets of the object in space. Through this technique, a set of possible object poses can be obtained for subsequent precise estimation and optimization, thereby improving the accuracy of pose estimation.

[0097] In the present invention, the median depth provides stable depth information by eliminating noise and extreme values in the depth image, providing a reliable starting point for pose estimation. The rotation and translation sampling technique ensures that the system can comprehensively explore the pose of the object by uniformly generating multiple sampling points around the object, covering various possible poses and positions of the object, especially in complex and occluded environments.

[0098] In a possible implementation manner, S5 specifically includes sub-steps S501 to S503:

[0099] S501: Take the position of the 3D point corresponding to the median depth of the 2D bounding box as the central position of the object to be grasped.

[0100] In the present invention, by using the 3D point corresponding to the median depth of the 2D bounding box as the central position of the object to be grasped, it helps to ensure that the pose estimation starts from the true center of the object, avoiding deviations caused by errors and making the subsequent pose estimation more accurate.

[0101] S502: Sample poses at multiple viewpoints on the icosahedron at the central position of the object to be grasped in a uniform sampling manner.

[0102] In the present invention, by uniformly sampling multiple viewpoints on the icosahedron at the central position of the object, comprehensive sampling can be carried out from multiple angles, which means that the algorithm can consider all possible rotation angles of the object, thus avoiding errors or omissions that may be brought by relying only on a single viewpoint and enhancing the reliability of pose estimation.

[0103] S503: Perform in-plane rotation sampling on the poses at each viewpoint to obtain multiple preliminary poses.

[0104] In the present invention, by performing in-plane rotation sampling on the poses at each sampled viewpoint to further optimize the rotation information of the object, multiple different preliminary poses can be generated, providing diverse candidate poses for subsequent precise optimization. This method is particularly suitable for dealing with complex environments and occlusion scenarios, ensuring that the pose estimation is more accurate and robust in practical applications.

[0105] S6: According to each preliminary pose and the target image, through a pose refinement network, refine the pose of the object to be grasped to obtain a refined pose.

[0106] Optionally, the pose refinement network is specifically a convolutional neural network.

[0107] It should be noted that the convolutional neural network (CNN) is a deep learning model widely used in fields such as image recognition, video analysis, and natural language processing. The CNN processes data by mimicking the working mode of the biological visual system. It consists of multiple convolutional layers, pooling layers, and fully connected layers. In the convolutional layer, the network uses filters (also known as convolutional kernels) to perform convolutional operations on the input image to extract local features in the image, such as edges and textures. The pooling layer reduces the dimension of the data by downsampling the image while retaining important information. Through multi-level convolutional and pooling operations, the CNN can gradually learn more advanced features and finally perform tasks such as classification or regression through the fully connected layer.

[0108] In the present invention, the CNN gradually extracts local features and global features in the image through convolutional layers and pooling layers, enabling the model to learn deep information of the image, such as the edges, textures, shapes of objects, etc. This feature extraction ability enables the pose refinement process to extract more fine-grained features from the target image, thereby effectively judging the matching degree between different poses.

[0109] In a possible implementation manner, S6 specifically includes sub-steps S601 to S604:

[0110] S601: Generate multiple rendered images of the object to be grasped according to each preliminary pose.

[0111] S602: Extract the feature map of the target image through a convolutional neural network.

[0112] S603: Compare the matching degrees between each rendered image and the feature map to obtain the scores of the poses in each rendered image.

[0113] Optionally, according to the following formula, compare the matching degrees between each rendered image and the feature map to obtain the scores of the poses in each rendered image:

[0114]

[0115] where s hyp represents the score of the pose in the rendered image, represents the sparse point coordinates of the y-th rendered image, represents the sparse point set of the y-th rendered image, min represents taking the minimum value, represents the sparse point coordinates of the z-th feature map, represents the sparse point set of the z-th feature map, represents the assumed rotation matrix, represents the transpose operation, represents the assumed translation vector, represents the sparse point set of the number of points.

[0116] It should be noted that both the rendered image and the feature map are represented as sparse point sets, and each point represents a key feature in the image. Suppose we have an image containing multiple objects, and the goal is to identify a specific cup from it. Multiple rendered images are generated through the preliminary pose, and then the feature map of the target image is extracted using the CNN. Calculate the matching degree between the sparse point set of each rendered image and the sparse point set of the feature map, and finally obtain the pose score of each rendered image. The pose represented by the rendered image with the highest score will be selected as the refined pose.

[0117] S604: Select the pose with the highest score as the refined pose.

[0118] In the present invention, by generating multiple rendered images and using a CNN to extract the feature map of the target image, the matching degree of different poses can be accurately compared. When faced with image data in different environments, the CNN can effectively process noise, occlusion, and images under different lighting conditions. This enables the pose refinement network to maintain high robustness and adaptability in various dynamic scenarios. Even in the presence of unseen objects or complex visual environments, it can still give accurate pose estimates. Through efficient convolutional operations and feature map extraction, the deep learning framework of the CNN can process a large amount of image data in a short time, thereby accelerating the calculation process of pose refinement.

[0119] Furthermore, the server returns the generated refined pose to the industrial control computer, and the industrial control computer determines the joint angles of the robotic arm according to the refined pose.

[0120] S7: Convert the refined pose from the camera coordinate system to the robotic arm coordinate system.

[0121] Optionally, according to the following formula, convert the refined pose from the camera coordinate system to the robotic arm coordinate system:

[0122]

[0123] where, represents the refined pose in the robotic arm coordinate system, represents the transformation matrix from the camera coordinate system to the robotic arm coordinate system (which can be obtained through the calibration process), represents the refined pose in the camera coordinate system.

[0124] In the present invention, converting the refined pose from the camera coordinate system to the robotic arm coordinate system can ensure that the pose information of the object can be accurately applied to the control system of the robotic arm. Since the camera and the robotic arm are in different coordinate systems, this conversion process can eliminate the coordinate differences between the two, enabling the robotic arm to perform precise operations based on the object position information obtained by the camera, thereby improving the accuracy and stability of the robot in the object grasping task and ensuring that the robotic arm can correctly position and execute the grasping action.

[0125] S8: On the robotic arm coordinate system, calculate the target grasping pose based on the predefined fixed transformation between the end effector and the object to be grasped.

[0126] It should be noted that the fixed transformation between the predefined end effector and the object to be grasped refers to the spatial relationship between the end effector (such as a robotic gripper, fixture, etc.) and the object to be grasped in the robotic system, which has been set in advance and remains unchanged. This transformation usually includes information such as the relative position and orientation (rotation) between the object and the end effector. Through this fixed transformation, it can be ensured that when the robotic arm performs a grasping task, the end effector can accurately contact the object, regardless of its position or posture. This fixed transformation is usually predefined in the robotic control system and ensures that the end effector can correctly grasp the object through precise coordinate system transformation.

[0127] Furthermore, in the operation of the robotic arm, the end effector is the last component of the robotic arm and is used to directly contact external objects, such as grippers, suction cups, or robotic hands. The object to be grasped is the target object that the robotic arm needs to operate on. The fixed transformation refers to the predefined and fixed geometric relationship between the end effector and the object to be grasped. The fixed transformation includes two parts: rotation and translation, and is usually represented by a 4x4 homogeneous transformation matrix.

[0128] In a possible implementation manner, S8 is specifically:

[0129] Calculate the target grasping pose according to the following formula:

[0130]

[0131] Where, represents the target grasping pose in the robotic arm coordinate system, represents the refined pose in the robotic arm coordinate system, represents the inverse operation of the fixed transformation between the predefined end effector and the object to be grasped.

[0132] In the present invention, by calculating the target grasping pose according to the fixed transformation between the predefined end effector and the object to be grasped, it can be ensured that the end effector of the robotic arm can always contact the object with the correct posture and position. The fixed transformation provides a stable spatial relationship between the object and the end effector, enabling the robotic arm to accurately interact with the object when performing a grasping task, regardless of how the specific position or posture of the object changes. In this way, the system can achieve more efficient and stable grasping operations, avoid mis-grasping caused by pose changes, and improve the success rate of grasping tasks and the accuracy of robotic operations.

[0133] S9: Convert the target grasping pose into the joint angles of the robotic arm through the inverse kinematics algorithm.

[0134] It should be noted that the inverse kinematics algorithm is a mathematical method used to calculate the joint angles or positions of a robot. Its purpose is to inversely solve the angles or displacements that each joint needs to reach according to the target position and posture of the end effector of the robotic arm. Usually, the robot control system requires the end effector (such as a robotic gripper) to reach a specific spatial position and orientation, and the inverse kinematics algorithm calculates how to adjust the angles or positions of the joints of the robotic arm based on the position and posture information of the end effector to achieve this goal. The inverse kinematics problem is very important in robotics because it converts the target in three-dimensional space into instructions in joint space, ensuring that the robotic arm can accurately perform tasks such as grasping and moving.

[0135] In the present invention, by converting the target grasping pose into the joint angles of the robotic arm through the inverse kinematics algorithm, it can ensure that the robotic arm can accurately perform the grasping task according to the predetermined target position and posture. The inverse kinematics algorithm calculates the required angles or displacements of each joint by inversely calculating based on the target position and posture of the end effector, enabling the robotic arm to precisely control the movement of each joint, thereby achieving high-precision grasping operations. This process can effectively convert the target in three-dimensional space into instructions in the joint space of the robotic arm, ensuring that the robotic arm can accurately and stably perform the grasping task in a complex environment, improving the efficiency and precision of robot operations.

[0136] S10: Control the robotic arm to grasp the object to be grasped according to the joint angles.

[0137] Specifically, the industrial control computer controls the robotic arm to grasp the object to be grasped according to the calculated joint angles.

[0138] In the present invention, by controlling the robotic arm to grasp through the industrial control computer according to the calculated joint angles, it can ensure that the robotic arm accurately performs the grasping task. This control method enables the robotic arm to accurately move according to the calculated joint angles, avoiding grasping failures caused by errors or instability, ensuring the accuracy and reliability of the grasping action. Especially in complex environments or when the object posture is irregular, it greatly improves the efficiency and success rate of robot grasping.

[0139] Refer to the attached Figure 2 to show the structural schematic diagram of an object grasping system provided by the present invention based on image segmentation and object pose estimation.

[0140] The present invention also provides an object grasping system 20 based on image segmentation and object pose estimation, which is applied to the above-mentioned object grasping method based on image segmentation and object pose estimation, and includes:

[0141] A processor 201;

[0142] A memory 202 stores computer-readable instructions thereon. When the computer-readable instructions are executed by the processor 201, an object grasping method based on image segmentation and object pose estimation as described in the method embodiment is implemented.

[0143] The object grasping system 20 based on image segmentation and object pose estimation provided by the present invention can execute the above object grasping method based on image segmentation and object pose estimation and achieve the same or similar technical effects. To avoid repetition, the present invention will not be elaborated herein.

[0144] The beneficial effects brought by the technical solution provided by the embodiment of the present invention at least include:

[0145] In the embodiment of the present invention, an object mask of the object to be grasped is generated through the SAM model and the DINOv2 model. When the object is partially occluded or there are multiple objects in the scene, the accuracy of pose estimation is improved, and the requirements of high-precision grasping can be met. According to the object mask, a 2D bounding box of the object to be grasped in the target image is determined. Based on the median depth of the 2D bounding box, the pose of the object to be grasped is initially estimated through the rotation and translation sampling technique to obtain multiple initial poses. According to each initial pose and the target image, the pose of the object to be grasped is refined and estimated through the pose refinement network to obtain a refined pose, which improves the calculation speed and can meet the requirements of real-time control.

[0146] It should be understood that the processor in the embodiment of the present invention may be a central processing unit (CPU), and the processor may also be other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field programmable gate arrays (FPGAs) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc.

[0147] It should also be understood that the memory in the embodiments of the present invention can be a volatile memory or a non-volatile memory, or can include both volatile and non-volatile memories. Among them, the non-volatile memory can be a read-only memory (ROM), a programmable ROM (PROM), an erasable PROM (EPROM), an electrically erasable PROM (EEPROM), or a flash memory. The volatile memory can be a random access memory (RAM), which is used as an external cache. By way of example but not limitation, many forms of random access memory (RAM) are available, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchlink DRAM (SLDRAM), and direct rambus RAM (DR RAM).

[0148] The above embodiments can be implemented in whole or in part by software, hardware (such as circuits), firmware, or any combination thereof. When implemented using software, the above embodiments can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions or computer programs. When the computer instructions or computer programs are loaded or executed on a computer, the processes or functions described in the embodiments of the present invention are generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired (such as infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any available medium that can be accessed by a computer or a data storage device such as a server or data center that contains one or more collections of available media. The available media can be magnetic media (such as floppy disks, hard disks, magnetic tapes), optical media (such as DVDs), or semiconductor media. The semiconductor media can be a solid-state drive.

[0149] It should be understood that the term "and / or" in this document is merely a description of the association relationship between associated objects, indicating that three relationships can exist. For example, A and / or B can represent: A exists alone, A and B exist simultaneously, and B exists alone. Here, A and B can be singular or plural. Additionally, the character " / " in this document generally represents an "or" relationship between the associated objects before and after, but it may also represent an "and / or" relationship, which can be specifically understood by referring to the context before and after.

[0150] In the present invention, "at least one" means one or more, and "a plurality" means two or more. "At least one of the following" or its similar expressions refer to any combination of these items, including any combination of single items or plural items. For example, at least one of a, b, or c can represent: a, b, c, a - b, a - c, b - c, or a - b - c, where a, b, and c can be single or multiple.

[0151] It should be understood that in various embodiments of the present invention, the magnitudes of the sequence numbers of the above processes do not imply the order of execution. The order of execution of each process should be determined by its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of the present invention.

[0152] Those of ordinary skill in the art can realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present invention.

[0153] Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the devices, apparatuses, and units described above can refer to the corresponding processes in the foregoing method embodiments, and will not be repeated here.

[0154] In several embodiments provided by the present invention, it should be understood that the disclosed devices, apparatuses, and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another device, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections to each other can be through some interfaces. The indirect couplings or communication connections of the devices or units can be in electrical, mechanical, or other forms.

[0155] The units described as separate components may or may not be physically separated. The components displayed as units may or may not be physical units, that is, they can be located in one place, or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0156] In addition, the functional units in various embodiments of the present invention can be integrated into one processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit.

[0157] When the above-mentioned function is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention. The foregoing storage medium includes: various media such as USB flash drives, external hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical discs that can store program codes.

[0158] An embodiment of the present invention provides a computer-readable storage medium, on which a computer program is stored, and is characterized in that when the program is executed by a processor, it implements the object grasping method based on image segmentation and object pose estimation as described in the method embodiment.

[0159] The computer-readable storage medium provided by the present invention can implement the steps and effects of the object grasping method based on image segmentation and object pose estimation in the above method embodiment. To avoid repetition, the present invention will not elaborate further.

[0160] The beneficial effects brought by the technical solution provided by the embodiments of the present invention at least include:

[0161] In the embodiments of the present invention, an object mask of the object to be grasped is generated through the SAM model and the DINOv2 model. When the object is partially occluded or there are multiple objects in the scene, the accuracy of pose estimation is improved, and the requirements for high-precision grasping can be met. According to the object mask, the 2D bounding box of the object to be grasped in the target image is determined. Based on the median depth of the 2D bounding box, the pose of the object to be grasped is initially estimated through the rotation and translation sampling technique to obtain multiple initial poses. According to each initial pose and the target image, through the pose refinement network, the pose of the object to be grasped is refined and estimated to obtain the refined pose, which improves the calculation speed and can meet the requirements of real-time control.

[0162] The above is only the specific implementation manner of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention can easily think of changes or substitutions, which should all be covered within the protection scope of the present invention. Therefore, the protection scope of the present invention should be subject to the protection scope of the claims.

[0163] The following points need to be explained:

[0164] (1) The accompanying drawings of the embodiments of the present invention only relate to the structures involved in the embodiments of the present invention, and other structures can refer to the general design.

[0165] (2) For clarity, in the accompanying drawings used to describe the embodiments of the present invention, the thickness of layers or regions is enlarged or reduced, that is, these drawings are not drawn to actual scale. It can be understood that when an element such as a layer, film, region, or substrate is referred to as being "on" or "under" another element, the element can be "directly" on or under the other element or there can be intermediate elements.

[0166] (3) Without conflict, the embodiments of the present invention and the features in the embodiments can be combined with each other to obtain new embodiments.

[0167] The above is only the specific implementation manner of the present invention, but the protection scope of the present invention is not limited thereto. The protection scope of the present invention shall be subject to the protection scope of the claims.

Claims

1. An object grasping method based on image segmentation and object pose estimation, characterized in that, Including: S1: Obtain the original image of the object to be grasped; S2: Preprocess the original image to obtain the target image; S3: Generate the object mask of the object to be grasped based on the target image through the SAM model and the DINOv2 model; S4: Determine the 2D bounding box of the object to be grasped in the target image according to the object mask; S5: Based on the median depth of the 2D bounding box, preliminarily estimate the pose of the object to be grasped through the rotation and translation sampling technique to obtain multiple preliminary poses; S6: According to each of the preliminary poses and the target image, refine the estimate of the pose of the object to be grasped through the pose refinement network to obtain the refined pose; S7: Convert the refined pose from the camera coordinate system to the robotic arm coordinate system; S8: On the robotic arm coordinate system, calculate the target grasping pose based on the predefined fixed transformation between the end effector and the object to be grasped; S9: Convert the target grasping pose into the joint angles of the robotic arm through the inverse kinematics algorithm; S10: Control the robotic arm to grasp the object to be grasped according to the joint angles.

2. The object grasping method based on image segmentation and object pose estimation according to claim 1, characterized in that, The specific content of S3 includes: S301: Segment the target image through the SAM model to generate multiple candidate segmentation regions; S302: Extract the image features of each of the candidate segmentation regions through the DINOv2 model; S303: Generate the object mask according to the image features of each of the candidate segmentation regions through the visibility ratio weight adjustment mechanism.

3. The object grasping method based on image segmentation and object pose estimation according to claim 2, characterized in that, The image features include: semantic features, appearance features, and geometric features.

4. The object grasping method based on image segmentation and object pose estimation according to claim 3, characterized in that, The specific content of S303 includes: S3031: Perform matching scoring on each of the candidate segmentation regions according to the semantic features to obtain the semantic matching scores of each of the candidate segmentation regions; S3032: Perform matching scoring on each of the candidate segmentation regions according to the appearance features to obtain the appearance matching scores of each of the candidate segmentation regions; S3033: Perform matching scoring on each of the candidate segmentation regions according to the geometric features to obtain the geometric matching scores of each of the candidate segmentation regions; S3034: Through the visibility ratio weight adjustment mechanism, perform weighted summation on the semantic matching scores, appearance matching scores, and geometric matching scores of each of the candidate segmentation regions to obtain the comprehensive matching scores of each of the candidate segmentation regions; S3035: Select the candidate segmentation region with the highest comprehensive matching score as the object mask.

5. The object grasping method based on image segmentation and object pose estimation according to claim 1, characterized in that, The specific content of S5 includes: S501: Take the position of the 3D point corresponding to the median depth of the 2D bounding box as the center position of the object to be grasped; S502: Sample the poses at multiple viewpoints on the icosahedron at the center position of the object to be grasped in a uniform sampling manner; S503: Perform in-plane rotation sampling on the poses at each of the viewpoints to obtain multiple of the preliminary poses.

6. The object grasping method based on image segmentation and object pose estimation according to claim 1, characterized in that, The pose refinement network is specifically a convolutional neural network.

7. The object grasping method based on image segmentation and object pose estimation according to claim 6, wherein The specific content of S6 includes: S601: Generate multiple rendered images of the object to be grasped according to each of the preliminary poses; S602: Extract the feature map of the target image through the convolutional neural network; S603: Compare the matching degrees between each of the rendering images and the feature map, and obtain the scores of the poses in each of the rendering images; S604: Select the pose with the highest score as the refined pose.

8. The object grasping method based on image segmentation and object pose estimation according to claim 1, characterized in that, The specific content of S8 is as follows: Calculate the target grasping pose according to the following formula: ; Among them, represents the target grasping pose in the robotic arm coordinate system, represents the refining pose in the robotic arm coordinate system, represents the inverse operation of the predefined fixed transformation between the end effector and the object to be grasped.

9. An object grasping system based on image segmentation and object pose estimation, characterized in that, including: a processor; a memory, on which computer-readable instructions are stored, and when the computer-readable instructions are executed by the processor, the object grasping method based on image segmentation and object pose estimation according to any one of claims 1 to 8 is implemented.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the program is executed by the processor, the object grasping method based on image segmentation and object pose estimation according to any one of claims 1 to 8 is implemented.

Citation Information

Patent Citations

  • Mechanical arm vision grabbing system and method based on self-supervised-learning neural network

    CN109702741A

  • Reflective part pose estimation method based on monocular vision

    CN114612393A

  • Pose estimation method and system for weak texture object

    CN114897982A

  • Object grabbing method, system and equipment based on mechanical arm and storage medium

    CN115213896A

  • Spatial positioning precision evaluation method and system, storage medium and computer

    CN115984388A

Cited By

  • Mechanical arm grabbing method and equipment based on few-sample semantic segmentation model and medium

    CN121105025A

  • Method, device and equipment for controlling mechanical arm and storage medium

    CN121315988A

  • Object grasping method and device, electronic equipment and storage medium

    CN122584361A