Seven-Degree-of-Freedom Grasping Pose Detection Method Based on RGB Images and Depth Images

By combining RGB images and depth images, the problem of point cloud data instability and lack of targeted grabbing is solved, and a stable and accurate seven-degree-of-freedom grabbing posture is generated, which improves the robustness of grabbing detection.

CN114140418BActive Publication Date: 2025-07-11NINGBO ARTIFICIAL INTELLIGENCE RES INST OF SHANGHAI JIAOTONG UNIV
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202111418398.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-11-26
Publication Date
2025-07-11
Estimated Expiration
2041-11-26

AI Technical Summary

Technical Problem

There are problems in the existing point cloud-based six-degree-of-freedom crawling detection method and the lack of targeted crawling generated by point cloud data instability and lack of targeted crawling posture accuracy and stability.

Method used

The seven-degree-of-freedom grasping posture detection method based on RGB images and depth images is adopted. By converting the depth image into point cloud data and projecting it into two-dimensional images, combining RGB images for target segmentation and feasible grab semantic segmentation, PENet is used for depth image completion, the normal vector and main curvature direction of feasible grab points are calculated, the heuristic algorithm is used for grab depth and width sampling, and the capture candidate classification is used for grabbing, and the targeted grab posture is finally generated.

Benefits of technology

The accuracy and stability of the grab pose are improved, the instability of point cloud data and the blindness of random sampling are overcome, and a stable and targeted seven-degree-of-freedom grab pose is generated.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114140418B_ABST
    Figure CN114140418B_ABST
Patent Text Reader

Abstract

The present invention discloses a seven - degree - of - freedom grasping pose detection method based on RGB images and depth images, which relates to the field of computer vision and includes: Step 1, converting the depth image into point cloud data and projecting it to obtain a three - channel image, i.e., the X - Y - Z image; Step 2, using ResNet - 50 to encode the information of the RGB image and the X - Y - Z image to obtain the target segmentation result and the feasible grasping semantic segmentation result; Step 3, completing the depth image to obtain a dense point cloud; Step 4, using the feasible grasping points and the dense point cloud to calculate the normal vector and two principal curvature directions of the feasible grasping points to form a grasping coordinate system; Step 5, sampling the grasping depth and grasping width of the feasible grasping points to generate a number of grasping candidates, and each grasping candidate corresponds to a grasping closed region; Step 6, inputting the points within the grasping closed region into PointNet to filter the grasping candidates to obtain the final set of grasping poses; Step 7, projecting the grasping candidates onto the target to generate the final grasping pose.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer vision, and particularly to a method for detecting a seven-degree-of-freedom grasping posture based on RGB images and depth images. Background Art

[0002] Robust robotic arm grasping is a basic requirement for robots in most industrial scenarios and daily life. The grasping of the entire robotic arm is divided into two parts: grasping detection and path planning. Grasping detection refers to generating the six-degree-of-freedom posture that the end of the robotic arm needs to reach by using scene information obtained from sensors such as monocular cameras, depth cameras, or binocular cameras. The six-degree-of-freedom posture refers to the position and coordinate system that the center of the end of the robotic arm needs to reach. Path planning refers to how to plan the movement path of the robotic arm in the workspace for the six-degree-of-freedom posture generated by grasping detection, so that the robotic arm does not collide with the scene and satisfies the kinematic constraints of the robotic arm.

[0003] In recent years, with the development of deep learning, vision-based robotic arm grasping detection algorithms have achieved rapid development. Vision-based robotic arm grasping detection schemes can be divided into three categories (see "Fang HS, Wang C, Gou M, etal. GraspNet-1 Billion: A Large-Scale Benchmark for General Object Grasping, CVPR 2020"). Among them, the first type of detection scheme takes the RGB images captured by monocular or multiocular camera sensors as input, and then detects a feasible grasping box in the 2D image. This grasping box contains the grasping position and an angle information representing the rotation in the plane. This type of algorithm limits the grasping to be perpendicular to the desktop, severely restricting the degrees of freedom of grasping, and making it difficult to grasp stacked objects in a cluttered scene. The second type of detection scheme is to detect the six-degree-of-freedom pose transformation of the target object, and transform the corresponding grasping of the target in the reference coordinate system to the coordinate system (see "Zhao W, Zhang S, Guan Z, etal. Learning Deep Network for Detecting 3D Object Keypoints and 6D Poses, CVPR 2020"). The problem with this type of algorithm is that it can only be used for grasping objects that already exist in the dataset. For new objects, 3D modeling needs to be carried out first, and then the grasping pose needs to be manually annotated, which will result in too high a cost for obtaining the dataset. The third type of detection scheme takes point cloud data as input, and uses the geometric and semantic information of the point cloud in 3D space to directly obtain the six-degree-of-freedom pose that the end of the robotic arm needs to reach through a single-stage or two-stage method (see "Liang H, Ma X, Li S, et al. PointNetGPD: Detecting Grasp Configurations from Point Sets, ICRA 2019"). The advantage of this type of method is that the trained model has good generality and can obtain unrestricted grasping postures. However, in most cases, this detection scheme only takes unstable point cloud data as input and cannot perform targeted grasping.

[0004] There are relatively few existing six-degree-of-freedom grasping pose detection solutions, and RGB data has not been applied to a solution for overcoming the instability of point cloud data and generating object-oriented grasps. In the patent application "A Robot Grasping Detection Method Based on Multi-Class Object Segmentation" (Chinese Invention Patent No. 112861667A) by Yu Xiuli et al., RGB images are used for image segmentation and semantic recognition, and a grasping rectangle containing the in-plane rotation angle is generated for the segmented object; in the patent application "A Robot Grasping Pose Estimation Method Based on an Object Recognition Deep Learning Model" (Chinese Invention Patent No. 01810803444) by Li Mingyang et al., a method of fusing two-dimensional visual information and three-dimensional visual information is used to obtain the point cloud of the target object, and then the point cloud of the target object is registered with the object point cloud template in the template library to estimate the pose of the target object; in the patent application "A Robot Grasping Detection Solution Based on Instance Segmentation under a Single-View Point Cloud" (Chinese Invention Patent No. 110363815A) by Qian Kun et al., RGB images are used for target segmentation, then the segmented target point set is mapped to the point cloud, and then an initial grasping coordinate system is generated based on the geometric structure of the original point cloud data at randomly sampled points, and finally the final six-degree-of-freedom grasping pose is generated through translation and filtering.

[0005] With the continuous development of deep learning, the role of RGB data in pose detection has gradually been explored. RGB data can predict points with specific semantic information on an image, such as key points on a human body (see "Sun K, Xiao B, Liu D, et al. Deep High-Resolution Representation Learning for Human Pose Estimation, CVPR 2019"), and can also be used to predict the grasping rotation matrix for each point on an image (see "Gou M, Fang H S, Zhu Z, et al. RGB Matters: Learning 7-DoF Grasp Poses on Monocular RGBD Images, ICRA 2021").

[0006] Therefore, for the grasping detection part, those skilled in the art are committed to developing a seven-degree-of-freedom grasping pose detection method based on RGB images and depth images, overcoming the instability of point cloud data and the defect of inability to perform targeted grasping in the point cloud-based grasping method, and improving the accuracy and stability of the generated grasping pose. Summary of the Invention

[0007] In view of the above defects of the prior art, the technical problem to be solved by the present invention is how to solve the problems of instability of point cloud data and lack of pertinence in the generated grasps in the existing six-degree-of-freedom grasp detection method based on point cloud, so as to improve the accuracy and stability of the finally generated grasp postures.

[0008] To achieve the above object, the present invention provides an improved grasp detection method using RGB data. For the grasp detection part, based on monocular RGB data and depth data, a robust seven-degree-of-freedom grasp posture for a parallel two-finger gripper is generated. Compared with the six-degree-of-freedom posture, the seven-degree-of-freedom grasp posture adds a grasp width.

[0009] A seven-degree-of-freedom grasp posture detection method based on RGB images and depth images provided by the present invention, the method includes the following steps:

[0010] Step 1: By converting the depth image into point cloud data and projecting the coordinates of the point cloud data into a two-dimensional image, a three-channel coordinate image, an X-Y-Z image, is obtained;

[0011] Step 2: Use ResNet-50 to encode the information of the RGB image and the X-Y-Z image, and then use a target segmentation decoding network and a feasible grasp semantic segmentation decoding network to decode the encoded information simultaneously, to obtain the target segmentation result and the feasible grasp semantic segmentation result of each pixel in the image, and feasible grasp points can be obtained from the feasible grasp semantic segmentation result;

[0012] Step 3: Use PENet and use the RGB image to complete the depth image, to obtain a completed dense depth image, and further obtain a dense point cloud;

[0013] Step 4: Use the feasible grasp points obtained from the feasible grasp semantic segmentation result in Step 2 and the dense point cloud obtained in Step 3 to calculate the normal vector and two principal curvature directions of the feasible grasp points, and form a grasp coordinate system;

[0014] Step 5: Use a heuristic algorithm to sample the grasp depth and grasp width of the feasible grasp points, to obtain a number of grasp candidates, the grasp depth of the grasp candidates is the largest and the grasp width is the smallest; each grasp candidate corresponds to a grasp closing region;

[0015] Step 6: Input the points in the grasp closing region of the grasp candidate into PointNet, filter out the infeasible grasp candidates, and obtain a final set of grasp postures;

[0016] Step 7: Combine the target segmentation result in Step 2, project the grasping candidates in the grasping pose set onto the target of the target segmentation result, and generate a targeted grasping pose.

[0017] Further, the gripper used in the method is a two-finger parallel gripper, and the grasping parameters of the seven degrees of freedom are expressed as: , where represents the position of the gripper in the world coordinate system, represents the rotation of the grasping coordinate system of the gripper around the x, y, and z axes of the world coordinate system, and w represents the end width of the gripper.

[0018] Further, in Step 1, the specific method of projecting the coordinates of the point cloud data onto the two-dimensional image to obtain the X-Y-Z image is as follows:

[0019]

[0020] where D represents the depth image, represents the camera internal parameters.

[0021] Further, in Step 2, use a multi-task semantic segmentation module to perform per-pixel object segmentation and feasible grasping semantic segmentation on the RGB image and the X-Y-Z image; the object segmentation is used to detect the category to which the pixel belongs, and the feasible grasping semantic segmentation is used to detect whether the pixel is suitable to be used as a grasping center;

[0022] The object segmentation decoding network and the feasible grasping semantic segmentation decoding network are composed of dense upsampling convolutional modules with different numbers of layers;

[0023] The loss function of the object segmentation decoding network uses an improved cross-entropy loss function , defined as:

[0024]

[0025] where N represents the total number of pixels in the image; represents the total number of categories; represents the weight of category c among all categories, and the calculation formula is ; represents the total number of pixels whose category truth value is c, is used to balance the situation of uneven category numbers; takes a value of 0 or 1. 0 means that the category is different from the category truth value corresponding to the pixel, and 1 means that the category is the same as the category truth value corresponding to the pixel; represents the confidence score that pixel x belongs to list c, and use To balance the difficulty levels of different samples, the loss weights of samples with higher confidence scores are made smaller, which is an adjustable parameter.

[0026] Furthermore, the feasible grasping semantic segmentation decoding network is a binary classification network, and a common cross-entropy loss function is used , and is specifically defined as:

[0027]

[0028] where N represents the total number of pixels in the image, represents the true class of pixel x, represents the measured confidence score, and is set to make the weight ratio of the loss of points with the label of graspable points larger.

[0029] Furthermore, the loss function of the multi-task semantic segmentation module is defined as:

[0030]

[0031] where and are adjustable parameters.

[0032] Furthermore, in step 3, a depth image completion module is used to complete the depth image;

[0033] The PENet algorithm adopts a dual-channel framework, and both use a deep convolutional neural network and a transposed convolutional way to construct a similar encoder-decoder network. One channel takes color information as the dominant input to obtain a color-dominated depth map; the other channel takes the original depth image as the dominant input, combines the color-dominated depth map, and obtains a depth-dominated depth map. Then, the obtained color-dominated depth map and the depth-dominated depth map are fused by a weighted method to obtain the preliminary dense depth image. Finally, DA-CSPN++ is used to refine the dense depth image to obtain the final completed dense depth image.

[0034] Furthermore, in step 4, the rotation matrix detection module is used to calculate the grasping coordinate system where the feasible grasping points are located;

[0035] Use the K-nearest neighbor algorithm to sample K nearest neighbor points near the feasible grasping point to form a point set, fit a plane closest to the point set, obtain the normal direction of the feasible grasping point according to the fitted plane, use the plane passing through the normal direction to cut the surface where the feasible grasping point is located, the surface will intersect with the plane to obtain a curve, the curve has a curvature at the feasible grasping point, and select the directions with the maximum curvature and the minimum curvature at the feasible grasping point among different curves as the two principal curvature directions of the feasible grasping point.

[0036] Further, in step 5, use the grasping depth and width detection module to determine the grasping depth and the grasping width of the feasible grasping point;

[0037] Taking the z value of the feasible grasping point as the interval center, use the heuristic algorithm to sample both sides of the interval center, and judge whether different z and w satisfy: 1) The gripper does not collide with the scene point cloud before closing; 2) The closing area of the gripper needs to contain the grasping center.

[0038] Further, in step 6 and step 7, use the grasping classification and assignment module to determine the targeted grasping posture;

[0039] Use the PointNet as an encoder to encode the information of the points in the grasping closing area of the generated grasping candidate, use the fully connected layer for classification, filter out the infeasible grasping candidates, and obtain the final set of grasping postures.

[0040] An improved grasping detection method using RGB data provided by the present invention has at least the following technical effects:

[0041] 1. Previous six - degree - of - freedom grasping detection methods solely use point - cloud data as input and obtain candidates for feasible grasping points by randomly sampling the point cloud. There are two problems: Firstly, since the point - cloud data obtained by a depth camera has a lot of noise and is sparse at some thin edges of the object, it leads to the inability to generate grasps when sampling noise points and at the thin edges of the object. Secondly, the distribution of feasible grasping points in the scene is not uniform. Therefore, the random sampling method will result in a large number of invalid operations, causing unnecessary computational expenses. Most existing six - degree - of - freedom grasping detections generate non - targeted grasps, and may even generate grasping postures in the background that meets the grasping space requirements, making it inapplicable to scenarios with high stability requirements. In the technical solution provided by the present invention, a multi - task semantic segmentation module is proposed. By combining RGB data and point - cloud data, pixel - by - pixel class labels and labels indicating whether it is graspable are obtained, which are used for subsequent targeted grasping generation and grasping posture generation respectively, overcoming the blindness of random sampling and the instability of the point cloud, and at the same time helping to form the final targeted grasp;

[0042] 2. The technical solution provided by the present invention introduces a depth - image completion algorithm into the six - degree - of - freedom grasping detection algorithm, which helps to generate a grasping coordinate system at the thin edges of the object that the sensor fails to capture, highlighting the solution to the instability problem of point - cloud data in the six - degree - of - freedom grasping detection method, and can detect stable and targeted seven - degree - of - freedom grasping postures.

[0043] The concept, specific structure and technical effects of the present invention will be further described below in conjunction with the drawings to fully understand the purpose, features and effects of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS

[0044] Figure 1 is the overall flowchart of a preferred embodiment of the present invention;

[0045] Figure 2 is Figure 1 the overall framework of the multi - task semantic segmentation module in the illustrated embodiment;

[0046] Figure 3 is Figure 1 the original scene, target segmentation, feasible - grasp semantic segmentation and generated targeted - grasp example diagrams in the illustrated embodiment. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0047] The following introduces multiple preferred embodiments of the present invention with reference to the accompanying drawings of the specification to make its technical content clearer and easier to understand. The present invention can be embodied in many different forms of embodiments, and the protection scope of the present invention is not limited to the embodiments mentioned in the text.

[0048] Existing six - degree - of - freedom grasping detection methods simply use point cloud data as input and obtain candidates for feasible grasping points by randomly sampling the point cloud data. This method has two problems: First, due to the large amount of noise in the point cloud data obtained by the depth camera and the sparsity at some thin edges of the object, it is impossible to generate grasps by sampling noise points and at the thin edges of the object. Second, the distribution of feasible grasping points in the scene is not uniform, and the random sampling method will lead to a large number of invalid operations, resulting in unnecessary computational expenses. At the same time, most existing six - degree - of - freedom grasping detections generate non - targeted grasps, and may even generate grasping postures in the background that meets the grasping space requirements, making the existing six - degree - of - freedom grasping detection methods inapplicable to scenarios with high stability requirements.

[0049] In view of the above - mentioned defects of the prior art, the technical problem to be solved by the present invention is how to solve the instability of point cloud data and the lack of pertinence of the generated grasps in the existing six - degree - of - freedom grasping detection method based on point cloud, so as to improve the accuracy and stability of the finally generated grasping postures.

[0050] To achieve the above object, the present invention provides an improved grasping detection method using RGB data. For the grasping detection part, based on monocular RGB data and depth data, a robust seven - degree - of - freedom grasping posture for a parallel two - finger gripper is generated. Compared with the six - degree - of - freedom posture, the seven - degree - of - freedom grasping posture adds the grasping width. Specifically, first, the RGB image and depth image of the scene are obtained, processed, and encoded using ResNet - 50. Then, through two decoding networks, target segmentation and feasible grasping semantic segmentation are performed. At the same time, the depth image is complemented using the RGB image to obtain a dense point cloud. The normal vector and two principal curvature directions of the feasible grasping points are generated in the dense point cloud as the grasping coordinate system. Then, a heuristic method is used to sample the depth and width of the grasp, and the grasping posture with the maximum grasping depth and the minimum grasping width is retained as the seven - degree - of - freedom grasping candidate. Finally, PointNet is used to classify the grasping candidates, filter out infeasible grasps, obtain the final feasible seven - degree - of - freedom grasping posture, and project it onto the corresponding target according to the grasping center to obtain a targeted grasping posture.

[0051] As Figure 1 shown, after obtaining the RGB image and depth image in advance, the method includes the following steps:

[0052] Step 1: Convert the depth image into point cloud data, and project the coordinates of the point cloud data into a two - dimensional image to obtain a three - channel coordinate image, an X - Y - Z image;

[0053] Step 2: Use ResNet-50 to encode the information of the RGB image and the X-Y-Z image, and then use the target segmentation decoding network and the feasible grasping semantic segmentation decoding network to decode the encoded information simultaneously to obtain the target segmentation result and the feasible grasping semantic segmentation result of each pixel in the image. The feasible grasping points can be obtained from the feasible grasping semantic segmentation result;

[0054] Step 3: Use PENet and the RGB image to complete the depth image, obtain the completed dense depth image, and further obtain the dense point cloud;

[0055] Step 4: Use the feasible grasping points obtained from the feasible grasping semantic segmentation result in Step 2 and the dense point cloud obtained in Step 3 to calculate the normal vector and two principal curvature directions of the feasible grasping points to form a grasping coordinate system;

[0056] Step 5: Use a heuristic algorithm to sample the grasping depth and grasping width of the feasible grasping points to obtain several grasping candidates. The grasping depth of the grasping candidate is the maximum and the grasping width is the minimum; each grasping candidate corresponds to a grasping closed area;

[0057] Step 6: Input the points in the grasping closed area of the grasping candidate into PointNet, filter out the infeasible grasping candidates, and obtain the final set of grasping postures;

[0058] Step 7: Combine the target segmentation result in Step 2, project the grasping candidates in the set of grasping postures onto the target of the target segmentation result, and generate targeted grasping postures.

[0059] Specifically, the gripper used in this method is a two-finger parallel gripper, and the grasping parameters with seven degrees of freedom are expressed as: , where represents the position of the gripper in the world coordinate system, represents the rotation of the grasping coordinate system of the gripper around the x, y, and z axes of the world coordinate system, and w represents the end width of the gripper.

[0060] Since the RGB image has rich semantic information and texture information in the two-dimensional space, and the point cloud data has semantic information and spatial geometric information in the three-dimensional space, the two are combined for the semantic segmentation task.

[0061] In Step 1, the specific method of projecting the coordinates of the point cloud data onto the two-dimensional image to obtain the X-Y-Z image is as follows:

[0062]

[0063] where D represents the depth image, represents the camera internal parameters.

[0064] Then, the obtained six-channel image (R, G, B, X, Y, Z) is used as the input image for Step 2.

[0065] In Step 2, a multi-task semantic segmentation module is used to perform per-pixel object segmentation and feasible grasping semantic segmentation on the RGB image and the X-Y-Z image (i.e., the six-channel image (R, G, B, X, Y, Z)); the object segmentation is used to detect the category to which the pixel belongs, and the feasible grasping semantic segmentation is used to detect whether the pixel is suitable to be used as a grasping center.

[0066] The object segmentation decoding network and the feasible grasping semantic segmentation decoding network are composed of dense upsampling convolutional modules with different numbers of layers;

[0067] The loss function of the object segmentation decoding network uses an improved cross-entropy loss function , which is defined as:

[0068]

[0069] where N represents the total number of pixels in the image; represents the total number of categories; represents the weight of category c among all categories, and the calculation formula is ; represents the total number of pixels whose category truth value is c, which is used to balance the situation of uneven category numbers; takes a value of 0 or 1. 0 means that the category is different from the category truth value corresponding to the pixel, and 1 means that the category is the same as the category truth value corresponding to the pixel; represents the confidence score that pixel x belongs to list c, and is used to balance the difficulty levels of different samples, and the loss weight of samples with higher confidence scores is made smaller, is an adjustable parameter.

[0070] The feasible grasping semantic segmentation decoding network is a binary classification network and uses a common cross-entropy loss function , which is specifically defined as:

[0071]

[0072] where N represents the total number of pixels in the image, represents the category truth value of pixel x, represents the measured confidence score, and is set to make the proportion of the loss of points with the label of graspable points larger.

[0073] The loss function of the multi-task semantic segmentation module is defined as:

[0074]

[0075] Among them, and are adjustable parameters.

[0076] The training data of the multi-task semantic segmentation module comes from the GraspNet-1Billion dataset. This dataset contains object segmentation labels. To obtain feasible grasping semantic segmentation labels, the grasping centers of the 6DOF grasping poses in this dataset can be projected onto the 2D image.

[0077] As Figure 2 shown, it is the overall framework of the multi-task semantic segmentation module in step 2. In this step, the multi-task semantic segmentation module is used to perform per-pixel object segmentation and feasible grasping semantic segmentation on the RGB image and the X-Y-Z image. Among them, object segmentation is used to detect the category to which the pixel belongs, and feasible grasping semantic segmentation is used to detect whether the pixel is suitable to be used as a grasping center. Specifically, ResNet-50 is used to encode the information of the RGB image and the X-Y-Z image, and then the object segmentation decoding network and the feasible grasping semantic segmentation decoding network obtain the object segmentation result and the feasible grasping semantic segmentation result of each pixel by decoding the encoded information. Both the object segmentation decoding network and the feasible grasping semantic segmentation decoding network adopt a dense upsampling convolutional network, but the number of layers is different.

[0078] Specifically, in step 3, the depth image completion module is used to complete the depth image.

[0079] The PENet algorithm adopts a two-channel framework, and both use a deep convolutional neural network and a transposed convolutional way to construct a similar encoder-decoder network. One channel takes the color information as the dominant input to obtain a color-dominated depth map; the other channel takes the original depth image as the dominant input, combines the color-dominated depth map, and obtains a depth-dominated depth map. Then, the obtained color-dominated depth map and depth-dominated depth map are fused by a weighted method to obtain a preliminary dense depth image. Finally, DA-CSPN++ is used to refine the dense depth image to obtain the final completed dense depth image.

[0080] Specifically, in step 4, the rotation matrix detection module is used to calculate the grasping coordinate system where the feasible grasping points are located.

[0081] Use the K-nearest Neighbor (KNN) algorithm to sample K nearest neighbor points near the feasible grasping point to form a point set, fit a plane closest to the point set, obtain the normal direction of the feasible grasping point based on the fitted plane, use the plane passing through the normal direction to cut the surface where the feasible grasping point is located, and the surface will intersect with the plane to obtain a curve. The curve has a curvature at the feasible grasping point. Select the directions with the maximum curvature and the minimum curvature at the feasible grasping point among different curves as the two principal curvature directions of the feasible grasping point.

[0082] Specifically, in step 5, a grasping depth and width detection module is used to determine the grasping depth and width of the feasible grasping point.

[0083] Taking the z value of the feasible grasping point as the center of the interval, use a heuristic algorithm to sample on both sides of the interval center, and judge whether different z and w satisfy: 1) The gripper does not collide with the scene point cloud before closing; 2) The closing area of the gripper needs to contain the grasping center.

[0084] Specifically, in steps 6 and 7, a grasping classification and assignment module is used to determine the targeted grasping posture.

[0085] Use PointNet as the encoder to encode the information of the points in the grasping closing area of the generated grasping candidates, use a fully connected layer for classification, filter out infeasible grasping candidates, and obtain the final set of grasping postures.

[0086] As Figure 3 shown, for Figure 1 the original scene, target segmentation, feasible grasping semantic segmentation, and generated targeted grasping example diagrams in the method shown, where Figure 3 the upper left figure is the original scene, the upper right figure is the target segmentation, the lower left figure is the feasible grasping semantic segmentation, and the lower right figure is the generated targeted grasping.

[0087] The above has described in detail the preferred specific embodiments of the present invention. It should be understood that those of ordinary skill in the art can make many modifications and variations according to the concept of the present invention without creative labor. Therefore, all technical solutions that can be obtained by those skilled in the art in the technical field of the present invention based on the concept of the present invention through logical analysis, reasoning, or limited experiments on the basis of the prior art should be within the protection scope determined by the claims.

Claims

1. A seven-degree-of-freedom grasping posture detection method based on RGB images and depth images, characterized in that The method includes the following steps: Step 1: Convert the depth image into point cloud data, and project the coordinates of the point cloud data into a two-dimensional image to obtain a three-channel coordinate image, an X-Y-Z image; Step 2: Use ResNet-50 to encode the information of the RGB image and the X-Y-Z image, and then use a target segmentation decoding network and a feasible grasping semantic segmentation decoding network to simultaneously decode the encoded information to obtain the target segmentation result and the feasible grasping semantic segmentation result of each pixel in the image. Feasible grasping points can be obtained from the feasible grasping semantic segmentation result; Step 3: Use PENet and the RGB image to complete the depth image to obtain a completed dense depth image, and further obtain a dense point cloud; Step 4: Use the feasible grasping points obtained from the feasible grasping semantic segmentation result in Step 2 and the dense point cloud obtained in Step 3 to calculate the normal vector and two principal curvature directions of the feasible grasping points to form a grasping coordinate system; Step 5: Use a heuristic algorithm to sample the grasping depth and grasping width of the feasible grasping points to obtain a number of grasping candidates, where the grasping depth of the grasping candidates is the maximum and the grasping width is the minimum; each grasping candidate corresponds to a grasping closed area; Step 6: Input the points in the grasping closed area of the grasping candidate into PointNet, filter out infeasible grasping candidates, and obtain a final set of grasping postures; Step 7: Combine the target segmentation result in Step 2, and project the grasping candidates in the set of grasping postures onto the target of the target segmentation result to generate targeted grasping postures; In Step 2, use a multi-task semantic segmentation module to perform per-pixel target segmentation and feasible grasping semantic segmentation on the RGB image and the X-Y-Z image; the target segmentation is used to detect the category to which the pixel belongs, and the feasible grasping semantic segmentation is used to detect whether the pixel is suitable to be used as a grasping center; The target segmentation decoding network and the feasible grasping semantic segmentation decoding network are composed of dense upsampling convolutional modules with different numbers of layers; The loss function of the target segmentation decoding network uses an improved cross-entropy loss function , which is defined as: Among them, N represents the total number of pixels in the image; represents the total number of categories; represents the weight of category c among all categories, and the calculation formula is ; represents the total number of pixels whose category truth value is c, which is used to balance the situation of uneven category numbers; takes a value of 0 or 1. 0 means that the category is different from the category truth value corresponding to this pixel, and 1 means that the category is the same as the category truth value corresponding to this pixel; represents the confidence score that pixel x belongs to list c, and is used to balance the difficulty levels of different samples, reducing the loss weight of samples with higher confidence scores, is an adjustable parameter; The feasible grasping semantic segmentation decoding network is a binary classification network that uses a common cross-entropy loss function , and is specifically defined as: where N represents the total number of pixels in the image, represents the ground truth of the category of pixel x, represents the measured confidence score, and is set to make the weight ratio of the loss where the label is a graspable point larger; The loss function of the multi-task semantic segmentation module is defined as: Among them, and are adjustable parameters; PENet adopts a two-channel framework, and both use a deep convolutional neural network and a transposed convolution method to construct a similar encoder-decoder network. One channel takes color information as the dominant input to obtain a color-dominated depth map; the other channel takes the original depth image as the dominant input, combines the color-dominated depth map, and obtains a depth-dominated depth map. Then, the obtained color-dominated depth map and depth-dominated depth map are fused by weighting to obtain a preliminary dense depth image. Finally, use DA-CSPN++ to refine the dense depth image to obtain the final completed dense depth image; In Step 4, use a rotation matrix detection module to calculate the grasping coordinate system where the feasible grasping points are located; Sample K nearest neighbor points near the feasible grasping point using the K-nearest neighbor algorithm to form a point set, fit a plane closest to the point set, obtain the normal direction of the feasible grasping point according to the fitted plane, use the plane passing through the normal direction to cut the surface where the feasible grasping point is located, and the surface will intersect with the plane to obtain a curve. The curve has a curvature at the feasible grasping point. Select the directions with the maximum curvature and the minimum curvature at the feasible grasping point among different curves as the two principal curvature directions of the feasible grasping point.

2. The seven-degree-of-freedom grasping pose detection method based on RGB images and depth images according to claim 1, wherein, The gripper used in the method is a two-finger parallel gripper, and the grasping parameters of the seven degrees of freedom are expressed as: , where represents the position of the gripper in the world coordinate system, represents the rotation of the grasping coordinate system of the gripper around the x, y, and z axes of the world coordinate system, and w represents the end width of the gripper.

3. The seven-degree-of-freedom grasping pose detection method based on RGB images and depth images according to claim 1, characterized in that In step 1, the specific method of projecting the coordinates of the point cloud data onto the two-dimensional image to obtain the X-Y-Z image is as follows: where D represents the depth image, represents the camera intrinsic parameters.

4. The seven-degree-of-freedom grasping pose detection method based on RGB images and depth images according to claim 1, characterized in that, In step 3, use the depth image completion module to complete the depth image.

5. The seven-degree-of-freedom grasping pose detection method based on RGB images and depth images according to claim 2, wherein, In step 5, use the grasping depth and width detection module to determine the grasping depth and the grasping width of the feasible grasping point; Taking the z value of the feasible grasping point as the interval center, use the heuristic algorithm to sample both sides of the interval center, and judge whether different z and w satisfy: 1) the gripper does not collide with the scene point cloud before closing; 2) the closing area of the gripper needs to contain the grasping center.

6. The seven-degree-of-freedom grasping pose detection method based on RGB images and depth images according to claim 1, wherein In steps 6 and 7, use the grasping classification and assignment module to determine the targeted grasping posture; Use the PointNet as the encoder to encode the information of the points in the grasping closing area of the generated grasping candidate, use the fully connected layer for classification, filter out the infeasible grasping candidates, and obtain the final set of grasping postures.

Citation Information

Patent Citations

  • Robot grabbing detection method based on instance segmentation under single-view-angle point cloud

    CN110363815A

  • Robot grabbing detection method based on multi-category target segmentation

    CN112861667A

  • Method and device and equipment for generating grasping trajectory of mechanical arm and storage medium

    CN110026987A

  • Construction method and device of three-dimensional map, robot and readable storage medium

    CN110298873A