6-DoF capture detection method based on depth image
Through the 6-DoF grab detection method based on depth images, the image and point cloud features fusion are used to generate enhanced point features, and the local features of the grab points are calculated through multiple modules, the problems of point cloud sparseness and local feature aggregation redundancy in the existing technology are solved, and more accurate grab prediction is achieved.
Patent Information
- Application Number
- CN202510067744.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-16
- Publication Date
- 2025-05-13
AI Technical Summary
The existing 6-DoF crawling detection method has sparse point clouds due to downsampling, which cannot fully express the local geometric details of the crawling point. The local feature aggregation method has redundant global geometric information, which affects the network's attention to local geometric information.
Using a 6-DoF grab and detection method based on depth images, pixel-by-pixel features are extracted through the image backbone network, point features are extracted in combination with point cloud backbone network, and pixel features are fused through position index to generate enhanced point features. Then, the local features of each grab point are calculated through the grab point sampling, proximity direction prediction and local feature aggregation module, and finally the grab width and grab score are predicted using the grab pose prediction module.
It improves the accuracy of the local geometric information of the crawling point, enhances the network's attention to local geometric information, and improves the reliability of crawling prediction.
Smart Images

Figure CN119991805A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of machine learning and computer vision, and aims to predict the 6-DoF grasping posture of an object, which can be used to support robot grasping interaction tasks in real scenes. Background Art
[0002] 6-DoF grasping posture detection aims to output a set of grasping posture parameters (including grasping point coordinates, rotational posture, and grasping width) of objects in the scene. The grasping point coordinates indicate the position of the gripper, the rotational posture indicates the rotational posture of the gripper when grasping the object, and the grasping width indicates the width of the gripper opening. Therefore, the prediction accuracy of the grasping detection model is very important for the reliable grasping of the robot in real scenes.
[0003] In order to reduce the computational burden, existing 6-DoF grasp detection methods downsample the original point cloud (such as farthest point sampling) to significantly reduce the number of points in the point cloud, which results in sparse neighborhood points of the grasp point, and thus the grasp point features learned by the model cannot fully express the local geometric details of the area where the grasp point is located. The resolution of the depth image is much higher than the resolution of the downsampled point cloud. The dense pixel learning features based on the depth image can make up for the lack of local geometric details of the features learned based on the sparse point cloud.
[0004] In addition, in the existing 6-DoF grasping pose detection network, local feature aggregation is usually performed based on neighborhood point features using MLP (multi-layer perceptron) and maximum pooling layers to obtain local geometric features of the grasping point. Since the neighborhood point features are obtained by downsampling encoding and upsampling decoding of the entire scene through the backbone network, the neighborhood point features used for local feature aggregation all contain global geometric information of the entire scene. Therefore, the existing local feature aggregation method is based on the premise that both the grasping point and its neighborhood points contain global geometric information. The redundant global geometric information may overwhelm the local geometric information, thereby affecting the network's attention to local geometric information.
[0005] In summary, the local geometric information of the grasping points obtained by the existing methods is not accurate enough, and the accuracy of the local geometric information has an impact on the reliability of grasping prediction. Summary of the invention
[0006] In order to solve the problem that the existing 6-DoF grasping detection network uses downsampled point clouds to predict grasping posture, resulting in the grasping point features learned by the network cannot fully express the local geometric details, and the existing local feature aggregation method cannot fully focus on the local geometric information of the grasping point and its neighborhood points, the present invention proposes a new 6-DoF grasping detection method based on depth images.
[0007] The technical solution adopted by the present invention is a 6-DoF grasping detection method based on a depth image. First, a depth camera acquires a single-view depth image of a scene. The depth image is subjected to pixel-by-pixel feature extraction through an image backbone network, and the original point cloud is reconstructed using camera internal parameters and the depth image. Next, the original point cloud is subjected to point cloud downsampling to obtain a sampled point cloud, and the position of the pixel in the depth image corresponding to each point in the sampled point cloud, i.e., the position index, is recorded. Subsequently, the obtained sampled point cloud is input into the point cloud backbone network to obtain point features. Next, according to the position index, the corresponding pixel features are retrieved for each point in the sampled point cloud, and the point features are fused with the retrieved pixel features to obtain enhanced point features. Then, the grasping point sampling module determines the graspable points based on the enhanced point features, i.e., determines the grasping point set; the approach direction prediction module determines the approach direction of each grasping point based on the enhanced point features; and the local feature aggregation module calculates the local features of each grasping point based on the offset features. Finally, the grasping posture prediction module uses the local features of the grasping point to predict the grasping width and grasping score.
[0008] The specific implementation steps are as follows: First, the grasping posture is defined as G = (p,u,r,d,w,q), is the grab point coordinate, is the direction vector of the approach direction of the grasping point, r is the rotation angle in the plane with u as the normal vector, d is the grasping depth, w is the grasping width, and q is the grasping score that characterizes the quality of the grasping posture.
[0009] Step 1), point cloud reconstruction and point cloud downsampling;
[0010] Point cloud reconstruction: for depth images The pixel p in d (x d ,y d ), where 1≤x d ≤H and 1≤y d ≤W, H is the height of the depth image, and W is the width of the depth image. When the camera optical center coordinates (c x ,c y ) and focal length (f x ,f y ), the coordinates (x) of the 3D point corresponding to the pixel in the depth image are calculated by formula (1) s ,y s ,z s ):
[0011]
[0012] where z d is the pixel p d (x d ,y d) is calculated by calculating the depth image I d The 3D points corresponding to all pixels in the original point cloud are obtained Arrange the pixels of the depth image in rows, and calculate the 3D point in the original point cloud P for the kth pixel in the pixel sequence. raw The index in is k, where 1≤k≤HW.
[0013] Point cloud downsampling: In order to improve computational efficiency, raw Downsample to produce a point cloud containing N points The downsampling process is as follows: First, N random numbers RN = {r n |1≤n≤N,1≤r n ≤HW}, then from P raw Select all the indexes r n points to form a new point cloud P v For point Contains two indexes, one is the point in the sampling point cloud P v The index i in the original point cloud P raw The index r in i For the original point cloud with index r i point Use formula (2) to calculate the coordinates of the corresponding pixel in the depth image
[0014]
[0015] Where Qu(·) represents the quotient of the division, and Re(·) represents the remainder of the division. v The points in the image are used to calculate their corresponding pixel coordinates and use the position index Record the coordinates of the pixel corresponding to each point, Step 2), calculation of enhanced point features;
[0016] Image backbone network and point cloud backbone network: For depth image I d , using the image backbone network to extract pixel features For the point cloud P v , use the point cloud backbone network to extract point features The image backbone network uses the SE-ResUNet feature extraction network, and the point cloud backbone network uses the ResUNet14 feature extraction network built based on MinkowskiEngine.
[0017] Pixel feature retrieval and feature fusion: Using the position index IM, point cloud P v Each point in Find the pixel in the depth image that corresponds to it Get with Corresponding pixel features Arrange point cloud P v The pixel features of all points in And compare it with the point feature F point Splicing in the channel dimension to obtain enhanced point features
[0018]
[0019] Among them, Concate means to With F point Concatenate in the channel dimension.
[0020] Step 3), grab point sampling;
[0021] The grasping point sampling module consists of a layer of MLP (multi-layer perceptron) that uses the enhanced point feature F fusion P v The grasping degree score and foreground / background attributes are predicted for each point in . FPS sampling is performed based on all foreground points with a grasping degree higher than 0.1 to obtain M grasping points. Based on F fusion Get the features of M grasping points
[0022] Step 4), approach direction prediction;
[0023] The approach direction prediction module consists of two layers of MLP. This module uses the grasp point feature F seed Predict the approach direction of the grasping point. Each grasping point is set to 300 approach directions. The approach direction prediction module predicts the confidence scores of the 300 approach directions. The approach direction corresponding to the maximum confidence score is the optimal approach direction of the grasping point. The optimal approach direction of all grasping points is recorded as 3 represents the dimension of the direction vector. In addition, the hidden layer features of the first layer MLP output close to the direction prediction module are combined with F seed Add together to get the new features of M grasping points
[0024] Step 5), local feature aggregation;
[0025] In the process of local feature aggregation, the input is the new feature F of the grasping point s ' eed And M grasping points P seedFor each grasp point, a cylinder is constructed based on the approach direction predicted in step 4) and a predefined radius. K points are sampled from the cylinder and are considered as the neighborhood points of the grasp point. The features of the K neighborhood points of each grasp point are arranged to obtain the grouped features of the M grasp points.
[0026] The new features F of the M grasping points s ' eed The dimension is adjusted to M×1×640, and K copies are made to obtain Then calculate the offset feature F offset , the calculation process is shown in formula (4):
[0027]
[0028] in represents element-by-element subtraction, MLP represents multi-layer perceptron, and Maxpool represents Max Pooling, i.e., the maximum pooling layer. F offset and F′ seed After splicing along the channel dimension, the local features of the grasping point are obtained by MLP calculation.
[0029] F local =MLP(Concate(F offset ,F′ seed )) (5)
[0030] Among them, Concate means concatenation along the channel dimension, F local Used for subsequent grasping pose prediction.
[0031] Step 6), grasp pose prediction;
[0032] The grasping pose prediction module consists of a layer of MLP. When performing grasping pose prediction, for each grasping point, along its approach direction, four grasping depths {0.01m, 0.02m, 0.03m, 0.04m} and 12 rotation angles are set, that is, 48 grasping pose anchor frames are set for each grasping point. The grasping pose prediction module uses the grasping point local feature F local Predict the grasp width and grasp score of each grasp pose anchor box.
[0033] The loss function is designed as follows:
[0034] The model updates parameters through back propagation, and the total loss of the crawl detection network is:
[0035] L=λ1L o +λ2L p +λ3L u+λ4L s +λ5L w (6)
[0036] Wherein λ1, λ2, λ3, λ4, λ5 are hyperparameters. In the present invention, λ1=1, λ2=20, λ3=100, λ4=50, λ5=15. o is the binary cross entropy loss for foreground / background segmentation at the grasp point sampling stage, L p is the regression loss of the grasping degree score output by the grasping point sampling module, L u is the regression loss of the approaching direction output by the approaching direction prediction module, L s is the regression loss of the grasp score output by the grasp pose prediction module, L w It is the regression loss of the grasping width output by the grasping posture prediction module. The above regression loss adopts Smooth L1 The loss function form is:
[0037]
[0038] The reasoning process of the crawl detection network is as follows:
[0039] The input of the grasp detection network is a single-view depth image, and the final output of the network is a set of grasp poses:
[0040] G scene = {G m |1≤m≤M} (8)
[0041] G m =(p m ,u m ,r m ,d m ,w m ,q m ) represents the grasping posture of the mth grasping point. m 、u m 、r m d m 、w m and q m They represent the three-dimensional coordinates of the grasping point, approach direction, rotation angle, grasping depth, grasping width and grasping score respectively.
[0042] When conducting a real robot grasping experiment, take the grasping posture G with the highest grasping score in the entire scene best Execute the crawling task, where:
[0043]
[0044] Compared with the prior art, the present invention uses the GraspNet-1Billion dataset to train and evaluate the model. The dataset contains 97,280 RGBD images from the real world, and the RGBD images are taken using realsense and kinect cameras. The dataset contains 100 training scenes and 90 test scenes, each scene is shot from 256 different perspectives, and all scenes have more than one billion annotated grasping poses. The test scenes are divided into three categories: "Seen", "Similar" and "Novel" according to the types of objects in the scene, which enables the dataset to support the evaluation of the model's generalization ability. The experimental evaluation indicator Average Precision AP is AP μ The average value of friction coefficient μ∈{0.2, 0.4, 0.6, 0.8, 1.0, 1.2}. AP μ The calculation method is: given the friction coefficient μ, calculate precision@T, which represents the accuracy of the grasping posture corresponding to the first T grasping scores (the ratio of the number of valid grasping postures to T), AP μ The precision@T is averaged when T is 1 to 50. The smaller the μ value, the more difficult the grasping task is, and the higher the accuracy requirement for the grasping posture predicted by the model. The force closure analysis algorithm is used to evaluate the effectiveness of the grasping posture. BRIEF DESCRIPTION OF THE DRAWINGS
[0045] Figure 1 6-DoF grasp detection method based on depth image.
[0046] Figure 26-Schematic diagram of DoF grasping posture.
[0047] Figure 3 Local feature aggregation module based on offset features.
[0048] Figure 4 Objects and their numbers used in the real-scene experiment. DETAILED DESCRIPTION
[0049] The present invention is described in detail below with reference to the accompanying drawings and embodiments.
[0050] The 6-DoF grasping detection method based on depth image proposed in the present invention is as follows Figure 1As shown in the figure. First, the depth camera acquires a single-view depth image of the scene. The depth image is pixel-by-pixel feature extracted through the image backbone network, and the original point cloud is reconstructed using the camera intrinsic parameters and the depth image. Next, the original point cloud is downsampled to obtain a sampled point cloud, and the position of the pixel in the depth image corresponding to each point in the sampled point cloud is recorded, that is, the position index. Subsequently, the obtained sampled point cloud is input into the point cloud backbone network to obtain point features. Next, according to the position index, the corresponding pixel features are retrieved for each point in the sampled point cloud, and the point features are fused with the retrieved pixel features to obtain enhanced point features. Then, the grasping point sampling module determines the graspable points based on the enhanced point features, that is, determines the grasping point set; the approach direction prediction module determines the approach direction of each grasping point based on the enhanced point features; and the local feature aggregation module calculates the local features of each grasping point based on the offset features. Finally, the grasping posture prediction module uses the local features of the grasping point to predict the grasping width and grasping score.
[0051] The grasping posture is defined as G = (p,u,r,d,w,q), such as Figure 2 As shown, is the grab point coordinate, is the direction vector of the approach direction of the grasping point, r is the rotation angle in the plane with u as the normal vector, d is the grasping depth, w is the grasping width, and q is the grasping score that characterizes the quality of the grasping posture.
[0052] Step 1), point cloud reconstruction and point cloud downsampling;
[0053] Point cloud reconstruction: for depth images The pixel p in d (x d ,y d ), where 1≤x d ≤H and 1≤y d ≤W, H is the height of the depth image, and W is the width of the depth image. When the camera optical center coordinates (c x ,c y ) and focal length (f x ,f y ), the coordinates of the 3D point corresponding to the pixel in the depth image (x s ,y s ,z s ):
[0054]
[0055] where z d is the pixel p d (x d ,y d ) is calculated by calculating the depth image I dThe 3D points corresponding to all pixels in the original point cloud can be obtained Arrange the pixels of the depth image in rows, and calculate the 3D point in the original point cloud P for the kth pixel in the pixel sequence. raw The index in is k, where 1≤k≤HW.
[0056] Point cloud downsampling: In order to improve computational efficiency, raw Downsample to produce a point cloud containing N points The downsampling process is as follows: First, N random numbers RN = {r n |1≤n≤N,1≤r n ≤HW}, then from P raw Select all the indexes r n points to form a new point cloud P v For point The point contains two indexes, one is the point in the sampling point cloud P v The index i in the original point cloud P raw The index r in i For the original point cloud with index r i point Use formula 2 to calculate the coordinates of the corresponding pixel in the depth image
[0057]
[0058] Where Qu(·) represents the quotient of the division, and Re(·) represents the remainder of the division. v The points in the image are used to calculate their corresponding pixel coordinates and use the position index Record the coordinates of the pixel corresponding to each point, Step 2), calculation of enhanced point features;
[0059] Image backbone network and point cloud backbone network: For depth image I d , using the image backbone network to extract pixel features For the point cloud P v , use the point cloud backbone network to extract point features The image backbone network uses the SE-ResUNet feature extraction network, and the point cloud backbone network uses the ResUNet14 feature extraction network built based on MinkowskiEngine.
[0060] Pixel feature retrieval and feature fusion: Using the position index IM, point cloud P v Each point in Find the pixel in the depth image that corresponds to it You can get Corresponding pixel features Arrange point cloud P v The pixel features of all points in And compare it with the point feature F point Splicing in the channel dimension to obtain enhanced point features
[0061]
[0062] Among them, Concate means to With F point Concatenate in the channel dimension.
[0063] Step 3), grab point sampling;
[0064] The grasping point sampling module consists of a layer of MLP (multi-layer perceptron) that uses the enhanced point feature F fusion P v The grasping degree score and foreground / background attributes are predicted for each point in . FPS sampling is performed based on all foreground points with a grasping degree higher than 0.1 to obtain M grasping points. Based on F fusion Get the features of M grasping points
[0065] Step 4), approach direction prediction;
[0066] The approach direction prediction module consists of two layers of MLP. This module uses the grasp point feature F seed Predict the approach direction of the grasping point. Each grasping point is set to 300 approach directions. The approach direction prediction module predicts the confidence scores of the 300 approach directions. The approach direction corresponding to the maximum confidence score is the optimal approach direction of the grasping point. The optimal approach direction of all grasping points is recorded as 3 represents the dimension of the direction vector. In addition, the hidden layer features of the first layer MLP output close to the direction prediction module are combined with F seed Add together to get the new features of M grasping points
[0067] Step 5), local feature aggregation;
[0068] The local feature aggregation process is as follows Figure 3 As shown, its input is the new feature F′ of the grasping point seed And M grasping points P seed For each grasp point, a cylinder is constructed based on the approach direction predicted in step 4) and a predefined radius. K points are sampled from the cylinder and are considered as the neighborhood points of the grasp point. Arranging the features of the K neighborhood points of each grasp point can obtain the grouping features of the M grasp points.
[0069] The new features F′ of M grasping points seed The dimension is adjusted to M×1×640, and K copies are made to obtain Then calculate the offset feature F offset , the calculation process is shown in formula 4:
[0070]
[0071] in represents element-by-element subtraction, MLP represents multi-layer perceptron, and Maxpool represents Max Pooling, i.e., the maximum pooling layer. F offset and F′ seed After splicing along the channel dimension, the local features of the grasping point are obtained by MLP calculation.
[0072] F local =MLP(Concate(F offset ,F′ seed )) (5)
[0073] Among them, Concate means concatenation along the channel dimension, F local Used for subsequent grasping pose prediction.
[0074] Step 6), grasp pose prediction;
[0075] The grasping pose prediction module consists of a layer of MLP. When performing grasping pose prediction, for each grasping point, along its approach direction, four grasping depths {0.01m, 0.02m, 0.03m, 0.04m} and 12 rotation angles are set, that is, 48 grasping pose anchor frames are set for each grasping point. The grasping pose prediction module uses the grasping point local feature F local Predict the grasp width and grasp score of each grasp pose anchor box.
[0076] The loss function is designed as follows:
[0077] The model updates parameters through back propagation, and the total loss of the crawl detection network is:
[0078] L=λ1L o +λ2L p +λ3L u +λ4L s +λ5L w (6)
[0079] Wherein λ1, λ2, λ3, λ4, λ5 are hyperparameters. In the present invention, λ1=1, λ2=20, λ3=100, λ4=50, λ5=15. o is the binary cross entropy loss for foreground / background segmentation at the grasp point sampling stage, L p is the regression loss of the grasping degree score output by the grasping point sampling module, L u is the regression loss of the approaching direction output by the approaching direction prediction module, L s is the regression loss of the grasp score output by the grasp pose prediction module, L w It is the regression loss of the grasping width output by the grasping posture prediction module. The above regression loss adopts Smooth L1 The loss function form is:
[0080]
[0081] The reasoning process of the crawl detection network is as follows:
[0082] The input of the grasp detection network is a single-view depth image, and the final output of the network is a set of grasp poses:
[0083] G scene = {G m |1≤m≤M} (8)
[0084] G m =(p m ,u m ,r m ,d m ,w m ,q m ) represents the grasping posture of the mth grasping point. m 、u m 、r m d m 、w m and q m They represent the three-dimensional coordinates of the grasping point, approach direction, rotation angle, grasping depth, grasping width and grasping score respectively.
[0085] When conducting a real robot grasping experiment, take the grasping posture G with the highest grasping score in the entire scene best Execute the crawling task, where:
[0086]
[0087] Experimental setup: downsampled point cloud P vThe number of midpoints is set to N = 20000, and the number of grasping points is set to M = 1024. The number of neighborhood points of each grasping point in the local feature aggregation stage is set to K = 16. In the local feature aggregation process, the cylinder radius is set to 0.05m, and the height range is [-0.02m, 0.04m] (with the grasping point as the origin and the approach direction as the positive direction). The network model is trained on NVIDIA's A6000 GPU for 10 cycles. The optimizer uses the Adam optimizer, with an initial learning rate of 0.001, and the learning rate decreases by 5% after each training cycle.
[0088] The AP of the present invention in the three types of scenes in the test set shot by the realsense camera reaches 75.81, 66.99 and 30.46 respectively, and the AP of the present invention in the three types of scenes in the test set shot by the kinect camera reaches 66.42, 56.17 and 22.26 respectively. The experimental results show that the present invention has a high AP value in all three types of scenes, and has excellent performance for novel scenes that are not in the training set, which shows that the present invention has good generalization for object grasping in an open environment. The experimental results are shown in Table 1, where the left side of " / " in the table is the index under the realsense camera, and the right side of " / " is the index under the kinect camera.
[0089] Table 1 AP accuracy comparison of the present invention and other methods on the GraspNet-1Billion dataset
[0090] method Seen resemblance novel PointNetGPD 25.96 / 27.59 22.68 / 24.38 9.23 / 10.66 GraspNet 27.56 / 29.88 26.11 / 27.84 10.55 / 11.51 GSNet 67.12 / 63.50 54.81 / 49.18 24.31 / 19.78 TransGrasp 39.81 / 35.97 29.32 / 29.71 13.83 / 11.41 VoteGrasp 34.1 / 37.5 33.0 / 35.9 16.9 / 18.5 P2CS-Grasp 67.90 / 64.41 57.67 / 53.79 25.41 / 21.37 Our 75.81 / 66.42 66.99 / 56.17 30.46 / 22.26
[0091] In order to further verify the reliability of the present invention in actual scenarios, the present invention uses a Tiago++ robot equipped with an RGBD camera to conduct a 6-DoF grasping experiment in a real scene. The Tiago++ robot uses an Astra S depth camera (which can obtain depth images in the range of 0.4 meters to 2 meters, and the image size is 640×480) to capture RGBD images, and uses MoveIt to plan the motion trajectory of the robotic arm, and finally uses a 7-DOF robotic arm with a gripper to perform the grasping action. The grasping experiment selected 10 objects to create 5 scenes, each containing 3-4 objects. If the robotic arm can grab an object and move it to a storage box, the grasping is considered successful; otherwise, it is considered a failure. The robot stops grasping when one of the following two conditions is met: 1) the number of attempted grasping in a single scene reaches 10 times, and 2) there is no object to grasp in the scene. The evaluation indicators of the real scene experiment are the success rate (the ratio of the number of successful grasping times to the total number of attempts) and the completion rate (the ratio of the number of objects successfully grasped in the experiment to the total number of objects). The objects used in the experiment are as follows Figure 4The experimental results are shown in Table 2. From the grasping experimental results in real scenes, it can be seen that the grasping posture predicted by the present invention in real scenes can better support the robot to perform stable grasping.
[0092] Explanation of the data calculation method in Table 2: Taking scene three as an example, there are 4 objects in this scene. After 5 attempts to grasp, all 4 objects are successfully grasped. The grasping success rate is 80% (4 / 5) and the completion rate is 100% (4 / 4).
[0093] Table 2 The grasping success rate and completion rate of the present invention in the real scene experiment
[0094]
[0095]
Claims
1. A 6-DoF grasping detection method based on depth image, characterized in that: First, the depth camera acquires a single-view depth image of the scene; the depth image is subjected to pixel-by-pixel feature extraction through the image backbone network, and the original point cloud is reconstructed using the camera intrinsic parameters and the depth image; then, the original point cloud is downsampled to obtain a sampled point cloud, and the position of the pixel in the depth image corresponding to each point in the sampled point cloud is recorded, i.e., the position index; then, the obtained sampled point cloud is input into the point cloud backbone network to obtain point features; Next, based on the position index, the corresponding pixel features are retrieved for each point in the sampling point cloud, and the point features are fused with the retrieved pixel features to obtain enhanced point features; then, the grasping point sampling module determines the graspable points based on the enhanced point features, that is, determines the grasping point set; the approach direction prediction module determines the approach direction of each grasping point based on the enhanced point features; the local feature aggregation module calculates the local features of each grasping point based on the offset features; finally, the grasping posture prediction module uses the local features of the grasping points to predict the grasping width and grasping score.
2. A 6-DoF grasping detection method based on depth image according to claim 1, characterized in that: First, the grasping posture is defined as G = (p, u, r, d, w, q), is the grab point coordinate, is the direction vector of the approach direction of the grasping point, r is the rotation angle in the plane with u as the normal vector, d is the grasping depth, w is the grasping width, and q is the grasping score that characterizes the quality of the grasping posture; Step 1), point cloud reconstruction and point cloud downsampling; Point cloud reconstruction: for depth images The pixel p in d (x d ,y d ), where 1≤x d ≤H and 1≤y d ≤W, H is the height of the depth image, W is the width of the depth image; when the camera optical center coordinates (c x , c y ) and focal length (f x , f y ), the coordinates (x) of the 3D point corresponding to the pixel in the depth image are calculated by formula (1) s ,y s , z s ): where z d is the pixel p d (x d ,y d ) is calculated by calculating the depth image I d The 3D points corresponding to all pixels in the original point cloud are obtained Arrange the pixels of the depth image in rows, and calculate the 3D point in the original point cloud P for the kth pixel in the pixel sequence. raw The index in is k, where 1≤k≤HW; Point cloud downsampling: In order to improve computational efficiency, raw Downsample to produce a point cloud containing N points The downsampling process is as follows: First, N random numbers RN = {r n |1≤n≤N,1≤r n ≤HW}, then from P raw Select all the indexes r n points to form a new point cloud P v ; For point Contains two indexes, one is the point in the sampling point cloud P v The index i in the original point cloud P raw The index r in i ; For the index r in the original point cloud i point Use formula (2) to calculate the coordinates of the corresponding pixel in the depth image Where Qu(·) represents the quotient of the division, Re(·) represents the remainder of the division; v The points in the image are used to calculate their corresponding pixel coordinates and use the position index Record the coordinates of the pixel corresponding to each point, Step 2), calculation of enhanced point features; Image backbone network and point cloud backbone network: For depth image I d , using the image backbone network to extract pixel features For the point cloud P v , use the point cloud backbone network to extract point features The image backbone network uses the SE-ResUNet feature extraction network, and the point cloud backbone network uses the ResUNet14 feature extraction network built based on MinkowskiEngine; Pixel feature retrieval and feature fusion: Using the position index IM, point cloud P v Each point in Find the pixel in the depth image that corresponds to it Get with Corresponding pixel features Arrange point cloud P v The pixel features of all points in And compare it with the point feature F point Splicing in the channel dimension to obtain enhanced point features Among them, Concate means to With F point Splicing in the channel dimension; Step 3), grab point sampling; The grasping point sampling module consists of a layer of MLP (multi-layer perceptron) that uses the enhanced point feature F fusion P v Predict the grasping degree score and foreground / background attributes for each point in ; perform FPS sampling based on all foreground points with grasping degrees higher than 0.1 to obtain M grasping points Based on F fusion Get the features of M grasping points Step 4), approach direction prediction; The approach direction prediction module consists of two layers of MLP. This module uses the grasp point feature F seed Predict the approach direction of the grasping point; set 300 approach directions for each grasping point, and the approach direction prediction module predicts the confidence scores of the 300 approach directions. The approach direction corresponding to the maximum confidence score is the optimal approach direction of the grasping point; the optimal approach direction of all grasping points is recorded as 3 represents the dimension of the direction vector; in addition, the hidden layer features of the first layer MLP output close to the direction prediction module are combined with F seed Add together to get the new features of M grasping points Step 5), local feature aggregation; In the process of local feature aggregation, the input is the new feature F′ of the grasping point seed And M grasping points P seed ; For each grasping point, a cylinder is constructed based on the approach direction predicted in step 4) and a predefined radius; K points are sampled from the cylinder, which are regarded as the neighborhood points of the grasping point; the features of the K neighborhood points of each grasping point are arranged to obtain the grouping features of the M grasping points The new features F′ of M grasping points seed The dimension is adjusted to M×1×640, and K copies are made to obtain Then calculate the offset feature F offset , the calculation process is shown in formula (4): in represents element-by-element subtraction, MLP represents multi-layer perceptron, and Maxpool represents Max Pooling, i.e., the maximum pooling layer. F offset and F′ seed After splicing along the channel dimension, the local features of the grasping point are obtained by MLP calculation. F local =MLP(Concate(F offset ,F′ seed )) (5) Among them, Concate means concatenation along the channel dimension, F local Used for subsequent grasping posture prediction; Step 6), grasp pose prediction; The grasping posture prediction module consists of a layer of MLP. When predicting the grasping posture, for each grasping point, four grasping depths {0.01m, 0.02m, 0.03m, 0.04m} and 12 rotation angles are set along its approach direction, that is, 48 grasping posture anchor frames are set for each grasping point. The grasping posture prediction module uses the local feature F local Predict the grasp width and grasp score of each grasp pose anchor box.
3. The 6-DoF grasping detection method based on depth image according to claim 1, characterized in that: The loss function is designed as follows: The model updates parameters through back propagation, and the total loss of the crawl detection network is: L=λ1L o +λ2L p +λ3L u +λ4L s +λ5L w (6) Among them, λ1, λ2, λ3, λ4, and λ5 are hyperparameters, λ1=1, λ2=20, λ3=100, λ4=50, λ5=15; L o is the binary cross entropy loss for foreground / background segmentation at the grasp point sampling stage, L p is the regression loss of the grasping degree score output by the grasping point sampling module, L u is the regression loss of the approaching direction output by the approaching direction prediction module, L s is the regression loss of the grasp score output by the grasp pose prediction module, L w It is the regression loss of the grasping width output by the grasping posture prediction module; the above regression loss adopts Smooth L1 The loss function form is:
4. The 6-DoF grasping detection method based on depth image according to claim 1, characterized in that: The reasoning process of the crawl detection network is as follows: The input of the grasp detection network is a single-view depth image, and the final output of the network is a set of grasp poses: G scene ={G m |1≤m≤M} (8) G m =(p m ,u m , r m , d m , w m ,q m ) represents the grasping posture of the mth grasping point; p m 、u m 、r m ,d m 、w m and q m They represent the three-dimensional coordinates of the grasping point, approach direction, rotation angle, grasping depth, grasping width, and grasping score respectively; When conducting a real robot grasping experiment, take the grasping posture G with the highest grasping score in the entire scene best Execute the crawling task, where: