A model-free grasping planning method and system based on deep vision
The depth vision-based method for model-free grasping addresses the challenges of complex robotic grasping by effectively distinguishing objects from backgrounds and adapting to similar colors, achieving high success rates and efficient pose generation.
Patent Information
- Application Number
- CN202310480343.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-04-28
- Publication Date
- 2025-07-15
- Estimated Expiration
- 2043-04-28
AI Technical Summary
In the prior art, the object grasping method requires a 3D model of the object as input, and it is difficult to adapt to complex grasping scenes with more objects, and the grab success rate is low when the object color is similar to the desktop.
A model-free grasping planning method based on depth vision is adopted, and the RGB image and depth image are obtained through the image acquisition module. Combined with point cloud data, a pose generation module based on image and point cloud is used to generate the grab pose, and the appropriate pose generation module is selected through the HSV color space judgment, and the robotic arm is controlled to perform the grab operation.
In the case of unknown object models, efficient crawling success rate is achieved, adapting to complex scenes and objects with similar colors to capture, improving the crawling efficiency and accuracy, and reducing dependence on the object 3D model.
Smart Images

Figure CN116277030B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of robot grasping, and more specifically, relates to a model-free grasping planning method and system based on depth vision. Background Art
[0002] Object grasping is a common means of robot operation. In the manipulator grasping task based on the object model, it is usually necessary to grasp various different types of objects in the grasping scene. The 3D model of the object needs to be manually made by professionals through certain technical means, with high costs, and it is difficult to obtain the 3D models of all objects. However, the traditional model-based grasping method requires the 3D model of the object as input and is difficult to adapt to complex grasping scenes with many objects.
[0003] In addition, the idea of grasping the target only through color segmentation is to segment through the color difference between the object and the desktop. When the object color is similar to the desktop color, it is difficult to perceive the object similar to the desktop color through the RGB image, and there are deviations when segmenting the object, resulting in grasping failure.
[0004] It can be seen that the existing technology has technical problems such as difficult model establishment, difficulty in adapting to complex grasping scenes with many objects, and low grasping success rate when the object color is similar to the desktop color. Summary of the Invention
[0005] In view of the above defects or improvement requirements of the existing technology, the present invention provides a model-free grasping planning method and system based on depth vision, thereby solving the technical problems existing in the existing technology, such as difficult model establishment, difficulty in adapting to complex grasping scenes with many objects, and low grasping success rate when the object color is similar to the desktop color.
[0006] To achieve the above object, according to one aspect of the present invention, a model-free grasping planning system based on depth vision is provided, including: an image acquisition module, a pose generation module, a processor, and a trajectory planning module;
[0007] The image acquisition module is used to acquire the RGB image, depth image, and point cloud of the object to be grasped and the desktop where it is located;
[0008] The pose generation module includes an image-based pose generation module and a point cloud-based pose generation module;
[0009] The image-based pose generation module is used to obtain the RGB image and the depth image, remove the desktop pixels in the RGB image to obtain the pixel region of the object, calculate the minimum circumscribed rectangle of the pixel region of the object to obtain the 2D pixel position of the object to be grasped in the RGB image, map the 2D pixel position to the depth image, and generate the grasping pose in the world coordinate system in combination with the camera parameter information;
[0010] A pose generation module based on point cloud is used to obtain point cloud. After downsampling the point cloud, the tabletop point cloud is removed, and then the remaining object point cloud is clustered to form an independent point cloud set. The minimum bounding rectangle box of the independent point cloud set is calculated, and the grasping pose in the world coordinate system is generated through the minimum bounding rectangle box and camera parameter information;
[0011] The processor is used to compare the HSV color space of the object to be grasped and the tabletop where the object is located. When the HSV color space of the object is within the HSV color space of the tabletop, the grasping pose generated by the pose generation module based on point cloud is input into the trajectory planning module. When the HSV color space of the object is not within the HSV color space of the tabletop, the grasping pose generated by the pose generation module based on image is input into the trajectory planning module;
[0012] The trajectory planning module is used to control the manipulator to move to the grasping pose and perform the grasping operation.
[0013] Furthermore, the pose generation module based on image includes:
[0014] The minimum circumscribed rectangle formation module is used to gray-scale the pixel region of the object to obtain the pixel point set of the object, set the initial angle so that the rectangle encloses the pixel point set of the object, rotate the rectangle, calculate the area of the rectangle at each rotation angle, and take the rectangle with the minimum area as the minimum circumscribed rectangle.
[0015] Furthermore, the pose generation module based on image further includes:
[0016] The grasping pose generation module is used to obtain the coordinates of the center point of the minimum circumscribed rectangle in the pixel coordinate system, calculate the coordinates of the two endpoints perpendicular to the two short sides of the minimum circumscribed rectangle of the center point in the pixel coordinate system, obtain the depth values of the center point and the two endpoints through the depth map. For the coordinates of the center point in the pixel coordinate system and its depth value, convert them to the coordinates of the center point in the camera coordinate system through the camera internal parameters. For the coordinates of the two endpoints in the pixel coordinate system and their depth values, convert them to the coordinates of the two endpoints in the camera coordinate system through the camera internal parameters, calculate the angle between the projection of the vector of the two endpoints on the X - O - Y plane and the X - axis, obtain the rotation angle of the grasping pose along the Z - axis of the world coordinate system, calculate the coordinates of the grasping center in the world coordinate system through the coordinates of the center point in the camera coordinate system, and the coordinates of the grasping center in the world coordinate system and the rotation angle of the grasping pose along the Z - axis of the world coordinate system form the grasping pose.
[0017] Furthermore, the pose generation module based on image further includes:
[0018] A pixel segmentation module, which is used to describe an RGB image in the RGB color space, remove the pixel region of the desktop in the RGB color space, retain the pixel region of the object in the RGB color space, or convert the RGB color space to the HSV color space, remove the pixel region of the desktop in the HSV color space, and retain the pixel region of the object in the HSV color space.
[0019] Further, the point cloud-based pose generation module includes:
[0020] A point cloud segmentation module, which is used to divide the three-dimensional space composed of the point cloud into multiple cubes, retain the center point of each cube to obtain the downsampled point cloud, randomly select N points from the downsampled point cloud as inliers, fit the inliers into an initial plane, and then traverse all the outliers outside the inliers in the downsampled point cloud. If the distance of an outlier from the initial plane is less than the threshold T, it is added to the inlier set. After multiple iterations, the inlier set with the largest number of inliers is the maximum plane point cloud, and the maximum plane point cloud is removed. The remaining outliers are the object point cloud.
[0021] Further, the point cloud-based pose generation module further includes:
[0022] A minimum bounding rectangle box formation module, which is used to divide the object point cloud into multiple independent point cloud clusters through Euclidean clustering. Each independent point cloud cluster corresponds to an object. For each independent point cloud cluster, calculate the coordinate mean and covariance matrix of the point cloud data. The coordinate mean is the centroid of the point cloud. The eigenvectors of the covariance matrix form a rotation transformation matrix, and the point cloud data is mapped into the coordinate system formed by the rotation transformation matrix and the translation vector corresponding to the centroid to generate an OBB rectangle box as the minimum bounding rectangle box.
[0023] Further, the point cloud-based pose generation module further includes:
[0024] A grasping pose generation module, which is used to combine the minimum bounding rectangle box with the camera parameter information, solve the rotation matrix from the OBB principal axis coordinate system to the camera coordinate system, convert the rotation matrix from the OBB principal axis coordinate system to the camera coordinate system and the translation vector obtained from the centroid coordinates of the point cloud to the world coordinate system to obtain the rotation matrix and translation vector in the world coordinate system, and combine the Euler angles solved from the rotation matrix in the world coordinate system with the translation vector in the world coordinate system to form a grasping pose.
[0025] According to another aspect of the present invention, a model-free grasping planning system based on depth vision is provided, including: an image acquisition module, a pose generation module, and a trajectory planning module;
[0026] The image acquisition module is used to acquire the RGB image and depth image of the object to be grasped and the desktop where it is located;
[0027] The pose generation module is used to obtain the RGB image and the depth image, remove the desktop pixels in the RGB image to obtain the remaining pixel region, calculate the minimum bounding rectangle of the remaining pixel region to obtain the 2D pixel position of the object to be grasped in the RGB image, map the 2D pixel position to the depth image, and generate the grasping pose in the world coordinate system in combination with the camera parameter information;
[0028] The trajectory planning module is used to control the manipulator to move to the grasping pose and perform the grasping operation.
[0029] According to another aspect of the present invention, a model-free grasping planning system based on depth vision is provided, including: an image acquisition module, a pose generation module and a trajectory planning module;
[0030] The image acquisition module is used to acquire the point cloud of the object to be grasped and the desktop where it is located;
[0031] The pose generation module is used to obtain the point cloud, downsample the point cloud, remove the desktop point cloud, then cluster the remaining object point cloud to form an independent point cloud set, calculate the minimum bounding rectangle box of the independent point cloud set, and generate the grasping pose in the world coordinate system through the minimum bounding rectangle box and the camera parameter information;
[0032] The trajectory planning module is used to control the manipulator to move to the grasping pose and perform the grasping operation.
[0033] According to another aspect of the present invention, a model-free grasping planning method is provided, including:
[0034] Acquire the RGB image, depth image and point cloud of the object to be grasped and the desktop where it is located;
[0035] When the HSV color space of the object is not within the HSV color space of the desktop, remove the desktop pixels in the RGB image to obtain the remaining pixel region, calculate the minimum bounding rectangle of the remaining pixel region to obtain the 2D pixel position of the object to be grasped in the RGB image, map the 2D pixel position to the depth image, and generate the grasping pose in the world coordinate system in combination with the internal and external camera parameters;
[0036] When the HSV color space of the object is within the HSV color space of the desktop, downsample the point cloud, remove the desktop point cloud, then cluster the remaining object point cloud to form an independent point cloud, calculate the minimum bounding rectangle box of the independent point cloud, and generate the grasping pose in the world coordinate system through the minimum bounding rectangle box and the internal and external camera parameters;
[0037] The manipulator moves to the grasping pose and performs the grasping operation.
[0038] Generally speaking, compared with the prior art through the above technical solutions conceived by the present invention, the following beneficial effects can be achieved:
[0039] (1) The system of the present invention uses a processor to determine whether the color of an object is similar to that of the desktop. When the HSV color space of the object is within the HSV color space of the desktop, the color of the object is similar to that of the desktop. At this time, the grasping pose generated by the point cloud-based pose generation module is used to execute the grasping task. Because it is found through experiments that the pose obtained by the point cloud-based pose generation module can be well used for the grasping planning task of an unknown object model, with an average grasping success rate of 83.3%, and it can adapt to objects with colors similar to the desktop background, and still can reach a grasping success rate of 82.9% when the color is similar to the desktop background. When the HSV color space of the object is not within the HSV color space of the desktop, it indicates that the color of the object is not similar to that of the desktop. At this time, the grasping pose generated by the image-based pose generation module is used to execute the grasping task. Because the experimental results show that the image-based pose generation module proposed by the present invention does not require the 3D model of the object as input and can be well used for the grasping planning task of an unknown object model, with an average grasping success rate of 84.6%. The present invention does not need to establish a model and can well adapt to complex grasping scenarios with a large number of objects, and has a high grasping success rate when the object color is similar to the desktop. According to the changes in the grasping scenario, the present invention selects the grasping poses generated by different methods to execute the grasping task, which can improve the efficiency and the grasping accuracy in different scenarios at the same time.
[0040] (2) The present invention accurately obtains the minimum area rectangle that can completely enclose the target contour by rotating the rectangle. Accurately finding the minimum circumscribed rectangle of the object is beneficial to forming an accurate grasping pose, thereby improving the grasping success rate. Because a parallel two-finger gripper is usually used at the end of the robotic arm to grasp the object. In the grasping pose, the center of the parallel two-finger gripper corresponds to the three-dimensional coordinates of the center point of the minimum circumscribed rectangle in the world coordinate system. To ensure that the parallel two-finger gripper can perform the grasping operation to the maximum extent, the grasping direction should be vertically downward along the short side of the minimum circumscribed rectangle. Therefore, the rotation angle of the grasping pose along the Z axis of the world coordinate system is calculated using the two endpoints of the center point perpendicular to the two short sides of the minimum circumscribed rectangle. When mapping the points on the pixel coordinate system to the camera coordinate system, since the RGB image has been registered with the depth image, the depth value corresponding to the pixel can be obtained through the depth image.
[0041] (3) The color information of the desktop in the scene is relatively fixed. Therefore, threshold segmentation of the desktop pixels can be adopted in the color space to separate the object pixels. When segmenting the desktop pixels, RGB or HSV segmentation can be used. The greatest advantage of the RGB color space is that it is suitable for hardware display systems, being intuitive and easy to understand. The HSV color space can better describe the way humans observe colors. The hue H and saturation S are closely related to the way people perceive colors, and changes in brightness do not affect the hue and saturation components of the image. Since HSV can more intuitively describe the human eye's perception of colors and is easier to track an object of a specific color than the RGB color space, the HSV space is preferentially selected when segmenting and grasping the target object in the present invention.
[0042] (4) The amount of scene point cloud data obtained by the RGB-D camera is large, and there are many redundant point clouds. If not processed and directly used as the input of the algorithm framework, it will cause a huge burden and waste of computing resources and the real-time performance is very poor. Therefore, preprocessing of downsampling the point cloud is required. After obtaining the preprocessed point cloud, the point cloud of the desktop needs to be segmented. Segmenting the desktop through the depth point cloud can be independent of its color, so it can better adapt to desktops of different colors and is applicable to scenarios where the color of the object is similar to that of the desktop. The present invention uses the method of voxel filtering to obtain the downsampled point cloud with the smallest density and the most sufficient information within the allowable error precision range of the algorithm. This method reduces the number of point clouds through downsampling and tries to preserve the shape features of the point clouds at the same time.
[0043] (5) When the present invention regards the inlier set with the largest number of points as the desktop, it is to prevent the problem that the largest plane is not the desktop. After separating the desktop, the remaining point cloud set is the point cloud of the target object, and this point cloud set needs to be divided into independent point cloud sets for each target object. There are two methods to find the bounding rectangle box, namely the OBB box (Oriented Bounding Box) and the AABB box (Axis Aligned Bounding Box). Among them, the OBB box is closer to the object than the AABB box. Therefore, the OBB rectangle box is adopted in the present invention.
[0044] (6) The pose generation module in the system of the present invention removes the desktop pixels in the RGB image to obtain the remaining pixel area, calculates the minimum circumscribed rectangle of the remaining pixel area, obtains the 2D pixel position of the object to be grasped in the RGB image, maps the 2D pixel position to the depth map, and generates the grasping pose in the world coordinate system in combination with the camera parameter information. The model-free pose generation technology based on RGB-D does not require the 3D model of the object as input, can be well used for the grasping planning task of unknown object models, adapts to complex grasping scenarios with more objects, has a high grasping success rate, and at the same time greatly reduces the time for generating the grasping pose.
[0045] (7) In the system of the present invention, the pose generation module separates the point cloud of the plane where the desktop is located, calculates the minimum bounding rectangle box of the target object and obtains the pose of the target object. The pose of the target object obtained by the model-free pose generation technology based on the point cloud can be well used for the grasping planning task of unknown object models, adapt to complex grasping scenarios with more objects, has a high grasping success rate, and can adapt to objects with colors similar to the desktop background. Even when the color is similar to the desktop background, it still has a high grasping success rate. Description of the Drawings
[0046] Figure 1 is a schematic diagram of a model-free grasping planning system based on depth vision provided by an embodiment of the present invention;
[0047] Figure 2 is a 4-DOF pose schematic diagram provided by an embodiment of the present invention;
[0048] Figure 3 is a top view of the grasping pose provided by an embodiment of the present invention;
[0049] Figure 4 In (a) is a schematic diagram of a voxel grid provided by an embodiment of the present invention;
[0050] Figure 4 In (b) is a schematic diagram of a voxel provided by an embodiment of the present invention. Detailed Embodiment
[0051] In order to make the purpose, technical solutions and advantages of the present invention clearer, the present invention will be further described in detail below with reference to the drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention. In addition, the technical features involved in the various embodiments of the present invention described below can be combined with each other as long as they do not conflict with each other.
[0052] As Figure 1 shown, a model-free grasping planning system based on depth vision includes: an image acquisition module, a pose generation module, a processor, and a trajectory planning module;
[0053] The image acquisition module is used to acquire the RGB image, depth image and point cloud of the object to be grasped and the desktop where it is located;
[0054] The pose generation module includes a pose generation module based on images and a pose generation module based on point clouds;
[0055] The image-based pose generation module is used to obtain the RGB image and the depth image, remove the desktop pixels in the RGB image to obtain the pixel region of the object, calculate the minimum bounding rectangle of the pixel region of the object, obtain the 2D pixel position of the object to be grasped in the RGB image, map the 2D pixel position to the depth image, and generate the grasping pose in the world coordinate system in combination with the camera parameter information;
[0056] The point cloud-based pose generation module is used to obtain the point cloud. After downsampling the point cloud, remove the desktop point cloud, then cluster the remaining object point cloud to form an independent point cloud set, calculate the minimum bounding rectangle box of the independent point cloud set, and generate the grasping pose in the world coordinate system through the minimum bounding rectangle box and the camera parameter information;
[0057] The processor is used to compare the HSV color space of the object to be grasped and the desktop where the object is located. When the HSV color space of the object is within the HSV color space of the desktop, the grasping pose generated by the point cloud-based pose generation module is input into the trajectory planning module. When the HSV color space of the object is not within the HSV color space of the desktop, the grasping pose generated by the image-based pose generation module is input into the trajectory planning module;
[0058] The trajectory planning module is used to control the robotic arm to move to the grasping pose and perform the grasping operation.
[0059] Embodiment 1
[0060] The end effector of the robotic arm of the model-free grasping planning system based on depth vision in the present invention is a typical parallel two-finger gripper.
[0061] To clearly describe the grasping operation, usually 4 parameters are used to describe the parallel two-finger gripper: the width hand_depth of the gripper closing area, the height hand_height of the gripper closing area, the length hand_width of the gripper closing area, and the thickness finger_width of the two fingers of the gripper.
[0062] In Embodiment 1 of the present invention, the parameters of the parallel two-finger gripper are:
[0063]
[0064] The maximum width of the object to be grasped is:
[0065] obj_width max = hand_width - 2 * finger_width = 70mm
[0066] The object set grasped by the grasping task of the present invention in the real environment is a variety of common items in life, such as fruits, biscuits, boxes, etc. For most objects, their 3D models have not been established and are difficult to obtain. A small number of them are similar in color to the desktop. The goal of the grasping task is to grasp all the objects on the desktop and collect them in a specified area.
[0067] The robotic arm first moves to the initial pose, and the grasping scene information is collected by the image acquisition module, including RGB images and depth information. Then, the visual information is input into the pose generation module to obtain the grasping pose of the target object. The present invention stipulates that the grasping pose is always vertically downward, that is, a 4-DOF grasping pose is calculated. Finally, the robotic arm is controlled by the trajectory planning module to move to the grasping pose and perform the grasping operation. In the image acquisition module, the robotic arm first moves to the acquisition pose, and simultaneously acquires the RGB and depth maps, and generates a point cloud through the camera SDK. In the pose generation module, the RGB image and depth information are used to calculate the grasping target pose and generate the grasping pose. In the trajectory planning module, the RRTConnect algorithm is used. First, the robotic arm reaches a position L above the positive z-axis of the 4-DOF pose of the grasping target through joint motion, and then the end of the robotic arm is controlled to move linearly along the z-axis direction of the grasping pose by a distance of L. At this time, the parallel two-finger gripper closes to grasp the object, and the object is moved to a specified place to complete the grasping operation.
[0068] As Figure 2 shown, the 4-DOF grasping pose can be described as where the coordinate point of the grasping pose in the world coordinate system is P(x, y, z), and the angle of rotation along the z-axis of the target pose is
[0069] When the robotic arm is performing grasping planning, it may cause phenomena such as missing the grasp or deviating from the grasp due to the deviation of the grasping pose, resulting in the failure of the planning. Therefore, certain metrics are needed to determine whether the grasping is successful, and the commonly used metric is the force closure model.
[0070] When using the two-finger gripper to grasp, there are only 2 contact points with the object. When the force applied by the gripper can offset the forces in other directions of the object, dynamic balance is achieved. According to Nguyen's theorem, the force closure condition can be described as: whether the connecting line between the contact points of the two-finger gripper and the target object is within the friction cone. If the connecting line is within the friction cone, it means that the force closure condition is met and the grasping planning is successful; otherwise, the force closure condition is not met and the grasping planning fails.
[0071] The pose generation module includes: an image-based pose generation module and a point cloud-based pose generation module.
[0072] In the present invention, the processor is used to compare the HSV color space of the object to be grasped and the desktop where the object is located. When the HSV color space of the object is within the HSV color space of the desktop, the grasping pose generated by the point cloud-based pose generation module is input into the trajectory planning module. When the HSV color space of the object is not within the HSV color space of the desktop, the grasping pose generated by the image-based pose generation module is input into the trajectory planning module.
[0073] The image-based pose generation module, on the basis of camera calibration, acquires the RGB image and depth map of the current scene, separates the desktop pixels through the prior knowledge of the known desktop color, and regards the remaining pixel regions with a certain scale as the grasping target; by calculating the minimum bounding rectangle of each remaining pixel region, the 2D pixel position of the target object in the RGB image is obtained; the 2D position is mapped to the depth map, and finally the pose of the object in the world coordinate system is generated through the depth information and the internal and external camera parameter information.
[0074] RGB is the most common color model for hardware display devices (such as PC monitors, mobile terminal displays, etc.). According to the human eye structure, all colors can be regarded as the superposition combination of three primary colors, namely red (R), green (G), and blue (B) in different proportions, that is, the three primary colors of light.
[0075] Normalize the RGB color space cube to a unit cube, then the RGB values [0, 255] are normalized to the interval [0, 1]. The origin O(0, 0, 0) of the coordinate system in the color space is black, and the vertex W(1, 1, 1) farthest from the origin corresponds to white. The gray scale distribution value from black to white is on the body diagonal . According to the RGB color space model, each color can be expressed as the brightness values of the three primary color planes from 0 to 255, and the changes and mutual superpositions of the three color channels can form 16,777,216 (i.e., 256 3 ) colors.
[0076] The greatest advantage of the RGB color space is that it is suitable for hardware display systems, intuitive and easy to understand. However, when describing colors, the three components of the three primary colors of light are highly correlated and the uniformity is poor. For example, when the brightness of the same color changes, all three components will change accordingly.
[0077] The HSV color space is a perception-based color model, which describes colors as three attributes: H (Hue, hue), S (Saturation, saturation), and V (Value, brightness). The specific description is as follows:
[0078] a) Hue: The wavelength of the light reflected or transmitted through the object, distinguished by color, such as red, green;
[0079] b) Brightness: The lightness or darkness of a color, such as dark red or bright red;
[0080] c) Saturation: The intensity of a color, such as deep red or light red.
[0081] The HSV color space can better describe the way humans perceive colors. The hue H and saturation S are closely related to the way people sense colors, and changes in brightness do not affect the hue and saturation components of an image. Its model corresponds to a cone in a cylindrical coordinate system.
[0082] In the conical HSV color space, the vertex of the cone has a brightness V = 0, the base has V = 1, and the brightness V increases linearly along the generatrix of the cone starting from the vertex; the color H is determined by the angle of rotation around the central axis of the cone. Taking R, G, B as examples, red corresponds to an angle of 0°, green corresponds to an angle of 120°, blue corresponds to an angle of 240°, and each color and its complementary color differ by 180°; the saturation S is 0 at the central axis of the cone and S = 1 on the lateral surface of the cone, and increases linearly from the inside to the outside.
[0083] Since colors objectively exist in reality and different color spaces are just descriptions from different angles, there is a unique correspondence between RGB and HSV color space parameters, and they can be mutually converted through the following transformation relationships:
[0084] a) Conversion from RGB to HSV color space:
[0085]
[0086]
[0087] v = max
[0088] where r, g, b are the components of the corresponding channels in the RGB color space, max = max(r, g, b), and min = min(r, g, b).
[0089] b) Conversion from HSV to RGB color space:
[0090] First, calculate the intermediate variables:
[0091]
[0092] where h, s, v are the values of the three channels in the HSV color space, and mod is the modulo symbol. The components of the three channels in the RGB color space are obtained through the intermediate variables as follows:
[0093]
[0094] Since HSV can more intuitively describe the human eye's perception of colors and is easier to track objects of a specific color than the RGB color space, the HSV space is adopted in the embodiments of the present invention when segmenting and grasping target objects.
[0095] According to relevant information consulted, the HSV values of some typical colors are shown in Table 1 as follows:
[0096] Table 1 Thresholds of the HSV color space for typical colors
[0097]
[0098] In Table 1, except that the hue of red has two ranges of 0 - 10 and 156 - 180, the hues of other colors have only one range.
[0099] Although the color of the desktop is off - white, which is an atypical color, the present invention designs a visualization software for segmenting the desktop of a specified color based on Qt. When segmenting, first, an RGB image needs to be collected as a sample, and then the software is used to drag the slider to change the channel values in the HSV space to segment the desktop color. The obtained interval of the HSV color space after segmenting the desktop is:
[0100]
[0101] It can be seen that the pixel points of the desktop part in the RGB image have been removed, and only the pixel points of the target object are retained. After gray - scaling the RGB image, there will be multiple pixel - point sets in the obtained grayscale image. The target - object point sets are distinguished by the Moore - Neighbor contour - finding algorithm. Its basic idea is to find all continuous and relatively large - scale pixel points in the binary image, extract the contours formed by them and return. The point set surrounded by each contour corresponds to the pixel - point set of an object. After distinguishing the object pixel points, the minimum - enclosing rectangle of each object pixel - point set can be calculated, that is, the pixel position of the object in the RGB image is obtained. The minimum - enclosing rectangle refers to the rectangle with the smallest area that can completely enclose the target contour. When solving, first, an initial angle is set so that the rectangle just encloses the object pixel - point set, and the rectangle is rotated in the coordinate system with a certain step size (0° - 90°), the area of the enclosing rectangle at each rotation angle is calculated, and the rectangle with the smallest area is found. Then, the center point P(u o , v o ) and the rotation angle θ are solved.
[0102] After obtaining the minimum - enclosing rectangle TR of the object through the RGB image, the coordinates of the center point of TR in the pixel coordinate system can be obtained as P(u o , v o) and the horizontal angle θ between the TR and the RGB image. In the grasping pose, the three-dimensional coordinates of the center corresponding point P of the parallel two-finger gripper in the world coordinate system. To ensure that the parallel two-finger gripper can perform the grasping operation to the maximum extent, the grasping direction should be vertically downward along the short side of the minimum circumscribed rectangle, as Figure 3 shown.
[0103] The coordinates of the two endpoints of the short side in the pixel coordinate system are calculated as A(u o , v o ) and B(u A , v A ) through P(u B , v B ) and θ. When mapping the points on the pixel coordinate system to the camera coordinate system, since the RGB image of the vision sensor has been registered with the depth map, the depth value z corresponding to the pixel can be obtained through the depth map. For the point (u, v) in the pixel coordinate system and its depth value z, the coordinates (x, y, z) of this pixel in the camera coordinate system can be solved by the following formula:
[0104]
[0105] where f x , f y , u0, and v0 are the camera internal parameters. (u0, v0) is the coordinate of the principal point of the camera, and the subscript of the principal point coordinate is 0. The subscript of the coordinate (u o , v o ) of point P is O. f x , f y are the camera focal lengths.
[0106] The coordinates of P(u o , v o ), A(u A , v A ), and B(u B , v B ) in the camera coordinate system are calculated as P P (x P , y P , z P ), P A (x A , y A , z A ), P B (x B , y B , z B ) respectively. In order to obtain the type 4-DOF pose through the coordinates of points P, A, and B, the coordinates of Pose in the world coordinate system are the grasping center P o (x o , yo , z o ), the rotation angle of the Pose along the z-axis of the world coordinate system is the vector projected onto the X-O-Y plane and the angle with the x-axis, that is Each parameter can be calculated by the following formula.
[0107]
[0108] The definition of atan2 is shown in the following formula:
[0109]
[0110] It can be seen that the atan2 function can calculate the angle between [-π, π] according to the input x and y coordinate values, and judge the quadrant where the angle is located according to the signs of x and y. While the inverse trigonometric function arctan usually can calculate two sets of solutions or no solution. Therefore, the atan2 function is more stable than the inverse trigonometric function arctan.
[0111] The length L of the short side of the minimum circumscribed rectangle TR can be expressed as the vector in three-dimensional space projected onto the X-O-Y plane of the length, as shown in the following formula:
[0112]
[0113] To ensure that the target can be grasped, it is necessary to ensure that the length obj_width of the gripper opening max is less than the minimum width of the object, that is, the length L of the short side of the minimum circumscribed rectangle TR. If obj_width max < L, the grasping planning can continue. If objwidth max ≥ L, it means that the force-closure grasping condition is not met, and this grasping is abandoned.
[0114] Pose generation module based on point cloud. In addition to the depth map, the depth visual information obtained by the Intel RealSense D435i camera can also calculate the depth point cloud through the camera SDK, and the spatial pose of the grasping target is calculated based on the depth point cloud. The amount of scene point cloud data obtained by the RGB-D camera is large, and there are many redundant point clouds. If not processed and directly used as the input of the algorithm framework, it will cause a huge burden and waste of computing resources, and the real-time performance is very poor. Therefore, it is necessary to preprocess the point cloud by downsampling. After obtaining the preprocessed point cloud, it is necessary to segment the point cloud of the desktop. A typical feature of the desktop in the grasping task is the largest plane in the field of view. Segmenting the desktop through the depth point cloud can be not affected by its color, so it can better adapt to desktops of different colors and is applicable to scenarios where the color of the object is similar to that of the desktop. After separating the desktop point cloud, cluster the remaining object point clouds, separate each grasping target point cloud into independent point clouds, and finally calculate the minimum bounding rectangle box of the independent point cloud, and solve the pose of the target object through the parameters of the rectangle box to obtain the grasping pose.
[0115] Since the point cloud obtained by the RGB-D camera SDK is a dense point cloud and the data volume is large, downsampling is required. The present invention uses the method of voxel grid filtering (abbreviated as voxel filtering), as Figure 4 (a) in Figure 4 and (b) in center shown. Its core idea is to divide the three-dimensional space into several small cubes with side length v, where Δx = Δy = Δz = v. For all the points in the cube, only the point P center closest to the center of the cube is retained, and other points are filtered out, so as to obtain the downsampled point cloud with the minimum density and the most sufficient information within the allowable error precision range of the algorithm. This method reduces the number of point clouds through downsampling and tries to preserve the shape features of the point clouds at the same time.
[0116] After downsampling the input dense point cloud, for the largest tabletop in the point cloud, voxel filtering greatly reduces the redundant point cloud. The largest plane after filtering is still the tabletop, so it does not affect tabletop segmentation. For the grasped target object, since the object has a certain volume, voxel filtering can still retain the key point cloud of the object, and has little impact on target recognition. However, voxel filtering also has some disadvantages. That is, when the object is small enough and smaller than the side length of the voxel grid, it may be filtered out. Therefore, the size of the voxel grid needs to be determined according to the side length of the grasped target. In the present invention, the voxel filtering grid size is used as an adjustable hyperparameter. When the grasped targets are generally small, a smaller voxel grid is used, so that small targets will not be filtered out. However, at the same time, less point cloud is filtered, resulting in slow data processing and poor real-time performance of the algorithm. When the targets generally have a large volume, a larger voxel grid size can be used, which can significantly reduce the calculation amount and improve the real-time performance of grasping.
[0117] After voxel filtering downsampling, the tabletop can be segmented as the background. In the grasping task, the tabletop is the largest plane within the field of view, and the RANSAC algorithm can be used for fitting and segmentation. This method assumes that the distribution in space can be described by the same model parameters. The point cloud set that conforms to this distribution model is the inlier, and the point cloud set that does not fit this model is the outlier. For the input point cloud data, the steps of the RANSAC algorithm are as follows:
[0118] 1) Randomly select N points as inliers and fit these N points into a specified model;
[0119] 2) Substitute the outliers into the fitted model to judge whether they belong to the inlier group and record the number of inliers;
[0120] 3) Specify the number of iterations N and repeat step 2) N times. The model with the largest number of inliers is the solution result.
[0121] According to the steps of the RANSAC algorithm, when segmenting the tabletop in the grasping task, 3 points in the point cloud are randomly selected as inliers each time, and an initial plane is generated with these 3 inliers. Then, all the outliers outside the inliers are traversed. If the distance of the outlier from the initial plane is less than the threshold T, it is added to the inlier set. After N iterations, the inlier set with the largest number of points in the whole process is the point cloud of the largest plane (i.e., the tabletop). The inliers are combined and separated, and the remaining outlier set is the point cloud of the grasped target object, where both T and N are adjustable hyperparameters.
[0122] When regarding the inlier set with the largest number of points as the tabletop, in order to prevent the problem that the largest plane is not the tabletop, an attitude constraint is established. The normalized equation of the largest plane fitted by the RANSAC algorithm is Ax + By + Cz + D = 0, where A 2+B 2 +C 2 If = 1, the plane normal vector is transformed into the world coordinate system to obtain Based on the prior knowledge that the known tabletop normal vector is always perpendicular to the ground, the fitted plane normal vector should be parallel to the z-axis in the world coordinate system, or the normal vector The included angle with the z-axis is less than a certain angle value θ. In the embodiment of the present invention, θ = 5° is taken, that is and The cosine value of the included angle:
[0123]
[0124] Since the problem of the present invention is taken in practical significance and The included angle range is In this interval, the cosine function cos decreases as the angle increases. Therefore, C ≥ cos5° = 0.99619469809.
[0125] After separating the tabletop, the remaining point cloud set is the point cloud of the target object. It is necessary to divide this point cloud set into independent point cloud sets for each target object. One of the commonly used point cloud clustering algorithms is Euclidean clustering. The core idea of Euclidean clustering is to classify all points with an Euclidean distance less than a certain threshold into the same category. The specific process is as follows: Select an unprocessed point. If the point has not been classified, a new category is established with this point, and points with an Euclidean distance less than the threshold T around it (called "neighboring points") are found and added to this category; if it has been classified, only the neighboring points around it need to be found, and then the newly added points to this category are continuously iteratively processed until there are no newly added points, so as to obtain all the points of this category. Finally, the point cloud will be divided into multiple point sets according to the category, and each point set represents a target object, and the pose of the point set is the pose of the required target object.
[0126] After dividing the object point cloud into several point cloud sets through Euclidean clustering, since the independent point cloud sets corresponding to each object are not regular, it is relatively difficult to directly obtain their poses. Therefore, the rectangular bounding box of the object point cloud set can be obtained first. There are two methods to find the bounding rectangular box, namely the OBB box (Oriented Bounding Box, oriented bounding box) and the AABB box (Axis Aligned Bounding Box, axis-aligned bounding box). Among them, the OBB box is closer to the object than the AABB box. Therefore, the OBB rectangular box is adopted in the present invention. The main steps are as follows:
[0127] (1) Take the single object point cloud set distinguished by Euclidean clustering as input and traverse it, obtain the coordinate information of each point, and calculate the coordinate mean and covariance matrix of the point cloud data, where the coordinate mean is the centroid P(x, y, z) of the point cloud;
[0128] (2) Calculate the eigenvectors and eigenvalues of the covariance matrix, and use the eigenvectors to form a rotation transformation matrix Rot;
[0129] (3) Map the point cloud data to a coordinate system consisting of the Rot rotation transformation and the translation transformation corresponding to the center of mass P (x, y, z) to generate an OBB rectangular box.
[0130] Through the above steps, the rotation matrix Rot from the OBB principal axis coordinate system to the camera coordinate system is solved, and the translation vector Trans(x, y, z) is obtained through the point cloud centroid coordinates P(x, y, z), and the two are combined into a rotation matrix T, and T is converted to the world coordinate system to obtain T w , where T w The rotation matrix is Rot w , the translation vector is Trans(x w ,y w , z w ), then T w The rotation matrix part of Rot w Solve the Euler angles to get the position of the target object, and finally obtain a 4-DOF grasping position that is always vertical to the tabletop and downward
[0131] Example 2
[0132] The RGB-D based methods are:
[0133] Collect RGB images and depth images of the object to be grasped and the desktop on which it is located;
[0134] Get the RGB image and the depth image, remove the desktop pixels in the RGB image, get the remaining pixel area, calculate the minimum bounding rectangle of the remaining pixel area, get the 2D pixel position of the object to be grasped in the RGB image, map the 2D pixel position to the depth image, and generate the grasping posture in the world coordinate system in combination with the camera parameter information;
[0135] Control the robot arm to move to the grasping position and perform the grasping operation.
[0136] Example 3
[0137] Point cloud based methods are:
[0138] Collect the point cloud of the object to be grasped and the desktop on which it is located;
[0139] After downsampling the point cloud, remove the tabletop point cloud, then cluster the remaining object point clouds to form independent point cloud sets, calculate the minimum bounding rectangle of the independent point cloud sets, and generate the grasping pose in the world coordinate system through the minimum bounding rectangle and camera parameter information;
[0140] Control the manipulator to move to the grasping pose and perform the grasping operation.
[0141] Example 4
[0142] Through the use of the RGB-D based pose generation method and point cloud based pose generation method of the present invention and the pose estimation methods DOPE and PointNetGPD used in the prior art for grasping experiments, the experimental results are analyzed in detail.
[0143] Most of the objects in the object set for the grasping experiment do not have their 3D models established, and there are a small number of objects with colors similar to the tabletop background. When designing the comparative experiment of the grasping task, the following 4 groups of comparative experiment scenarios are set:
[0144] 1) Scenario 1: 5 known model objects;
[0145] 2) Scenario 2: 3 known models + 2 unknown model objects;
[0146] 3) Scenario 3: 5 unknown model objects;
[0147] 4) Scenario 4: 5 unknown model objects, including 2 objects with colors similar to the tabletop background;
[0148] Compare the method of the present invention with the existing pose estimation methods DOPE and PointNetGPD. Set each group of experiments to be repeated 20 times, with 7 grasps for each task. The evaluation index is the grasping success rate, calculated by the number of successful grasps / total number of grasps. In addition, since the manipulator movement time is determined by the actual pose of the object, to control variables, only the average time consumption for generating the grasping pose is used to evaluate the real-time performance of the algorithm. The experimental results are shown in Table 2.
[0149] Table 2 Grasping comparison experiment results
[0150]
[0151] *Note: The RGB-D based method needs to collect one image sample, and this time has been averaged into each experiment.
[0152] Based on the experimental results in Table 2, the following analysis is made: 1) When DOPE obtains the pose of the target object, it has a high grasping success rate for objects with known 3D models. However, as the number of objects with unknown models increases, the grasping success rate drops significantly. If there are no objects with known 3D models in the scene, this method completely fails and cannot obtain the pose of objects with unknown models; 2) PointNetGPD can also obtain the 6-DOF grasping pose of the target object when the 3D model of the object is unknown, but this method takes a long time to calculate the optimal 6-DOF pose; 3) The present invention uses an RGB-D based pose generation method for grasping, which does not require the input of the 3D model of the object, can grasp objects with unknown models in the scene, and has an average grasping success rate of 84.6%. Compared with DOPE, the grasping success rate is increased by 49.8%, and compared with PointNetGPD, the grasping success rate is decreased by 3.0%. Since this method only calculates the 4-DOF grasping pose, the time consumption is reduced by 73.1% compared with PointNetGPD, greatly improving the speed of grasping pose generation. It can be seen that this method is more suitable for grasping scenarios that require high real-time performance; 4) The pose generation method based on point cloud used in the present invention for grasping is limited by the accuracy of the depth sensor. Compared with the RGB-D based method, the average grasping success rate is reduced by 1.5%, and the pose acquisition time is increased by 1.1 times. However, in a scene containing two objects with colors similar to the desktop background, the grasping success rate of the point cloud based method is increased by 11.7% compared with the RGB-D based method, and the average time consumption is reduced by 43.6% compared with PointNetGPD. It can be seen that this method is more suitable for scenarios where the object color is very similar to the desktop background.
[0153] Example 1 has modules for generating two grasping poses simultaneously. Examples 2 and 3 only have one method for generating a grasping pose. It can be seen that the model-free pose generation method based on RGB-D in Example 2 calculates the 4-DOF pose of the target object according to the pixel information of the separated object. The experimental results show that the model-free pose generation method based on RGB-D proposed in the present invention does not require a 3D model of the object as input and can be well used for the grasping planning task of an unknown object model, achieving an average grasping success rate of 84.6%. The model-free pose generation method based on point cloud in Example 3 separates the point cloud of the desktop where the desktop is located, calculates the minimum bounding rectangle of the target object and obtains the 4-DOF pose of the target object. The experimental results show that the pose of the target object obtained by the model-free pose generation method based on point cloud proposed in the present invention can be well used for the grasping planning task of an unknown object model, achieving an average grasping success rate of 83.3%, and can adapt to objects with a color similar to the desktop background, and still achieve a grasping success rate of 82.9% when the color is similar to the desktop background. Example 1 includes modules corresponding to the two methods, selects different modules to calculate the grasping pose according to the actual situation, and guides the robotic arm grasping task through the grasping pose. The experimental results show that the method proposed in the present invention performs well when grasping an object with an unknown model in a real environment, and has high grasping efficiency and real-time performance, and the method based on point cloud can adapt to objects with a color similar to the desktop background. According to the change of the grasping scene, different grasping poses are selected to execute the grasping task, which can improve the grasping accuracy in different scenes while improving the efficiency.
[0154] It is easy for those skilled in the art to understand that the above are only preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent replacements, and improvements made within the spirit and principle of the present invention should be included in the protection scope of the present invention.
Claims
1. A model-free grasping planning system based on depth vision, characterized in that, Including: An image acquisition module, a pose generation module, a processor, and a trajectory planning module; The image acquisition module is used to acquire the RGB image, depth image, and point cloud of the object to be grasped and the desktop where it is located; The pose generation module includes an image-based pose generation module and a point cloud-based pose generation module; The image-based pose generation module is used to obtain the RGB image and depth image, remove the desktop pixels in the RGB image to obtain the pixel area of the object, calculate the minimum bounding rectangle of the pixel area of the object to obtain the 2D pixel position of the object to be grasped in the RGB image, map the 2D pixel position to the depth image, and generate the grasping pose in the world coordinate system in combination with the camera parameter information; The point cloud-based pose generation module is used to obtain the point cloud, remove the desktop point cloud, then cluster the remaining object point cloud to form an independent point cloud set, calculate the minimum bounding rectangle box of the independent point cloud set, and generate the grasping pose in the world coordinate system through the minimum bounding rectangle box and the camera parameter information; The processor is used to compare the HSV color spaces of the object to be grasped and the desktop where the object is located. When the HSV color space of the object is within the HSV color space of the desktop, the grasping pose generated by the point cloud-based pose generation module is input into the trajectory planning module. When the HSV color space of the object is not within the HSV color space of the desktop, the grasping pose generated by the image-based pose generation module is input into the trajectory planning module; The trajectory planning module is used to control the manipulator to move to the grasping pose and execute the grasping operation.
2. The model-free grasping planning system based on depth vision according to claim 1, characterized in that, The image-based pose generation module includes: The minimum bounding rectangle formation module is used to gray-scale the pixel area of the object to obtain the pixel point set of the object, set the initial angle so that the rectangle encloses the pixel point set of the object, rotate the rectangle, calculate the area of the rectangle at each rotation angle, and use the rectangle with the minimum area as the minimum bounding rectangle.
3. A model-free grasping planning system based on depth vision according to claim 2, characterized in that The image-based pose generation module further includes: The grasping pose generation module is used to obtain the coordinates of the center point of the minimum bounding rectangle in the pixel coordinate system, calculate the coordinates of the two endpoints perpendicular to the two short sides of the minimum bounding rectangle at the center point in the pixel coordinate system, obtain the depth values of the center point and the two endpoints through the depth image. For the coordinates of the center point in the pixel coordinate system and its depth value, convert them to the coordinates of the center point in the camera coordinate system through the camera internal parameters. For the coordinates of the two endpoints in the pixel coordinate system and their depth values, convert them to the coordinates of the two endpoints in the camera coordinate system through the camera internal parameters, calculate the angle between the projection of the vector of the two endpoints on the X-O-Y plane and the X axis to obtain the rotation angle of the grasping pose along the Z axis of the world coordinate system, calculate the coordinates of the grasping center in the world coordinate system through the coordinates of the center point in the camera coordinate system, and the coordinates of the grasping center in the world coordinate system and the rotation angle of the grasping pose along the Z axis of the world coordinate system form the grasping pose.
4. A model-free grasping planning system based on depth vision according to claim 3, wherein, The image-based pose generation module further includes: A pixel segmentation module, which is used to describe an RGB image in the RGB color space, remove the pixel region of the desktop in the RGB color space, retain the pixel region of the object in the RGB color space, or convert the RGB color space to the HSV color space, remove the pixel region of the desktop in the HSV color space, and retain the pixel region of the object in the HSV color space.
5. A model-free grasping planning system based on depth vision according to claim 1, characterized in that, The point cloud-based pose generation module includes: A point cloud segmentation module, which is used to evenly divide the three-dimensional space composed of the point cloud into multiple cubes, retain the center point of each cube to obtain the downsampled point cloud, randomly select N points from the downsampled point cloud as inliers, fit the inliers into an initial plane, and then traverse all the outliers outside the inliers in the downsampled point cloud. If the distance from an outlier to the initial plane is less than the threshold T, it is added to the inlier set. After multiple iterations, the inlier set with the largest number of inliers obtained is the maximum plane point cloud, and the maximum plane point cloud is removed. The remaining outliers are the object point cloud.
6. A model-free grasping planning system based on depth vision according to claim 5, characterized in that The point cloud-based pose generation module further includes: A minimum bounding rectangle box formation module, which is used to divide the object point cloud into multiple independent point cloud clusters through Euclidean clustering. Each independent point cloud cluster corresponds to an object. For each independent point cloud cluster, calculate the coordinate mean and covariance matrix of the point cloud data, where the coordinate mean is the centroid of the point cloud, form a rotation transformation matrix with the eigenvectors of the covariance matrix, map the point cloud data to the coordinate system composed of the rotation transformation matrix and the translation vector corresponding to the centroid, and generate an OBB rectangle box as the minimum bounding rectangle box.
7. The model-free grasping planning system based on depth vision according to claim 6, characterized in that, The point cloud-based pose generation module further includes: A grasping pose generation module, which is used to combine the minimum bounding rectangle box with the camera parameter information, solve the rotation matrix from the OBB principal axis coordinate system to the camera coordinate system, convert the rotation matrix from the OBB principal axis coordinate system to the camera coordinate system and the translation vector obtained through the centroid coordinates of the point cloud to the world coordinate system to obtain the rotation matrix and translation vector in the world coordinate system, and combine the Euler angles solved from the rotation matrix in the world coordinate system with the translation vector in the world coordinate system to form a grasping pose.
8. A model-free grasping planning method based on depth vision, characterized in that, It includes: Collect the RGB image, depth image and point cloud of the object to be grasped and its desktop; When the HSV color space of the object is not within the HSV color space range of the desktop, remove the desktop pixels in the RGB image to obtain the remaining pixel region, calculate the minimum circumscribed rectangle of the remaining pixel region to obtain the 2D pixel position of the object to be grasped in the RGB image, map the 2D pixel position to the depth image, and generate a grasping pose in the world coordinate system by combining the internal and external camera parameters; When the HSV color space of the object is within the HSV color space range of the desktop, remove the desktop point cloud, then cluster the remaining object point cloud to form independent point clouds, calculate the minimum bounding rectangle box of the independent point clouds, and generate a grasping pose in the world coordinate system through the minimum bounding rectangle box and the internal and external camera parameters; The robotic arm moves to the grasping pose and performs a grasping operation.
Citation Information
Patent Citations
Grabbing method for industrial stacked parts, terminal equipment and readable storage medium
CN112109086A
Robot grabbing pose detection method based on domain migration under single-view-angle point cloud
CN112489117A