Object grasping method and system based on point cloud deep learning
By using a point cloud-based deep learning method, the grasping position of the target object is calculated using a deep learning network and a hand-eye transformation matrix. This solves the problems of low efficiency in 2D vision grasping and large errors in traditional 3D point cloud methods when the lighting changes, and achieves high-precision automated grasping.
Patent Information
- Application Number
- CN202511142769.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-15
- Publication Date
- 2025-11-04
- Estimated Expiration
- 2045-08-15
AI Technical Summary
Existing 2D vision grasping methods cannot acquire depth information, resulting in low grasping efficiency and accuracy, while traditional 3D point cloud methods have large grasping errors when the ambient lighting changes.
A point cloud-based deep learning approach is adopted. By acquiring the point cloud data of the target object, point cloud registration is performed using a deep learning network. Combined with the hand-eye transformation matrix and the most recent iteration algorithm, the grasping position of the target object in the base coordinate system of the robotic arm is calculated.
It improves the success rate and accuracy of grasping, reduces grasping deviations caused by inaccurate matching, realizes an automated grasping process, and adapts to different types of objects and scenarios.
Smart Images

Figure CN120715903B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of computer vision and industrial automation, and particularly relates to an object grasping method and system based on point cloud deep learning. BACKGROUND
[0002] With the continuous development of industrial automation and machine vision technology, object grasping and positioning play an increasingly important role in industrial production and logistics transportation. At present, the common technical solutions are a grasping system based on 2D vision and a grasping system based on traditional 3D point cloud. The grasping method using 2D is suitable for grasping planar objects, and the grasping efficiency and accuracy are low due to the lack of depth information to obtain the attitude in RX and RY directions. The traditional 3D point cloud method has high dependence on the integrity of the target object point cloud when processing, and the imaging may be missing when the environment light changes, resulting in large grasping errors.
[0003] Based on this, the present application provides an object grasping method and system based on point cloud deep learning. SUMMARY
[0004] In order to overcome the above-mentioned problem of large grasping error under the condition of change of environment light, the present application provides an object grasping method and system based on point cloud deep learning.
[0005] The purpose of the present application is achieved by adopting the following technical solutions:
[0006] In a first aspect, the present application provides an object grasping method based on point cloud deep learning, comprising:
[0007] Obtaining a target object point cloud of a target object, using a point cloud registration model based on a deep learning network to perform point cloud matching on the target object point cloud to obtain an initial matching position, and then optimizing the matching position by a nearest iteration algorithm to obtain an accurate transformation matrix of the target object relative to a template in a camera coordinate system;
[0008] Based on a pre-acquired hand-eye transformation matrix, calculating a conversion relationship of the point cloud attitude of the target object relative to a tool coordinate system through teaching; the hand-eye transformation matrix is used to indicate the transformation relationship between a sensor coordinate system and an end-of-arm coordinate system of a robot arm;
[0009] According to the conversion relationship and the accurate transformation matrix, calculating a grasping position coordinate of the target object in a base coordinate system of the robot arm.
[0010] Preferably, the hand-eye transformation matrix is obtained in the following manner:
[0011] According to the motion transformation of the end of the robot arm and the motion transformation observed by the sensor, the hand-eye transformation matrix is solved and obtained.
[0012] Preferably, the hand-eye transformation matrix is obtained by transforming the motion of the end of the robot arm and the motion observed by the sensor, and comprises:
[0013] Camera intrinsic calibration is performed, corner points are extracted and refined to sub-pixel accuracy, and Zhang's calibration algorithm is used to obtain camera internal parameters and distortion coefficients as the camera calibration result;
[0014] Image distortion correction is performed, and the camera calibration result is used to correct the image distortion, and the sub-pixel accuracy corner point coordinates are re-extracted;
[0015] A world coordinate system is defined on the calibration board, and 3D coordinates of the corner points are obtained;
[0016] A perspective projection model is established using the PnP algorithm, and the rotation and translation matrix of the world coordinate system to the camera coordinate system is solved and used as the camera pose;
[0017] Based on the robot arm pose and the camera pose, the hand-eye transformation matrix is obtained by solving the equation.
[0018] Preferably, the hand-eye transformation matrix is obtained by solving the equation based on the robot arm pose and the camera pose, and comprises:
[0019] Based on the robot arm pose and the camera pose, the hand-eye calibration matrix is obtained by converting the relationship;
[0020] The accuracy value of the hand-eye calibration matrix is evaluated by calculating the re-projection error, and when the accuracy value meets the preset accuracy requirement, the hand-eye calibration matrix is used as the hand-eye transformation matrix;
[0021] The accuracy value of the hand-eye calibration matrix is evaluated by calculating the re-projection error, and when the accuracy value meets the preset accuracy requirement, the hand-eye calibration matrix is used as the hand-eye transformation matrix;
[0022] Using two sets of pose data, one set is used to estimate the camera pose of the other set, and the pose and 3D point set are re-projected to obtain the image plane 2D point set coordinates. The mean square error value between the projection point coordinates and the corner point coordinates is calculated and used as the accuracy value.
[0023] Preferably, when the sensor is fixed at the end of the robot arm, the hand-eye transformation matrix is obtained by the eye-in-hand calibration method, and the conversion relationship is as follows:
[0024] ;
[0025] When the sensor holder is fixed at a position outside the robot arm, the hand-eye transformation matrix is obtained by the eye-in-hand calibration method, and the conversion relationship is as follows:
[0026]
[0027] wherein, a transformation matrix of a target object coordinate system to a sensor coordinate system, a transformation matrix of a sensor coordinate system to a robot end coordinate system, a transformation matrix of a robot end coordinate system to a robot base coordinate system, a transformation matrix of a sensor coordinate system to a robot base coordinate system, a transformation matrix of a robot base coordinate system to a robot end coordinate system; T 0 and T 1 represent transformation matrices in different two states.
[0028] Preferably, the conversion relationship of the point cloud posture of the target object relative to the tool coordinate system is calculated by teaching based on the pre-acquired hand-eye transformation matrix, comprising:
[0029] When the eye is on the hand, the conversion relationship of the point cloud posture of the target object relative to the tool coordinate system is calculated by teaching as follows:
[0030] ;
[0031] When the eye is outside the hand, the conversion relationship of the point cloud posture of the target object relative to the tool coordinate system is calculated by teaching as follows: ;
[0032] wherein, a transformation matrix of a target object point cloud coordinate system to a tool coordinate system, a transformation matrix of a target object point cloud coordinate system to a robot base coordinate system, a transformation matrix of a robot base coordinate system to a tool coordinate system, a transformation matrix of a target object cloud coordinate system to a sensor coordinate system, a transformation matrix of an end coordinate system to a robot base coordinate system when taking a photo;
[0033] The conversion relationship and the accurate transformation matrix are used to calculate the grasping position coordinates of the target object in the robot base coordinate system, comprising:
[0034] When the eye is on the hand, the grasping position is calculated by the following formula:
[0035] ;
[0036] When the eye is outside the hand, the grasping position is calculated by the following formula:
[0037] ;
[0038] wherein, a transformation matrix of a robot base coordinate system to an end coordinate system when taking a photo, a transformation matrix from the coordinate system of the robot arm end to the coordinate system of the sensor, a transformation matrix from the coordinate system of the sensor to the coordinate system of the target point cloud, a transformation matrix from the coordinate system of the robot arm base to the coordinate system of the sensor.
[0039] Preferably, the target point cloud of the target object is obtained in the following manner:
[0040] The sensor is used to obtain raw point cloud data. When the eye is on the hand, the sensor is arranged at the end of the robot arm. When the eye is outside the hand, the sensor is arranged at a fixed position outside the robot arm. The target object is extracted from the raw point cloud data to obtain the target point cloud.
[0041] Preferably, the target object is extracted from the raw point cloud data to obtain the target point cloud, including:
[0042] According to the preset XYZ direction maximum value, the point cloud data is filtered to filter out the redundant points in the scene and retain the point cloud data in the region of interest;
[0043] The normal vector estimate value of each retained point in its neighborhood is obtained, and the point cloud is filtered using the normal vector information;
[0044] For the filtered point cloud data, points within a predetermined range are divided into multiple connected clusters. From the multiple connected clusters obtained by clustering segmentation, points that meet the characteristics of the target object are filtered and extracted as the target point cloud;
[0045] Preferably, for the filtered point cloud data, points within a predetermined range are divided into multiple connected clusters, including:
[0046] S131, select a point as a starting point;
[0047] S132, mark the starting point as visited and add it to a separate target cluster;
[0048] S133, select points within a predetermined range as candidate points with the point as the center, and filter out unvisited points to join the target cluster;
[0049] S134, for each point in the target cluster, repeat S133 until all points are visited;
[0050] S135, select a remaining unvisited point as a new starting point, and repeat S132 until all points are clustered to form multiple connected clusters.
[0051] Preferably, the point cloud matching of the target point cloud using the point cloud registration model based on the deep learning network obtains an initial matching position, including:
[0052] constructing a deep learning network;
[0053] obtaining a data set, the data set being a preset number of sample target object point clouds collected by a sensor device, then performing denoising and downsampling on the obtained target object point clouds to obtain a point cloud registration model;
[0054] using the data set as a training set, training the deep learning network to obtain a point cloud registration model, and inputting the target object point cloud into the point cloud registration model to obtain an initial matching position.
[0055] In a second aspect, the present application also provides an object grasping system based on point cloud deep learning, comprising an electronic device and a point cloud acquisition device, the electronic device comprising a memory and at least one processor, the memory storing a computer program, and the at least one processor being configured to implement the method of any one of the first aspect when executing the computer program; the point cloud acquisition device is used to obtain a target object point cloud of a target object and transmit it to the electronic device.
[0056] In combination with the above technical solutions and solved technical problems, the present application provides an object grasping method based on point cloud deep learning. First, 3D data (point cloud) of a target object and a coordinate conversion relationship (hand-eye transformation matrix) between a sensor and an end of a mechanical arm are obtained. Then, a deep learning network is used to match the point cloud of the target object to obtain an initial matching position, and a nearest iteration algorithm is used to optimize the position to obtain an accurate transformation matrix of the target object in a camera coordinate system relative to a template. Then, based on the hand-eye transformation matrix, a teaching is used to calculate a conversion relationship of a point cloud posture of the target object in a tool coordinate system. Finally, according to the conversion relationship and the accurate transformation matrix, a specific grasping position coordinate of the target object in a mechanical arm base coordinate system is calculated to guide the mechanical arm to move to the position to complete grasping.
[0057] The technical scheme to be protected has the advantages and positive effects that: on the one hand, the introduction of the hand-eye transformation matrix solves the coordinate conversion problem between visual information and the motion control of the robot arm, and through the point cloud registration of the deep learning network, the complex features of the target object can be learned, and compared with the traditional method, the position of the target object can be determined more accurately. Subsequent optimization combined with the iterative algorithm further improves the matching accuracy, thereby obtaining more accurate target object position and attitude information, effectively reducing the deviation caused by inaccurate matching, and improving the success rate and accuracy of the grabbing. On the other hand, the attitude of the target object in the camera coordinate system is converted into the attitude in the robot arm base coordinate system, so that the robot arm can understand and utilize this information to plan the grabbing action. The whole process realizes the automatic grabbing process from data acquisition, point cloud registration, position calculation to subsequent robot arm grabbing guidance, not only reduces manual intervention and complex parameter setting, improves operation efficiency and reliability, but also makes the method suitable for different types of object and scene grabbing tasks. BRIEF DESCRIPTION OF DRAWINGS
[0058] The present application will be further described below in conjunction with the drawings and embodiments.
[0059] Figure 1 is a flowchart of an object grabbing method based on point cloud deep learning provided by an embodiment of the present application.
[0060] Figure 2 is a flowchart of hand-eye transformation matrix acquisition provided by an embodiment of the present application.
[0061] Figure 3 is a flowchart of hand-eye transformation matrix acquisition provided by an embodiment of the present application.
[0062] Figure 4 is a flowchart of target object point cloud acquisition provided by an embodiment of the present application.
[0063] Figure 5 is a flowchart of connected cluster acquisition provided by an embodiment of the present application.
[0064] Figure 6 is a flowchart of initial matching position acquisition provided by an embodiment of the present application.
[0065] Figure 7 is a structural block diagram of an electronic device provided by an embodiment of the present application.
[0066] Figure 8 is a schematic diagram of the calibration board picture of 21 sets of calibration data and the acquisition pose provided by an embodiment of the present application.
[0067] Figure 9a and Figure 9bA coordinate system distribution diagram of the eye-in-hand and the eye-out-of-hand cases is provided by the embodiment of the application. DETAILED DESCRIPTION
[0068] The application will be further described below in conjunction with the drawings and the specific embodiments. It should be noted that the embodiments described below or the technical features between the embodiments can be combined in any manner to form new embodiments without conflict. The embodiments of the application will be described below with reference to the drawings and preferred embodiments, and those skilled in the art can easily understand other advantages and effects of the application from the disclosure. The application can also be implemented or applied by different specific implementation procedures, and the details in the specification can be modified or changed based on different views and applications without departing from the spirit of the application. It should be understood that the preferred embodiments are only for illustrating the application, and are not intended to limit the protection scope of the application.
[0069] The technical field and related terms of the embodiments of the application will be briefly described below to facilitate understanding by those skilled in the art.
[0070] The object grasping system is an automatic technical device combining mechanical arms, sensors, intelligent algorithms and other technologies, which can automatically identify, locate and grasp target objects. The point cloud data of the object can be obtained by a 3D sensor, and the position and attitude of the object are determined by analyzing and processing the point cloud.
[0071] The more common technical solutions are a 2D vision-based grasping system and a traditional 3D point cloud-based grasping system. The existing 2D grasping method is suitable for planar object grasping, but it cannot obtain the RX and RY attitudes due to the lack of depth information. The object coordinates obtained are used for grasping, and the efficiency and accuracy are low. Compared with the 2D vision technology, the 3D point cloud method in the related art can obtain the RX and RY attitudes through depth information, but it depends highly on the completeness of the target object point cloud. Changes in the environment light will cause imaging loss, and the object coordinates obtained are used for grasping, which has a large error.
[0072] To solve the above technical problems, the application provides an object grasping method based on point cloud deep learning, which uses a 3D sensor to obtain scene data, performs point cloud matching based on a deep learning method, and obtains object coordinates that can be used to realize high-precision stable grasping. The method will be described first, and then the system will be described.
[0073] Method embodiment.
[0074] Reference Figure 1 , Figure 1 A flowchart of the object grasping method based on point cloud deep learning is provided by the embodiment of the application.
[0075] The embodiment of the present application provides an object grasping method based on point cloud deep learning, and the method comprises S100 to S300:
[0076] S100, acquiring target object point cloud of a target object, performing point cloud matching on the target object point cloud by using a point cloud registration model based on a deep learning network to obtain an initial matching position, and then optimizing the matching position by using a nearest iteration algorithm to obtain an accurate transformation matrix of the target object relative to a template in a camera coordinate system;
[0077] In some embodiments, the target object point cloud of the target object is acquired in the following manner:
[0078] Raw point cloud data is acquired by using a sensor, the sensor is arranged at the end of the mechanical arm when the eye is on the hand, and the sensor is arranged at a fixed position outside the mechanical arm when the eye is off the hand; target object extraction is performed on the raw point cloud data to acquire the target object point cloud.
[0079] According to different application requirements and mechanical arm structures, the sensor (for example, a 3D camera) is arranged at the end of the mechanical arm (the eye is on the hand) or a fixed position outside the mechanical arm (the eye is off the hand). When the "eye on the hand" mode is used, the sensor moves together with the end of the mechanical arm, can acquire point cloud data of objects near the end of the mechanical arm in real time, and is suitable for a scene in which the mechanical arm needs to flexibly grasp objects at different positions; and when the "eye off the hand" mode is used, the sensor is fixed and does not move, provides a stable view angle, and can acquire point cloud data of the whole working area, and is suitable for a task of grasping objects in a fixed working area.
[0080] Referring to Figure 4 , Figure 4 FIG. 1 is a flowchart of a process for acquiring target object point cloud provided by the embodiment of the present application.
[0081] In some embodiments, the target object extraction is performed on the raw point cloud data to acquire the target object point cloud, and the target object extraction comprises the following steps:
[0082] S110, filtering point cloud data according to preset XYZ direction maximum values, filtering out redundant points in a scene, and retaining point cloud data in a region of interest;
[0083] S120, acquiring normal vector estimation values of each point in a neighborhood of the point, and filtering point cloud by using the normal vector information;
[0084] S130, dividing points with a distance in a preset range into a plurality of connected clusters for the filtered point cloud data; and screening and extracting point cloud meeting target object characteristics from the plurality of connected clusters obtained by clustering segmentation as the target object point cloud.
[0085] The coordinate filtering (S110) refers to filtering the point cloud data according to the set XYZ direction extreme value, filtering out the redundant points in the scene, and retaining the point cloud in the region of interest. In the normal filtering process (S120), the normal vector is understood as a vector perpendicular to the point cloud surface, which can represent the local characteristics of the point cloud surface; the normal vector of each point in its neighborhood is estimated, and the current point cloud is filtered by using the normal vector information, which can solve the point cloud sticking problem caused by the shooting angle. In the clustering segmentation process (S130), points within a certain range are divided into multiple connected clusters, and the specific process is as follows: a point is selected as a starting point, the starting point is marked as visited, and is added to a cluster; the point is selected as the center, points within a certain range are selected as candidate points, and unvisited points are screened and added to the cluster; for each point in the cluster, the above operation is repeated until all points are visited; an unvisited point is selected, and the above steps are repeated until all points are clustered.
[0086] The advantage of the embodiment is that the target object extraction steps such as coordinate filtering, normal filtering and clustering segmentation can quickly and effectively filter out the target object point cloud from a large amount of original point cloud data, reduce the data processing amount and calculation time, and enable the mechanical arm to continuously and stably complete the grasping task of multiple target objects.
[0087] Referring to Figure 5 , Figure 5 is a flowchart for obtaining a connected cluster provided by the embodiment of the application.
[0088] In some embodiments, the points within a predetermined range are divided into multiple connected clusters (S130) for the filtered point cloud data, including:
[0089] S131, a point is selected as a starting point;
[0090] S132, the starting point is marked as visited, and is added to a separate target cluster;
[0091] S133, the point is selected as the center, points within a predetermined range are selected as candidate points, and unvisited points are screened and added to the target cluster;
[0092] S134, for each point in the target cluster, S133 is repeated until all points are visited;
[0093] S135, a remaining unvisited point is selected as a new starting point, S132 is repeated until all points are clustered, and multiple connected clusters are formed.
[0094] Only when the distance between two points is less than or equal to the preset range, the points are considered to be in the neighborhood and are divided into the same connected cluster. The setting of the preset range can be determined according to the size of the target object, the resolution of the point cloud data, and the distance between objects in the actual application scenario, to ensure that the points belonging to the same object can be reasonably gathered together, while avoiding the points of different objects being mistakenly merged into a cluster.
[0095] The advantage of the embodiment is that by dividing the filtered point cloud data into multiple connected clusters, the specific position and shape of the target object in the point cloud can be accurately identified. Each connected cluster represents a possible object or a local area of an object, so that the robot arm can accurately position the target object according to the geometric features and position information of the connected cluster.
[0096] Since each connected cluster is formed based on the neighborhood relationship and distance threshold between points, background point clouds or noise point clouds that do not match the characteristics of the target object will be divided into different clusters or filtered out separately, ensuring that the robot arm only performs grasping operations on connected clusters that match the characteristics of the target object, avoiding the phenomenon of false recognition of background or other objects, and improving the grasping efficiency.
[0097] Referring to Figure 6 , Figure 6 is a flowchart of an embodiment provided by the present application for obtaining an initial matching position.
[0098] In some embodiments, the point cloud matching of the target object point cloud using the point cloud registration model based on the deep learning network obtains an initial matching position, including:
[0099] S140, constructing a deep learning network;
[0100] S150, obtaining a data set, the data set being a preset number of sample target object point clouds collected by a sensor device, and then being obtained after the target object point cloud is traversed and denoised and down-sampled;
[0101] S160, taking the data set as a training set, training the deep learning network to obtain a point cloud registration model, and inputting the target object point cloud into the point cloud registration model to obtain an initial matching position.
[0102] The deep learning network can be a REGTR deep learning network. The trained deep learning network can be used to match the actually obtained point cloud with the template point cloud to achieve registration. During the training process, a certain number and scene of initial target object point cloud data are collected by a sensor device. The collected initial target object point cloud data are traversed, denoised, and down-sampled to remove noise in the environment, and pre-processed to form a data set. The data set is put into the REGTR deep learning network for training to obtain a trained model file.
[0103] Extracting down-sampling key points from the input point cloud and , and the corresponding features are , . These key points and features are transmitted to the transformer cross-encoding layer. The transformer layer includes multi-head self-attention and cross-attention, which facilitates global information aggregation. The transformer encoder outputs the feature , . The transformed position of the output key point is , . The correspondence can be obtained from and , that is, , and the other direction is the same. Meanwhile, the overlap score , is output, which represents the probability of each key point being located in the overlap region. The required transformation matrix is estimated directly from the correspondence in the overlap region.
[0104] The template point cloud (i.e., template) refers to pre-defined and saved point cloud data that describes the shape, structure, and features of the target object, which is equivalent to the blueprint of the target object. When an object needs to be grasped, the actually obtained point cloud data is compared with the template point cloud to find the most matching part, thereby determining the position and pose of the object in space. The creation of the template point cloud is based on the accurate scanning of the target object, ensuring that it can accurately reflect the geometric characteristics of the object. In practical applications, by matching the template point cloud with the actually obtained point cloud data, the target object can be quickly and accurately identified.
[0105] Therefore, by using the trained point cloud registration model to perform point cloud matching on the target object point cloud, the initial matching position of the point cloud can be obtained. The ICP (Iterative Closest Point) algorithm is used for matching to calculate the accurate position. By using the ICP algorithm, the accurate transformation matrix can be quickly obtained while ensuring the matching accuracy, so that the grasping position coordinates can be calculated in time.
[0106] S200, based on the pre-acquired hand-eye transformation matrix, the point cloud posture of the target object is calculated through teaching to obtain the conversion relationship relative to the tool coordinate system; the hand-eye transformation matrix is used to indicate the transformation relationship between the sensor coordinate system and the end coordinate system of the robot arm; the tool coordinate system is the pose of the end of the robot arm, such as the tool of the gripper of the robot arm, which is described and defined by Cam.
[0107] The hand-eye transformation matrix is obtained by solving the hand-eye transformation matrix according to the motion transformation of the end of the robot arm and the motion transformation observed by the sensor.
[0108] Referring to Figure 2 , Figure 2 is a flowchart of a hand-eye transformation matrix acquisition method provided by the embodiment of the application.
[0109] In some embodiments, the hand-eye transformation matrix is obtained by:
[0110] S210, camera intrinsic calibration is performed, corner points are extracted and refined to sub-pixel accuracy, and Zhang's calibration algorithm is used to obtain camera internal parameters and distortion coefficients as camera calibration results;
[0111] The camera is calibrated by using Zhang's calibration algorithm, then the corner points are extracted and refined to sub-pixel accuracy by shooting the calibration board image, and the internal parameters (such as focal length) and distortion coefficients of the camera are obtained.
[0112] S220, image distortion correction is performed, and the camera calibration results are used to correct the image, and the sub-pixel accuracy corner point coordinates are re-extracted;
[0113] The camera intrinsic parameters and distortion coefficients obtained by calibration are used to correct the image, and the sub-pixel accuracy corner point coordinates of the corrected image can be re-extracted.
[0114] S230, define the world coordinate system on the calibration board, and obtain the 3D coordinates of the corner points;
[0115] The world coordinate system is defined on the calibration board, and the accurate 3D coordinates of the corner points are obtained. This step establishes a physical coordinate system through the calibration board.
[0116] S240, a perspective projection model is established using a PnP algorithm, and a rotation and translation matrix from the world coordinate system to the camera coordinate system is solved and used as the camera pose;
[0117] Specifically, a perspective projection model can be established based on the obtained corner 2D image coordinates and corresponding 3D world coordinates by using a PnP algorithm. A rotation matrix and a translation matrix from the world coordinate system to the camera coordinate system are solved by the perspective projection model, that is, the pose of the camera is obtained. The camera pose reflects the position and direction of the camera in space, which is used for subsequent hand-eye calibration.
[0118] In S250, a hand-eye transformation matrix is obtained by solving equations based on the robot pose and the camera pose.
[0119] The motion transformation data of the robot end at different poses and the corresponding camera pose data can be considered as containing the relative motion relationship of the robot and the camera in space. The robot pose and camera pose data are substituted into the equation model of hand-eye calibration (such as AX = XB), and the hand-eye transformation matrix is obtained by mathematical solving method. The matrix describes the geometric transformation relationship between the sensor coordinate system and the robot end coordinate system, which is used to realize the combination of visual information and robot motion control.
[0120] The advantage of the embodiment is that the camera intrinsic parameter calibration and distortion correction, and the PnP algorithm for calculating the camera pose ensure the accuracy and reliability of the visual information.
[0121] Referring to Figure 3 , Figure 3 is a flowchart of obtaining a hand-eye transformation matrix provided by an embodiment of the application.
[0122] In some embodiments, the hand-eye transformation matrix is obtained by solving equations based on the robot pose and the camera pose (S250), including:
[0123] In S251, a hand-eye calibration matrix is obtained by a conversion relationship based on the robot pose and the camera pose.
[0124] In S252, the accuracy value of the hand-eye calibration matrix is evaluated by calculating the re-projection error, and the hand-eye calibration matrix is taken as the hand-eye transformation matrix when the accuracy value meets the preset accuracy requirement.
[0125] It can be understood that a hand-eye calibration equation (such as AX = XB) is constructed based on the motion transformation (A) of the robot end and the motion transformation (B) observed by the sensor. Wherein, X is the hand-eye transformation matrix to be solved. The robot pose data (describing the coordinate information of the robot end at different positions and attitudes) and the camera pose data (the rotation and translation matrix from the world coordinate system to the camera coordinate system obtained by the PnP algorithm) are substituted into the above equation, and the hand-eye calibration matrix is obtained by mathematical method.
[0126] In some embodiments, the accuracy value of the hand-eye calibration matrix is evaluated by calculating the re-projection error, including:
[0127] Two sets of pose data are used to estimate the camera pose of one set as the true value, and the pose and 3D point set are reprojected to get the image plane 2D point set coordinates. The mean square error value (MSE error value) between the projected point coordinates and the corner point coordinates is calculated as the accuracy value.
[0128] Two sets of pose data are selected, one set as the true value reference, and the other set used to estimate the camera pose. The two sets of data contain the coordinate information of the robot arm and camera in different positions and attitudes. The pose data and 3D point set are converted from one coordinate system to another coordinate system using the hand-eye calibration matrix, and are reprojected to the image plane to get the 2D point set coordinates, completing the re-projection. The difference between the 2D point set coordinates obtained by re-projection and the actual corner point coordinates is compared, and the mean square error (MSE) method is used to calculate the error value between the projected point coordinates and the corner point coordinates, i.e. the re-projection error. This error reflects the accuracy of the hand-eye calibration matrix, and the smaller the error, the more accurate the matrix in describing the geometric relationship between the sensor and the end of the robot arm.
[0129] The benefits of this embodiment are that by constructing equations and combining the pose data of the robot arm and camera, the hand-eye calibration matrix can be accurately solved. The re-projection error evaluation and optimization steps further ensure the high accuracy of the matrix, allowing the robot arm to accurately adjust its attitude based on visual information and achieve high-precision grasping of objects, effectively reducing the grasping deviation caused by hand-eye coordinate system conversion errors. In addition, using the re-projection error as an accuracy evaluation indicator directly reflects the accuracy of the hand-eye transformation matrix in predicting the position and attitude of objects in actual applications.
[0130] In some embodiments, when the sensor is fixed at the end of the robot arm, the hand-eye transformation matrix is obtained by eye-in-hand calibration, and the conversion relationship is as follows:
[0131] ;
[0132] When the sensor holder is fixed at a position outside the robot arm, the hand-eye transformation matrix is obtained by eye-in-hand calibration, and the conversion relationship is as follows:
[0133] ;
[0134] wherein, is the transformation matrix from the target object coordinate system to the sensor coordinate system, is the transformation matrix from the sensor coordinate system to the robot arm end coordinate system, is the transformation matrix from the robot arm end coordinate system to the robot arm base coordinate system, is the transformation matrix from the sensor coordinate system to the robot arm base coordinate system, is the transformation matrix from the base coordinate frame of the robot arm to the end coordinate frame of the robot arm. T 0 and T 1 represent the transformation matrix of (the end of the robot arm or the target object) in two different states (position and attitude), is the transformation matrix from the target object coordinate frame to the sensor coordinate frame in the first state; is the inverse transformation matrix from the target object coordinate frame to the sensor coordinate frame in the second state; is the inverse transformation matrix from the end coordinate frame of the robot arm to the base coordinate frame of the robot arm in the first state; is the transformation matrix from the end coordinate frame of the robot arm to the base coordinate frame of the robot arm in the second state; is the inverse transformation matrix from the base coordinate frame of the robot arm to the end coordinate frame of the robot arm in the first state; is the transformation matrix from the base coordinate frame of the robot arm to the end coordinate frame of the robot arm in the second state.
[0135] It can be considered that the hand-eye calibration problem can be represented as the equation AX=XB, where A is the motion transformation of the end of the robot arm, B is the motion transformation observed by the sensor, and X is the hand-eye transformation matrix to be solved. For the two different ways of "eye on hand" and "eye off hand", the form of the conversion relationship is different, which needs to be processed respectively.
[0136] For the case of eye on hand, the coordinate frame distribution is as shown in Figure 9a , where the coordinate frame is only a schematic diagram and does not represent the actual coordinate frame direction. The sensor is fixed to the end of the robot arm and moves with the robot arm. The calibration board is fixed to the object plane, and the calibration image is obtained by moving the robot arm. The target is to solve the relationship between the sensor coordinate frame (cam) and the end coordinate frame (end). The conversion relationship is .
[0137] is the transformation matrix from the target object coordinate frame (tar) to the base coordinate frame (base) of the robot arm;
[0138] is the transformation matrix from the target object coordinate frame to the sensor coordinate frame;
[0139] is the transformation matrix from the sensor coordinate frame to the end coordinate frame of the robot arm;
[0140] is the transformation matrix from the end coordinate frame of the robot arm to the base coordinate frame of the robot arm.
[0141] The relative position of the calibration object (tar, i.e. the target object) to the base of the robot arm is unchanged, so is fixed, thus can be eliminated by two sets of pose relationship, i.e. , and is arranged as , i.e. , wherein is the transformation of the target object coordinate system to the camera coordinate system in different poses, is the transformation of the robot end coordinate system to the robot base coordinate system in different poses.
[0142] For the case of eye outside the hand, the coordinate system distribution is shown in Figure 9b , where the coordinate system is only a schematic diagram and does not represent the actual coordinate system direction. The sensor is fixed at a fixed position other than the end of the robot arm and does not move with the robot arm. The calibration plate is fixed to the end of the robot arm, and calibration images are obtained by moving the robot arm. The target is to solve the transformation matrix of the robot base coordinate system to the camera. The conversion relationship is .
[0143] is the transformation matrix of the target object coordinate system to the robot end coordinate system;
[0144] is the transformation matrix of the target object coordinate system to the sensor coordinate system;
[0145] is the transformation matrix of the sensor coordinate system to the robot base coordinate system;
[0146] is the transformation matrix of the robot base coordinate system to the robot end coordinate system.
[0147] The relative position of the calibration object to the end of the robot arm is fixed, thus is fixed, thus can be eliminated by two sets of pose relationship, i.e. . Arranged as , i.e. , wherein is the transformation of the target object coordinate system to the camera coordinate system in different poses, is the transformation of the robot end coordinate system to the robot base coordinate system in different poses.
[0148] Therefore, whether the eye is outside the hand or inside the hand, the hand-eye calibration problem is converted to solve , and (hand-eye transformation matrix).
[0149] In some embodiments, the conversion relationship of the point cloud pose of the target object relative to the tool coordinate system is calculated by teaching based on the pre-acquired hand-eye transformation matrix, comprising:
[0150] When the eye is on the hand, the transformation relationship of the point cloud pose of the target object relative to the tool coordinate system is calculated by the following formula:
[0151] ;
[0152] When the eye is outside the hand, the transformation relationship of the point cloud pose of the target object relative to the tool coordinate system is calculated by the following formula: .
[0153] Wherein, is the transformation matrix of the target object point cloud coordinate system to the tool coordinate system, is the transformation matrix of the target object point cloud coordinate system to the robot base coordinate system, is the transformation matrix of the robot base coordinate system to the tool coordinate system, is the transformation matrix of the target object point cloud coordinate system to the sensor coordinate system, is the transformation matrix of the end coordinate system to the robot base coordinate system when taking a photo.
[0154] S300, according to the transformation relationship and the accurate transformation matrix, the grasping position coordinates of the target object in the robot base coordinate system are calculated; the grasping position coordinates are used to guide the robot to move to the grasping position to complete grasping.
[0155] In some embodiments, the calculation of the grasping position coordinates of the target object in the robot base coordinate system according to the transformation relationship and the accurate transformation matrix (S300) comprises:
[0156] When the eye is on the hand, the grasping position is calculated by the following formula:
[0157]
[0158] When the eye is outside the hand, the grasping position is calculated by the following formula:
[0159] ;
[0160] In the above formula, is the transformation matrix of the robot base coordinate system to the end coordinate system when taking a photo, is the transformation matrix of the robot end coordinate system to the sensor coordinate system, is the transformation matrix of the sensor coordinate system to the target object point cloud coordinate system, is the transformation matrix of the robot base coordinate system to the sensor coordinate system.
[0161] As an example, taking a workpiece grasping as an example, the workpiece is grasped by using the above object grasping method, comprising:
[0162] Determination of hand-eye calibration results, taking the eye on the hand as an example, 21 sets of calibration data, according to the calibration board picture and the acquisition pose. Referring to Figure 8 , Figure 8 is the schematic diagram of the 21 sets of calibration data provided by the embodiment of the application.
[0163] Calibration results are:
[0164]
[0165] Demonstration grasping matrix is:
[0166]
[0167] Demonstration shooting matrix is:
[0168]
[0169] The calculated demonstration result matrix is:
[0170]
[0171] Get the target object matching conversion matrix :
[0172]
[0173] Get the final grasping position:
[0174] Through calculation, the final grasping position is
[0175]
[0176] In specific applications, the above technical solutions can be applied to industrial production lines, and different shapes and sizes of parts can be recognized and grasped to meet the needs of different industrial automatic production lines.
[0177] The technical scheme provided by the method embodiment is to first acquire 3D data (point cloud) of a target object and a coordinate conversion relationship (hand-eye transformation matrix) between a sensor and an end of a mechanical arm. Then, a deep learning network is used to match the point cloud of the target object to obtain an initial matching position, and the position is optimized through a nearest iteration algorithm to obtain an accurate transformation matrix of the target object relative to a template in a camera coordinate system. Based on the hand-eye transformation matrix, the conversion relationship of the point cloud posture of the target object in the tool coordinate system is calculated through teaching. Finally, according to the conversion relationship and the accurate transformation matrix, the specific grasping position coordinates of the target object in the base coordinate system of the mechanical arm are calculated to guide the movement of the mechanical arm to the position to complete grasping.
[0178] The beneficial effects of the method embodiment are as follows: on the one hand, the introduction of the hand-eye transformation matrix solves the coordinate conversion problem between visual information and mechanical arm motion control, and the point cloud registration is performed through the deep learning network, which can learn the complex features of the target object, and compared with the traditional method, the position of the target object can be more accurately determined. The subsequent optimization of the nearest iteration algorithm further improves the matching accuracy, thereby obtaining more accurate target object position and posture information, effectively reducing the grasping deviation caused by inaccurate matching, and improving the success rate and accuracy of grasping. On the other hand, the posture of the target object in the camera coordinate system is converted into the posture in the base coordinate system of the mechanical arm, so that the mechanical arm can understand and utilize the information to plan the grasping action. The entire process realizes an automatic grasping process from data acquisition, point cloud registration, position calculation to subsequent mechanical arm grasping guidance, not only reduces manual intervention and complex parameter setting, improves operation efficiency and reliability, but also makes the method suitable for grasping tasks of different types of objects and scenes.
[0179] System embodiment.
[0180] The embodiment provides an object grasping system based on point cloud deep learning, and specific embodiments and technical effects thereof are consistent with those described in the above method embodiments, and some contents will not be described again.
[0181] The object grasping system comprises an electronic device and a point cloud acquisition device. The electronic device comprises a memory and at least one processor, the memory stores a computer program, and the at least one processor is configured to implement the method described in any one of the method embodiments when executing the computer program. The point cloud acquisition device is used to acquire target object point cloud of a target object and transmit the target object point cloud to the electronic device. The point cloud acquisition device comprises, for example, a 3D camera and a bracket for fixing the 3D camera.
[0182] Reference Figure 7 , Figure 7is a structural block diagram of an electronic device provided by an embodiment of the present application. The electronic device 10 may, for example, include at least one memory 11, at least one processor 12, and a bus 13 connecting different platform systems.
[0183] The memory 11 can include a (computer) readable medium in the form of volatile memory, such as a random access memory (RAM) 111 and / or a cache memory 112, and can further include a read-only memory (ROM) 113.
[0184] The memory 11 further stores a computer program, which can be executed by the processor 12, so that the processor 12 implements the steps of any of the above methods.
[0185] The memory 11 can further include a utility 114 having at least one program module 115, such as an operating system, one or more application programs, other program modules, and program data, each of or a combination of which can include the implementation of a network environment.
[0186] Correspondingly, the processor 12 can execute the above computer program and can execute the utility 114.
[0187] The processor 12 can employ one or more application specific integrated circuits (ASICs), DSPs, programmable logic devices (PLDs), complex programmable logic devices (CPLDs), field programmable gate arrays (FPGAs), or other electronic elements.
[0188] The bus 13 can represent one or more of several types of bus structures, including a memory bus or memory controller, a peripheral bus, a graphics acceleration port, a processor or local bus using various bus architectures, or any bus structure that supports the transfer of data between different components of a device.
[0189] The electronic device 10 can also communicate with one or more external devices such as a keyboard or a pointing device, a Bluetooth device, etc., and can also communicate with one or more devices that enable interaction with the electronic device 10 (for example, a router, a modem, etc.), and / or enable communication with one or more other computing devices. Such communication can occur via the input / output interface 14. Further, the electronic device 10 can communicate with one or more networks (such as a local area network (LAN), a wide area network (WAN), and / or the public network, such as the Internet) via the network adapter 15. The network adapter 15 can communicate with the other modules of the electronic device 10 via the bus 13. It should be appreciated that although not shown, other hardware and / or software modules could be used in conjunction with the electronic device 10 in some embodiments, including but not limited to: microcode, device drivers, redundant processing units, external disk drive arrays, RAID systems, tape drives, and data archival storage systems, etc.
[0190] It should be noted that in the embodiments of the present application, "at least one" means one or more, and "multiple" means two or more. "And / or" describes the association relationship of the associated objects, which means that there can be three relationships, for example, A and / or B, which can represent the following three cases: A exists alone, A and B exist together, and B exists alone, where A and B can be singular or plural. The character " / " generally represents an "or" relationship between the front and rear associated objects. "At least one of the following" or similar expressions means any combination of these items, including single or multiple combinations of any combination. For example, at least one of a, b or c, can represent a, b, c, a and b, a and c, b and c, or a and b and c, where a, b and c can be single or multiple. It should be noted that "at least one" can also be interpreted as "one or more".
[0191] The terms "first", "second", and the like in the specification and claims of the present application and the above-described drawings are configured to distinguish similar objects, and are not necessarily configured to describe a particular order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances, so that the embodiments of the present application described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "include" and "have" and any variations thereof are intended to cover non-exclusive inclusion, for example, a process, method, system, product or device that includes a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but can include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.
[0192] The application is described from the use purpose, efficiency, progress and novelty, which meets the function improvement and use requirements emphasized by the patent law. The above description and drawings are only the preferred embodiments of the application, and are not limited to the application. Therefore, all similar, identical, equivalent replacement or modification within the scope of the patent application of the application shall be within the scope of the patent application protection of the application.
Claims
1. A method for object grasping based on point cloud deep learning, characterized in that, The method comprises the following steps: obtaining a target object point cloud of a target object, performing point cloud matching on the target object point cloud by using a point cloud registration model based on a deep learning network to obtain an initial matching position, and then optimizing the matching position by using a nearest iteration algorithm to obtain an accurate transformation matrix of the target object relative to a template in a camera coordinate system; calculating a conversion relationship of a point cloud posture of the target object relative to a tool coordinate system by teaching based on a pre-obtained hand-eye transformation matrix; the hand-eye transformation matrix is used to indicate a transformation relationship between a sensor coordinate system and an end coordinate system of a robot arm; solving a grasping position coordinate of the target object in a base coordinate system of the robot arm according to the conversion relationship and the accurate transformation matrix; the hand-eye transformation matrix is obtained in the following manner: solving the hand-eye transformation matrix according to a motion transformation of an end of the robot arm and a motion transformation observed by the sensor; the solving of the hand-eye transformation matrix according to the motion transformation of the end of the robot arm and the motion transformation observed by the sensor comprises the following steps: performing camera intrinsic parameter calibration, extracting a corner point and refining the corner point to sub-pixel accuracy, and obtaining camera internal parameters and distortion coefficients by using a Zhang calibration algorithm as a camera calibration result; performing image distortion correction, removing distortion of an image by using the camera calibration result, and reextracting a sub-pixel accuracy corner point coordinate; defining a world coordinate system on a calibration board and obtaining a 3D coordinate of the corner point; establishing a perspective projection model by using a PnP algorithm, solving a rotation and translation matrix of the world coordinate system to the camera coordinate system, and taking the rotation and translation matrix as a camera pose; obtaining the hand-eye transformation matrix by solving an equation based on the robot arm pose and the camera pose.
2. The object grasping method according to claim 1, characterized by, the obtaining of the hand-eye transformation matrix by solving the equation based on the robot arm pose and the camera pose comprises the following steps: obtaining a hand-eye calibration matrix by a conversion relationship based on the robot arm pose and the camera pose; evaluating an accuracy value of the hand-eye calibration matrix by calculating a re-projection error, and taking the hand-eye calibration matrix as the hand-eye transformation matrix when the accuracy value meets a preset accuracy requirement; the evaluation of the accuracy value of the hand-eye calibration matrix by calculating the re-projection error comprises the following steps: using two sets of pose data, taking one set as a true value to estimate a camera pose of the other set, re-projecting the pose and a 3D point set to obtain a 2D point set coordinate on an image plane, calculating a mean square error value between a projection point coordinate and a corner point coordinate, and taking the mean square error value as the accuracy value.
3. The object grasping method according to claim 1, characterized by, when the sensor is fixed on the end of the robot arm, the hand-eye transformation matrix is obtained by an eye-on-hand calibration manner, and the conversion relationship is as follows: when a sensor frame is fixed at a position outside the robot arm, the hand-eye transformation matrix is obtained by an eye-out-of-hand calibration manner, and the conversion relationship is as follows: wherein, Ttarget2sensor is a transformation matrix from the target coordinate system to the sensor coordinate system, Tsensor2end is a transformation matrix from the sensor coordinate system to the end-of-arm coordinate system, Tend2base is a transformation matrix from the end-of-arm coordinate system to the base-of-arm coordinate system, Tsensor2base is a transformation matrix from the sensor coordinate system to the base-of-arm coordinate system, Tbase2end is a transformation matrix from the base-of-arm coordinate system to the end-of-arm coordinate system; T 0 and T 1 indicates the transformation matrix in the different two states.
4. The object grasping method according to claim 3, characterized by, the calculation of the conversion relationship of the point cloud posture of the target object relative to the tool coordinate system by teaching based on the pre-obtained hand-eye transformation matrix comprises the following steps: when the eye is on the hand, the conversion relationship of the point cloud posture of the target object relative to the tool coordinate system is calculated by teaching by using the following formula: ; When the eye is outside the hand, the conversion relationship of the point cloud posture of the target object relative to the tool coordinate system is shown by the following formula: ; wherein, is a transformation matrix from the target object point cloud coordinate system to the tool coordinate system, is a transformation matrix from the target object point cloud coordinate system to the robot base coordinate system, is a transformation matrix from the robot base coordinate system to the tool coordinate system, is a transformation matrix from the target object point cloud coordinate system to the sensor coordinate system, is a transformation matrix from the end coordinate system at the time of shooting to the robot base coordinate system; the solving of the grasping position coordinate of the target object in the base coordinate system of the robot arm according to the conversion relationship and the accurate transformation matrix comprises the following steps: when the eye is on the hand, the grasping position is solved by using the following formula: ; when the eye is out of the hand, the grasping position is solved by using the following formula: ; wherein, is a transformation matrix from the base coordinate system of the robot arm to the end coordinate system at the time of photographing, is a transformation matrix from the end coordinate system of the robot arm to the sensor coordinate system, is a transformation matrix from the sensor coordinate system to the target point cloud coordinate system, is a transformation matrix from the base coordinate system of the robot arm to the sensor coordinate system.
5. The object grasping method according to claim 1, characterized by, the obtaining of the target object point cloud of the target object comprises the following steps: Raw point cloud data is acquired by using a sensor, the sensor is arranged at the end of a robot arm when the eye is on the hand, and the sensor is arranged at a fixed position outside the robot arm when the eye is outside the hand; target object extraction is performed on the raw point cloud data to obtain target object point cloud.
6. The object grasping method according to claim 5, characterized by, The target object extraction on the raw point cloud data to obtain the target object point cloud comprises: Filtering the point cloud data according to preset XYZ direction maximum values, filtering out redundant points in the scene, and retaining point cloud data in the region of interest; Obtaining a normal vector estimation value of each retained point in its neighborhood, and filtering the point cloud by using the normal vector information; For the filtered point cloud data, points within a preset range are divided into multiple connected clusters; target object point cloud is extracted from the multiple connected clusters obtained by clustering segmentation and meets the characteristics of the target object; The step of dividing the points within the preset range into multiple connected clusters for the filtered point cloud data comprises: S131, selecting a point as a starting point; S132, marking the starting point as visited and adding it to a separate target cluster; S133, selecting points within a preset range as candidate points with the point as the center, and screening out unvisited points to join the target cluster; S134, for each point in the target cluster, repeating S133 until all points are visited; S135, selecting a remaining unvisited point as a new starting point, repeating S132 until all points are clustered to form multiple connected clusters.
7. The object grasping method according to claim 6, characterized in that, The step of performing point cloud matching on the target object point cloud by using a point cloud registration model based on a deep learning network to obtain an initial matching position comprises: Constructing a deep learning network; Obtaining a data set, the data set is a preset number of sample target object point clouds collected by a sensor device, then the obtained target object point cloud is traversed and denoised and downsampled to obtain; The data set is used as a training set to train the deep learning network to obtain a point cloud registration model; the target object point cloud is input into the point cloud registration model to obtain an initial matching position.
8. A point cloud deep learning based object grasping system, characterized in that, The electronic device comprises a memory and at least one processor, the memory stores a computer program, and the at least one processor is configured to execute the computer program to implement the method of any one of claims 1-7; the point cloud acquisition device is used to obtain target object point cloud of a target object and transmit it to the electronic device.
Citation Information
Patent Citations
Three-degree-of-freedom parallel robot hand-eye calibration method based on 3D visual sensor
CN111872922A
Hand-eye calibration method based on calibration plate three-dimensional point cloud
CN115861445A