A workpiece pose estimation method and device for unordered sorting scenarios

By using RGB-D images and pose estimation models in disordered sorting scenarios, combining geometric constraints and 3D-3D mapping relationships, the accuracy and real-time problems of workpiece pose estimation in disordered sorting scenarios are solved, and more efficient workpiece sorting is achieved.

CN115359119BActive Publication Date: 2025-05-30SOUTH CHINA UNIV OF TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210824877.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-07-14
Publication Date
2025-05-30
Estimated Expiration
2042-07-14

AI Technical Summary

Technical Problem

In disorderly sorting scenarios, existing workpiece pose estimation methods are difficult to effectively solve the perceptual difficulty caused by target occlusion, lighting changes and complex object stacking, which makes it difficult to meet the accuracy and real-time pose estimation.

Method used

A workpiece pose estimation method for disorderly sorting scenes is adopted. By acquiring RGB-D images, the object detection module is used to obtain the surrounding frame coordinate information, the crop depth map is converted into a point cloud, and the RGB image is input to the pose estimation model to obtain the RGB-point cloud fusion features. Through geometric constraints and 3D-3D mapping relationships, the constraint registration of key points and center points is carried out to output the final pose of the target workpiece.

Benefits of technology

It improves the accuracy and stability of pose estimation, can handle complex situations in disordered sorting scenarios more effectively, and enhances the success rate of workpiece sorting tasks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115359119B_ABST
    Figure CN115359119B_ABST
Patent Text Reader

Abstract

The present invention discloses a workpiece pose estimation method and device for an unordered sorting scenario. The method includes: obtaining an RGB-D image of the scene workpiece; inputting the RGB image into a target detection module to obtain coordinate information; cropping the RGB image and the depth image according to the coordinate information and inputting them into a pose estimation model to obtain an RGB-point cloud fusion feature; obtaining the coordinates of key points and a center point in the camera coordinate system according to the RGB-point cloud fusion feature; performing constraints on the coordinates of the key points and the center point; registering the constrained coordinates of the key points and the center point with the corresponding key point coordinates and center point coordinates in a preset model coordinate system, and outputting the final pose of the target workpiece. The present invention introduces geometric feature constraints of the target object into the model, such as the constraints of predicted key points and the center point, which can improve the accuracy of the prediction result. The present invention can be widely applied to the field of pose recognition technology.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of pose recognition, and in particular to a workpiece pose estimation method and device for an unordered sorting scenario. Background Art

[0002] Workpiece sorting is one of the basic application fields of industrial robots. Currently, industrial robots can already complete workpiece sorting tasks in structured and semi-structured scenarios well. However, in an unordered scenario, due to the unknown and uncertain nature of the scenario, there are mutual stacking and occlusion between targets, which increases the difficulty for the robot to perceive the environment and targets, and reduces the success rate of the robot sorting task.

[0003] The core of the workpiece sorting problem in an unordered scenario lies in the perception of workpiece targets, that is, it is necessary to use visual perception means to achieve the recognition of targets and the estimation of their position and pose. Due to the accuracy problem of sensors, the change of illumination in the scenario, and the stacking of complex objects, the existing methods still cannot well solve the pose estimation problem.

[0004] From the perspective of actual application scenarios, the existing pose estimation methods are mainly of two types: traditional pose estimation based on feature matching and pose estimation method based on deep learning. Among them, the former predicts the pose information of the target by matching the features of the scene point cloud and the target model point cloud. This method requires the establishment of a template library and needs to perform online feature search and matching on a large number of point clouds, with low efficiency and it is difficult to meet the real-time requirements. Especially when there are multiple objects stacked in the scenario, this method causes serious errors in pose estimation due to mis-matching, and then leads to the failure of sorting.

[0005] For the above reasons, deep learning-based methods have received more attention in recent years. Through training on a large amount of data, the required generalization ability is obtained, and the speed of online detection can be achieved by using GPU acceleration. It mainly includes direct regression and two-stage methods. The former refers to the network predicting the target pose and performing end-to-end training on the target pose; the latter first estimates intermediate variables related to geometry and then uses 2D-3D and 3D-3D mapping relationships to obtain the final pose. However, inferring the pose of an object from multi-viewpoint cloud data and images collected by a depth camera essentially uses the information in the projection space to restore the 3D physical space information. The above methods only rely on the network to predict the pose end-to-end and lack explicit global geometric constraints internally. Therefore, the prediction stability of the network is poor. Especially when there is ambiguity in the local features of the object (such as workpieces with high shape regularity and symmetry, the projection result of a single view can be interpreted as multiple three-dimensional structures or poses), the network cannot correctly regress, resulting in incorrect pose estimation, which is unacceptable for highly reliable applications. Secondly, deep learning-based methods require a large amount of data for supervised training. Existing methods manually annotate by taking pictures of the scene images. This process is time-consuming, inefficient, and prone to annotation deviations due to human factors, which in turn causes deviations in network prediction. Summary of the Invention

[0006] To at least partly solve one of the technical problems existing in the prior art, an object of the present invention is to provide a workpiece pose estimation method and device for an unordered sorting scenario.

[0007] The technical solution adopted by the present invention is as follows:

[0008] A workpiece pose estimation method for an unordered sorting scenario includes the following steps:

[0009] Obtain the RGB-D image of the scene workpiece, where the RGB-D image includes an RGB image and a depth image;

[0010] Input the RGB image into the target detection module to obtain the coordinate information of the bounding box of the target workpiece;

[0011] Crop the RGB image and the depth image according to the obtained coordinate information, convert the cropped depth map into a point cloud, and input the point cloud and the cropped RGB image into the pose estimation model to obtain the RGB-point cloud fusion feature;

[0012] Obtain the key point coordinates and the center point coordinates in the camera coordinate system according to the RGB-point cloud fusion feature;

[0013] Constrain the obtained key point coordinates with the true values of the key point coordinates, constrain the center point coordinates with the true values of the center point coordinates, and perform geometric constraints on the edges between the key point coordinates and the center point coordinates;

[0014] Through the 3D-3D mapping relationship, the coordinates of the constrained key points and the center point are registered with the corresponding key point coordinates and center point coordinates in the preset model coordinate system, and the final pose of the target workpiece is output.

[0015] Furthermore, the pose estimation model is obtained through the following training method:

[0016] Obtain rendering data, obtain real data and perform annotation, and obtain a training set and a validation set according to the rendering data;

[0017] Input the RGB-D images in the training set into the target detection module to obtain the rectangular box information of the target;

[0018] Crop the RGB image and the depth image in the RGB-D image according to the rectangular box information, and input the cropped RGB image and depth image into the pose estimation model;

[0019] Use a backpropagation voting mechanism to obtain the key point coordinates, the center point coordinates, and obtain the pose result;

[0020] Use the added geometric constraints to train the pose estimation model;

[0021] Use transfer learning to fine-tune the pose estimation model according to the real data;

[0022] Use the validation set to validate the trained pose estimation model.

[0023] Furthermore, the rendering data and the real annotated data are obtained through the following methods:

[0024] Use the rendering data production toolbox of blenderproc to obtain the rendering data set; the rendering data set consists of two parts. The first part is that the object generates uniformly sampled viewpoints of the Fibonacci circle, and at the same time, a uniform camera spin angle is used at each sampled viewpoint. The second part is the stacked data, where the camera spins randomly at a vertical angle to sample the data, and the scene is transformed with a random background and random lighting; annotate the real data, including:

[0025] Fix the camera at the end of the robotic arm, and obtain the pose of the camera relative to the base coordinate of the robotic arm through the eye-in-hand hand-eye calibration procedure; among them, the camera pose of the real data is obtained through the following formula:

[0026]

[0027] In the formula, is the pose of object j relative to the camera in the i-th frame, is the pose of object j in the camera coordinate system in the 0-th frame, is the inverse matrix of the camera pose of the i-th frame.

[0028] Further, after obtaining the data, the obtained data is filtered to delete target objects with a low visibility ratio.

[0029] Among them, the definition of the visibility ratio is as follows:

[0030]

[0031] In the formula, p total refers to the sum of all pixels when the object is completely visible in a certain perspective, and p show refers to the actually visible pixels.

[0032] Further, the key point coordinates and the center point coordinates are obtained in the following way:

[0033] The pose estimation model regresses to obtain the key point offset and the center point offset based on the target reference point;

[0034] For a specific key point, a prediction set of this key point is obtained. By adopting a backpropagation voting mechanism for the obtained key point set, the key point coordinates are obtained;

[0035] The expression of the backpropagation voting mechanism is as follows:

[0036]

[0037] In the formula, where i is the target reference point in the scene, and j represents the j-th key point, represents the confidence of the target reference point in the j-th key point; is the coordinate value of the j-th key point in the camera coordinate system obtained by voting; is the coordinate value of the j-th key point estimated by the target reference point, and n is the total number of target reference points selected in the scene;

[0038] The confidence calculation formula is as follows:

[0039]

[0040]

[0041] In the formula, is the threshold, represents the set of coordinate values of the j-th key point estimated by the selected target reference point.

[0042] Further, during the training process, a combined loss function is used to compare the difference between the value predicted by the model regression and the true annotation value, and the pose estimation model is trained according to the obtained difference.

[0043] Among them, the combined loss function includes a target segmentation loss function, a key point offset loss function based on the target reference point, a center point offset loss function based on the target reference point, a key point center point edge vector offset loss function, a key point loss function, and a center point loss function.

[0044] Furthermore, the expression of the combined loss function is as follows:

[0045] Loss = w 1 Lk p_v + w 2 L center_v + w 3 L seg + w 4 L kp + w 5 L center + w 6 L center_kp_v

[0046] In the formula, w 1 , w 2 , w 3 , w 4 , w 5 , w 6 are the weights of the loss function;

[0047] Among them, the key point offset loss function is:

[0048]

[0049] In the formula, M is the target key point, N is the randomly selected scene point cloud, Ω k is the set of ambiguous key points, of i jk* is the true offset of the key point, of i j is the actual offset of the i-th point in the scene to the j-th key point. The II formula means that the formula holds only when p i belongs to the points on the target instance I;

[0050] The center point offset loss function is:

[0051]

[0052] In the formula, of i* is the true offset of the center point;

[0053] The target segmentation loss function:

[0054] L seg = -α(1 - q i )γ log(q i ), q i = c i .l i

[0055] where α is the balance hyperparameter, c i represents the confidence that the i-th point belongs to each category, γ represents the attention hyperparameter, l i is the true class label;

[0056] The key point loss function is:

[0057]

[0058] kp jk* represents the true coordinates of the key point; kp j represents the actual coordinates of the key point during the training process;

[0059] The center point loss function is:

[0060]

[0061] where is the true coordinate of the center point, Δx i is the actual coordinate of the center point during the training process;

[0062] The key point center point edge vector offset loss function is:

[0063]

[0064] where cof jk* is the true offset between the key point set and the center point, cof j is the offset between the actual key point set and the center point.

[0065] Furthermore, after obtaining the key point coordinates and the center point coordinates through the voting mechanism, the pose of the target is obtained in the following way:

[0066] The pose estimation model predicts the key point coordinates and the center point coordinates in the camera coordinate system, and at the same time registers with the key points and the center point on the model to obtain the homogeneous transformation matrix of the target object; the pose of the target is obtained through the following formula:

[0067]

[0068] X = x 1 , x 2 , x 3 , …, x n

[0069] P = P1 , P 2 , P 3 , ..., p n

[0070] Among them, X is the set of key points in the camera coordinate system, and P is the set of corresponding key points in the object coordinate system; R is the rotation matrix estimated from the camera to the object coordinate system, t is the translation amount estimated from the camera to the object coordinate system, and N p is the number of corresponding points.

[0071] Furthermore, the verification of the trained pose estimation model using the validation set includes:

[0072] According to the obtained homogeneous transformation matrix, the model point cloud is transformed through the homogeneous transformation matrix, that is, the model point cloud is transformed from the model coordinate system to the camera coordinate system;

[0073] Predict the segmented target scene point cloud through the pose estimation model, calculate the average Euclidean distance between the model point cloud and the target scene point cloud in the camera coordinate system, and judge the accuracy of the pose estimation according to the value of the average Euclidean distance.

[0074] Another technical solution adopted by the present invention is:

[0075] A workpiece pose estimation device for an unordered sorting scenario includes:

[0076] At least one processor;

[0077] At least one memory for storing at least one program;

[0078] When the at least one program is executed by the at least one processor, the at least one processor implements the above method.

[0079] The beneficial effect of the present invention is that the present invention introduces geometric feature constraints of the target object in the model, such as the constraints of the predicted key points and the center point, which can improve the accuracy of the prediction results. Description of the Drawings

[0080] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following introduces the relevant technical solution drawings in the embodiments of the present invention or the prior art. It should be understood that the drawings introduced below only facilitate the clear expression of some embodiments of the technical solutions in the present invention. For those skilled in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0081] Figure 1 is the flowchart of the steps of a workpiece pose estimation method for an unordered sorting scenario in an embodiment of the present invention;

[0082] Figure 2 is the flowchart for creating real data annotations in the embodiments of the present invention;

[0083] Figure 3 is the flowchart of the pose estimation module in the embodiments of the present invention;

[0084] Figure 4 is the visualization diagram of Fibonacci sphere sampling of the rendered dataset in the embodiments of the present invention;

[0085] Figure 5 is the visualization result diagram in the embodiments of the present invention. Detailed implementation manners

[0086] The embodiments of the present invention will be described in detail below. The examples of the embodiments are shown in the accompanying drawings, where the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below by referring to the accompanying drawings are exemplary and are only used to explain the present invention and should not be construed as a limitation to the present invention. For the step numbers in the following embodiments, they are only set for the convenience of description and explanation, and no limitation is imposed on the order between the steps. The execution order of each step in the embodiments can be adaptively adjusted according to the understanding of those skilled in the art.

[0087] In the description of the present invention, it should be understood that for the orientation description, such as the orientation or positional relationship indicated by up, down, front, back, left, right, etc. is based on the orientation or positional relationship shown in the accompanying drawings, and is only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore should not be construed as a limitation to the present invention.

[0088] In the description of the present invention, the meaning of several is one or more, the meaning of multiple is two or more, greater than, less than, exceeding, etc. are understood as not including the present number, and above, below, within, etc. are understood as including the present number. If there is a description of first and second, it is only for the purpose of distinguishing technical features and should not be understood as indicating or implying relative importance or implicitly indicating the quantity of the indicated technical features or implicitly indicating the sequence of the indicated technical features.

[0089] In the description of the present invention, unless otherwise clearly defined, words such as setting, installing, connecting, etc. should be understood in a broad sense, and those skilled in the art can reasonably determine the specific meanings of the above words in the present invention in combination with the specific content of the technical solution.

[0090] As Figure 1 shown, this embodiment provides a workpiece pose estimation method for an unordered sorting scenario, including the following steps:

[0091] S1. Obtain the RGB-D image of the scene workpiece, where the RGB-D image includes an RGB image and a depth image.

[0092] S2. Input the RGB image into the target detection module (YOLO) to obtain the coordinate information of the bounding box of the target workpiece.

[0093] S3. Crop the RGB image and the depth image according to the obtained coordinate information, convert the cropped depth map into a point cloud, and input the point cloud and the cropped RGB image into the pose estimation model to obtain the RGB-point cloud fusion feature.

[0094] S4. Obtain the key point coordinates and the center point coordinates in the camera coordinate system according to the RGB-point cloud fusion feature.

[0095] S5. Constrain the obtained key point coordinates with the ground truth of the key point coordinates, constrain the center point coordinates with the ground truth of the center point coordinates, and perform geometric constraints on the edges between the key point coordinates and the center point coordinates.

[0096] S6. Through the 3D-3D mapping relationship, register the constrained key point coordinates and center point coordinates with the corresponding key point coordinates and center point coordinates in the preset model coordinate system, and output the final pose of the target workpiece.

[0097] As an alternative implementation, it further includes the steps of constructing a grasping platform and verifying the pose accuracy.

[0098] Hardware platform: Construct a grasping platform with the eye outside the hand, and obtain the relationship between the robot and the camera coordinate system and the gripper coordinate system through camera calibration, hand-eye calibration, and gripper calibration.

[0099] Software platform: The method is a two-stage method. In the first stage, the rectangular box information of the target is obtained through the YOLO algorithm of the target detection framework. In the second stage, the original RGB image and depth image are cropped according to the obtained rectangular box information, and the homogeneous transformation matrix of the target in the camera coordinate system is obtained by inputting them into the network.

[0100] Software and hardware connection platform: Solve the problem of cross-platform language calls. The robot hardware platform is written in C# language, and the software platform is written in the PyTorch framework in Python language. Through the Flask module in Python, the homogeneous transformation matrix obtained by the software platform is uploaded to the web page, and the robot hardware platform captures the homogeneous transformation matrix in the web page to achieve the mutual call of cross-platform languages.

[0101] Among them, the pose estimation model is obtained through the following steps A1-A8:

[0102] A1. Obtain rendering data, obtain real data and perform annotation, and obtain the training set and validation set according to the rendering data.

[0103] Obtain the rendering data through the rendering data production toolbox of blenderproc. The composition of the rendering dataset consists of two parts. The first part is to generate a Fibonacci circle for the object to uniformly sample viewpoints, as Figure 4 shown. At the same time, a uniform camera spin angle is adopted at each sampled viewpoint. The second part is the stacked data. The camera samples data with a random spin at a vertical angle, and the scene is transformed with a random background and random lighting. The reason for dividing the rendering data into two parts is as follows: If only the stacked data is used, when the blenderproc object simulates gravity fall, the pose distribution of the object will be concentrated and biased towards certain poses. For example, for a standard part corner code, when falling, it will probably contact the ground with the object surface with a plane. However, the probability of some other poses, such as contacting the ground only with a certain edge, is very small. Therefore, if only the stacked data is input into the network, it is difficult for the network to fully learn the features of the object. Secondly, the purpose of adopting a uniform camera spin angle at each sampled viewpoint is also to further improve the network's learning of the object features. After determining the viewpoint, the camera can also spin along the direction vector connecting the viewpoint and the center point. The features of the object are different at different spin angles. Therefore, increasing the shooting with different camera spin angles can further improve the network's learning of the object features.

[0104] See Figure 2 , for the automatic annotation of real data: Fix the camera at the end of the robotic arm. Through the hand-eye calibration program of eye-in-hand, obtain the pose of the camera relative to the base coordinate of the robotic arm. The camera pose for obtaining real data is through the following formula:

[0105]

[0106] where is the pose of object j relative to the camera at the i-th frame, is the pose of object j in the 0th frame in the camera coordinate system. The key step in obtaining pose information lies in getting the accurate pose of the object in the first frame. Currently, it is difficult for existing pose annotation tools to achieve good accuracy, and more often it is based on manual annotation. There are the following problems in obtaining the accurate pose of the object in the first frame: there are a large number of missing points in the scene point cloud of the target obtained by the RGB-D camera, which makes it difficult to directly perform ICP registration between the target and the scene point cloud to obtain the accurate pose of the object. Therefore, a reconstruction method is used to obtain a dense fused point cloud, and then, based on the known size of the fiducial marker in the scene, it is restored to the normal scale and registered with the model and the target CAD model to obtain the pose of the target object in the camera coordinate system. After obtaining the data, the data needs to be screened. In the case of stacking, it is necessary to define the visible ratio of the object. If the visible ratio of the object is too small, on the one hand, it is difficult for the network to predict the pose with too few features. Secondly, if the visible ratio is too low, it is difficult to perform workpiece sorting. Therefore, it is necessary to screen out the target objects with too low visible ratios. The definition of the visible ratio is shown in the following formula:

[0107]

[0108] p total refers to the sum of all pixels when the object is completely visible in a certain perspective, and p show refers to the actually visible pixels.

[0109] As an optional implementation, filter out the image objects with a visible ratio u greater than 0.5 and input them into the network.

[0110] A2. Make a set of key points.

[0111] The point cloud features and image features extracted by the network are ambiguous, resulting in poor network training effects. This ambiguity is because objects in three-dimensional space are projected onto a two-dimensional space from different angles, and the resulting depth maps and images are the same. This ambiguity situation especially occurs in symmetric objects. The underlying reason is that the definition of the key points predicted by the network is ambiguous, causing the above problems. When a symmetric object rotates a certain angle around a symmetry axis or symmetry plane, the depth map and the image are the same, but the coordinates of the key points are different, resulting in the problem of key point ambiguity. This makes it difficult for the network to learn and converge correctly. To eliminate the problem of key point definition ambiguity, define the singular block of the symmetric object. The singular block is defined as the smallest unit composed of object objects. By rotating, translating, and mirroring the singular block, the original object can be formed. Using the above method, record the transformation matrix for obtaining the original object by rotating and translating the singular block, and multiply the benchmark key point coordinates obtained by the fps algorithm by the transformation matrix to obtain the set of key points of the object.

[0112] A3. Input the RGB-D images in the training set into the object detection module to obtain the rectangular box information of the object.

[0113] The algorithm stage of the present invention includes two stages. The first stage is to obtain the rectangular box coordinates of the object on the image. By performing object detection, the specific position of the object on the image is first determined, reducing the interference factors of the scene background, and greatly improving the convergence speed and accuracy of the subsequent pose estimation network.

[0114] A4. Crop the RGB image and depth image in the RGB-D image according to the rectangular box information, and input the cropped RGB image and depth image into the pose estimation model.

[0115] A5. Use a backpropagation voting mechanism to obtain the key point coordinates, center point coordinates, and obtain the pose result.

[0116] See Figure 3 , adopt a backpropagation voting mechanism to obtain the coordinates of the key points and the center point. At the same time, register with the key points and the center point on the model to obtain the homogeneous transformation matrix of the target object. Obtained through the following formula:

[0117]

[0118] where i is the target reference point in the scene, and j represents the jth key point, represents the confidence of the target reference point in the jth key point. is the coordinate value of the jth key point in the camera coordinate system obtained by voting. The above confidence calculation formula is as follows:

[0119]

[0120]

[0121] is the threshold. In this embodiment, experiments are carried out with .

[0122] The pose estimation network predicts the key point and center point coordinates in the camera coordinate system, and at the same time registers with the key points and the center point on the model to obtain the homogeneous transformation matrix of the target object. The object pose is optimized by the following formula:

[0123]

[0124] X = x 1 , x 2 , x 3 , …, x n

[0125] P = p1 , p 2 , p 3 , …, p n

[0126] Where X is the set of key points in the camera coordinate system, and P is the set of corresponding key points in the object coordinate system under the X point set. The pose of the object is obtained by minimizing the Euclidean distance between two sets of corresponding points.

[0127] It is necessary to verify the obtained pose, and the verification index is ADD-S. The formula is as follows:

[0128]

[0129] In the formula, v ∈ O represents all points included in the object point cloud model, m is the number of points in the point cloud, R and t are the true values, and R * , t * are the values predicted by the network.

[0130] See Figure 5 , Table 1-3. Table 1 shows the comparison results of the ROC curve area for ADDS index calculation. Table 2 shows the comparison results of the ROC curve area for ADDS index calculation when the viewpoint camera spins arbitrarily and the camera shoots at a fixed angle in the rendering dataset production scheme. Table 3 shows the time and accuracy of obtaining key points by different voting schemes.

[0131] Table 1

[0132]

[0133]

[0134] Table 2

[0135]

[0136] Table 3

[0137]

[0138] Among them, during the training process, a combined loss function is used to compare the difference between the value predicted by the network regression and the true annotation value, and the pose estimation model is trained according to the obtained difference.

[0139] The expression of the combined loss function is as follows:

[0140] Loss = w 1 Lk p_v + w 2 L center_v + w 3 L seg + w 4 Lk p + w5 L center +w 6 L center_kp_v

[0141] In the formula, w 1 、w 2 、w 3 、w 4 、w 5 、w 6 are the weights of the loss function.

[0142] Among them, the key point offset loss function is:

[0143]

[0144] In the formula, Ω k is the set of polysemous key points, of i jk* is the true offset of the key point, M is the target key point, N is the randomly selected scene point cloud, and supervised learning is performed on these scene point clouds. The number of the above selected scene point clouds is 4096. The II formula means that the above formula holds only when p i belongs to the points on the target instance I.

[0145] The above center point offset loss function is:

[0146]

[0147] In the formula, of i* is the true offset of the center point.

[0148] The above target segmentation loss function is:

[0149] L seg =-α(1 - q i ) γ log(q i )q i =c i .l i

[0150] In the formula, α is the balancing hyperparameter, and α = 1 is selected here. c i represents the confidence that the i-th point belongs to each category. γ represents the attention hyperparameter. In this embodiment, γ = 2 is selected.

[0151] The above key point loss function is:

[0152]

[0153] In the formula, kp jk* represents the true coordinates of the key point, and supervised learning is performed on the key point coordinates.

[0154] The above center point loss function is as follows:

[0155]

[0156] In the formula, is the true coordinate of the center point, and supervised learning is performed on the center point coordinates.

[0157] The above key point center point edge vector offset loss function is as follows:

[0158]

[0159] In the formula, cof jk* is the true offset between the key point set and the center point.

[0160] A6. Use the increased geometric constraints to train the pose estimation model.

[0161] Using the obtained key point coordinates and center point coordinates, the increased geometric constraints include the key point loss function, the center point loss function, and the key point center point edge vector offset loss function. Compare the difference between the value predicted by the network regression and the true value, and train the pose estimation network model.

[0162] A7. Use transfer learning to fine-tune the pose estimation model according to the real data.

[0163] After training on a large number of rendered data sets, use the obtained weights as the initial weights to fine-tune on the real data, so that the network has a certain generalization ability on the real data.

[0164] A8. Use the validation set to validate the trained pose estimation model.

[0165] Input the real data into the YOLO network to obtain the rectangular frame position information of the target on the two-dimensional image. Crop the original image and the depth image according to the obtained rectangular frame position information, and input them into the pose estimation network module to obtain the pose information of the object in the camera coordinate system.

[0166] In summary, compared with the prior art, this embodiment has the following advantages and beneficial effects:

[0167] (1) Compared with the existing pose estimation methods, the method of this embodiment uses the RGB-point cloud fusion feature for pose estimation and improves the following problems. Among them: 1) Introduce the geometric feature constraints of the target object in the model, such as the constraints of the predicted key points and the center point. 2) Define the key point set according to the shape of the object to eliminate the ambiguity of the image features and the point cloud features. Define the key point set according to the shape of the object, which improves the accuracy of the prediction result.

[0168] (2) The method of this embodiment proposes a real - data annotation method, effectively solving the problem that it is difficult to obtain real data for pose estimation.

[0169] This embodiment also provides a workpiece pose estimation device for an unordered sorting scenario, including:

[0170] At least one processor;

[0171] At least one memory for storing at least one program;

[0172] When the at least one program is executed by the at least one processor, the at least one processor implements Figure 1 The method shown.

[0173] A workpiece pose estimation device for an unordered sorting scenario in this embodiment can execute a workpiece pose estimation method for an unordered sorting scenario provided by the method embodiment of the present invention, can execute any combination of implementation steps of the method embodiment, and has the corresponding functions and beneficial effects of the method.

[0174] This application embodiment also discloses a computer program product or a computer program. The computer program product or the computer program includes computer instructions, and the computer instructions are stored in a computer - readable storage medium. The processor of the computer device can read the computer instructions from the computer - readable storage medium, and the processor executes the computer instructions, so that the computer device executes Figure 1 The method shown.

[0175] In some alternative embodiments, the functions / operations mentioned in the block diagram may not occur in the order mentioned in the operation diagram. For example, depending on the functions / operations involved, two consecutive blocks shown can actually be executed substantially simultaneously or the blocks can sometimes be executed in the reverse order. In addition, the embodiments presented and described in the flowcharts of the present invention are provided by way of example for the purpose of providing a more comprehensive understanding of the technology. The disclosed method is not limited to the operations and logical flows presented herein. Alternative embodiments are foreseeable, where the order of various operations is changed and the sub - operations described as part of a larger operation are executed independently.

[0176] In addition, although the present invention has been described in the context of functional modules, it should be understood that, unless otherwise stated to the contrary, one or more of the functions and / or features described may be integrated in a single physical device and / or software module, or one or more functions and / or features may be implemented in separate physical devices or software modules. It should also be understood that a detailed discussion of the actual implementation of each module is not necessary for an understanding of the present invention. Rather, given the attributes, functions, and internal relationships of the various functional modules in the devices disclosed herein, the actual implementation of the modules will be understood within the ordinary skills of an engineer. Thus, those skilled in the art can implement the present invention as set forth in the claims without undue experimentation. It should also be understood that the specific concepts disclosed are merely illustrative and are not intended to limit the scope of the present invention, which is determined by the full scope of the appended claims and their equivalents.

[0177] If the described functions are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention. The aforementioned storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROMs), random access memories (RAMs), magnetic disks, or optical discs that can store program codes.

[0178] The logic and / or steps represented in the flowchart or otherwise described herein, for example, can be considered as a predefined sequence of executable instructions for implementing logical functions, and can be specifically implemented in any computer-readable medium for use by an instruction execution system, apparatus, or device (such as a computer-based system, a system including a processor, or other systems that can fetch and execute instructions from the instruction execution system, apparatus, or device), or in conjunction with such instruction execution systems, apparatus, or devices. For the purposes of this specification, a "computer-readable medium" can be any device that can contain, store, communicate, propagate, or transport a program for use by or in conjunction with an instruction execution system, apparatus, or device.

[0179] More specific examples (nonexhaustive list) of computer-readable media include the following: an electrical connection (electronic device) having one or more wirings, a portable computer diskette (magnetic device), a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber device, and a portable compact disc read-only memory (CDROM). Additionally, the computer-readable media can even be paper or other suitable media on which the program can be printed, since the program can be obtained electronically, for example, by optically scanning the paper or other media, followed by editing, interpretation, or otherwise processing as appropriate, and then stored in a computer memory.

[0180] It should be understood that the various parts of the present invention can be implemented by hardware, software, firmware, or a combination thereof. In the above-described embodiments, the multiple steps or methods can be implemented by software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented by hardware, as in another embodiment, any one or a combination of the following techniques well known in the art can be used: discrete logic circuits having logic gate circuits for implementing logical functions on data signals, application specific integrated circuits having appropriate combinational logic gate circuits, programmable gate arrays (PGAs), field programmable gate arrays (FPGAs), and the like.

[0181] In the foregoing description of the present specification, the descriptions referring to the terms "one embodiment / example", "another embodiment / example", or "certain embodiments / examples", etc. mean that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described can be combined in any one or more embodiments or examples in a suitable manner.

[0182] Although the embodiments of the present invention have been shown and described, those of ordinary skill in the art can understand that various changes, modifications, substitutions, and variations can be made to these embodiments without departing from the principles and spirit of the present invention, and the scope of the present invention is defined by the claims and their equivalents.

[0183] The above has specifically described the preferred embodiments of the present invention, but the present invention is not limited to the above embodiments. Those skilled in the art can also make various equivalent deformations or substitutions without departing from the spirit of the present invention, and these equivalent deformations or substitutions are all included within the scope defined by the claims of this application.

Claims

1. A workpiece pose estimation method for disordered sorting scenarios, characterized in that, it includes the following steps: Obtain the RGB-D image of the scene workpiece, where the RGB-D image includes an RGB image and a depth image; Input the RGB image into the target detection module to obtain the coordinate information of the bounding box of the target workpiece; Crop the RGB image and the depth image according to the obtained coordinate information, convert the cropped depth image into a point cloud, and input the point cloud and the cropped RGB image into the pose estimation model to obtain the RGB-point cloud fusion feature; Obtain the key point coordinates and the center point coordinates in the camera coordinate system according to the RGB-point cloud fusion feature; Constrain the obtained key point coordinates and the true values of the key point coordinates, constrain the center point coordinates and the true values of the center point coordinates, and geometrically constrain the edges between the key point coordinates and the center point coordinates; Through the 3D-3D mapping relationship, register the constrained key point coordinates and center point coordinates with the corresponding key point coordinates and center point coordinates in the preset model coordinate system, and output the final pose of the target workpiece; The pose estimation model is obtained by training in the following way: Obtain rendering data, obtain real data and perform annotation, and obtain a training set and a validation set according to the rendering data; Input the RGB-D images in the training set into the target detection module to obtain the rectangular box information of the target; Crop the RGB image and the depth image in the RGB-D image according to the rectangular box information, and input the cropped RGB image and depth image into the pose estimation model; Use a backpropagation-enabled voting mechanism to obtain the key point coordinates, the center point coordinates, and obtain the pose result; Use the added geometric constraints to train the pose estimation model; Use transfer learning to fine-tune the pose estimation model according to the real data; Use the validation set to verify the trained pose estimation model; The rendering data and the real annotated data are obtained in the following way: Use the rendering data production toolbox of blenderproc to obtain the rendering data set; the composition of the rendering data set contains two parts. The first part is that the object generates a Fibonacci circle to uniformly sample viewpoints, and at the same time uses a uniform camera spin angle at each sampled viewpoint. The second part is the stacked data, where the camera randomly spins and samples data at a vertical angle, and the scene is transformed with a random background and random lighting; Annotate the real data, including: Fix the camera at the end of the robotic arm, and obtain the pose of the camera relative to the robotic arm base coordinate through the hand-eye calibration procedure of eye-in-hand; among them, the camera pose of the real data is obtained through the following formula: In the formula, is the pose of object j relative to the camera at the i-th frame, is the pose of object j in the camera coordinate system at the 0-th frame, is the inverse matrix of the camera pose at the i-th frame.

2. The workpiece pose estimation method for disordered sorting scenarios according to claim 1, characterized in that, After obtaining the data, screen the obtained data and delete the target objects with a low visible ratio; Among them, the definition of the visible ratio is as follows: Where p total refers to the sum of all pixels when all pixels of an object are visible from a certain perspective, and p show refers to the actually visible pixels.

3. The workpiece pose estimation method for disordered sorting scenarios according to claim 1, characterized in that, The key point coordinates and the center point coordinates are obtained in the following way: The pose estimation model regresses the key point offset and the center point offset based on the target reference point; For a specific key point, a prediction set of the key point is obtained. By adopting a backpropagation voting mechanism for the obtained key point set, the key point coordinates are obtained. The expression of the backpropagation voting mechanism is as follows: In the formula, where \(i\) is the target reference point in the scene, and \(j\) represents the \(j\)th key point, representing the confidence of the target reference point in the \(j\)th key point; is the coordinate value of the \(j\)th key point in the camera coordinate system obtained by voting; is the coordinate value of the \(j\)th key point estimated by the target reference point, and \(n\) is the total number of target reference points selected in the scene; The confidence calculation formula is as follows: In the formula, is the threshold, represents the set of coordinate values of the j-th key point estimated by the selected target reference point.

4. A workpiece pose estimation method for an unordered sorting scenario according to claim 1, characterized in that, During the training process, a combined loss function is used to compare the values predicted by the model regression with the true annotation values, and the pose estimation model is trained according to the obtained differences. Among them, the combined loss function includes an object segmentation loss function, a key point offset loss function based on the object reference point, a center point offset loss function based on the object reference point, a key point center point edge vector offset loss function, a key point loss function, and a center point loss function.

5. A workpiece pose estimation method for an unordered sorting scenario according to claim 4, characterized in that, The expression of the combined loss function is as follows: Loss=w 1 L kp_v +w 2 L center_v +w 3 L seg +w 4 L kp +w 5 L center +w 6 L center_kp_v where w 1 , w 2 , w 3 , w 4 , w 5 , w 6 are the weights of the loss function; Among them, the key point offset loss function is: where M is the target key point, N is the randomly selected scene point cloud, and Ω k is the set of ambiguous key points, is the true offset of the key point, is the actual offset of the i-th point in the scene to the j-th key point; Equation II means that the equation holds only when p i belongs to the points on the target instance I; The center point offset loss function is: where, of i* is the true offset of the center point; The object segmentation loss function is: L seg = -α(1 - q i ) γ log(q i ), q i = c i .l i where α is the balance hyperparameter, c i represents the confidence that the i-th point belongs to each category, γ represents the attention hyperparameter, l i is the true class label; The key point loss function is: kp jk* Represents the true coordinates of the key point; kp j Represents the actual coordinates of the key point during the training process; The center point loss function is: In the formula, is the true coordinate of the center point, and Δx i is the actual coordinate of the center point during the training process; The key point center point edge vector offset loss function is: where cof jk* is the true offset between the set of key points and the center point, and cof j is the offset between the actual set of key points and the center point.

6. A workpiece pose estimation method for an unordered sorting scenario according to claim 3, characterized in that, After obtaining the key point coordinates and the center point coordinates through the voting mechanism, the pose of the target is obtained through the following method: The pose estimation model predicts the key point coordinates and the center point coordinates in the camera coordinate system, and at the same time registers them with the key points and the center point on the model to obtain the homogeneous transformation matrix of the target object. The pose of the target is obtained through the following formula: X = x 1 , x 2 , x 3 , …, x n P = p 1 , p 2 , p 3 , …, p n Among them, X is the key point set in the camera coordinate system, P is the key point set corresponding to the points in X in the object coordinate system; R is the rotation matrix estimated from the camera to the object coordinate system, t is the translation amount estimated from the camera to the object coordinate system, and N p is the number of corresponding points.

7. A workpiece pose estimation method for an unordered sorting scenario according to claim 6, characterized in that, The verification of the trained pose estimation model using the validation set includes: According to the obtained homogeneous transformation matrix, the model point cloud is transformed through the homogeneous transformation matrix, that is, the model point cloud is transformed from the model coordinate system to the camera coordinate system. The pose estimation model predicts and segments the target scene point cloud, calculates the average Euclidean distance between the model point cloud and the target scene point cloud in the camera coordinate system, and judges the accuracy of the pose estimation according to the value of the average Euclidean distance.

8. A workpiece pose estimation device for an unordered sorting scenario, characterized in that, including: At least one processor; At least one memory for storing at least one program; When the at least one program is executed by the at least one processor, the at least one processor implements the method according to any one of claims 1-7.

Citation Information

Patent Citations

  • Deep learning and geometric algorithm combined non-cooperative target relative pose estimation method

    CN111862126A

  • Real-time 6D pose estimation method and computer readable storage medium

    CN114359377A