A 6D Pose Estimation Method and System Based on Deep Learning

Through deep learning methods, the image and point cloud data of the target object are processed, and the accuracy and inefficiency of 6D pose estimation in the prior art are solved, and the more efficient and lower-cost pose estimation of the target object is achieved, which enhances the robustness and adaptability of the system.

CN115457128BActive Publication Date: 2025-06-10GUANGDONG UNIV OF TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211072450.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-09-02
Publication Date
2025-06-10
Estimated Expiration
2042-09-02

AI Technical Summary

Technical Problem

The prior art has low prediction accuracy and prediction efficiency when estimating the target object 6D pose, and is relatively expensive, making data acquisition difficult, especially in complex scenarios.

Method used

Using a deep learning-based method, by obtaining the markless and marked-coded images of the target object, segmenting the background noise to obtain sparse point cloud data, processing it into dense point cloud data and 3D models, using point cloud completion neural network and 6D pose estimation neural network for training and optimization, and finally 6D pose estimation is performed.

Benefits of technology

The prediction accuracy and prediction efficiency of target object 6D pose estimation are improved, production costs are reduced, the difficulty of obtaining complete point cloud data is simplified, and robustness and generalization are enhanced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115457128B_ABST
    Figure CN115457128B_ABST
Patent Text Reader

Abstract

The present invention provides a 6D pose estimation method and system based on deep learning. The method first acquires the unmarked code image and the marked code image of the target object; then processes the marked code image of the target object, and then uses a preset point cloud completion neural network model to obtain the complete point cloud data and 3D model of the surface of the target object; then inputs the 3D model, the complete point cloud data of the surface and the unmarked code image of the target object into a preset 6D pose estimation neural network model for training and iterative optimization; finally, uses the optimized 6D pose estimation neural network model to perform 6D pose estimation on the target object to be predicted. The method performs 6D pose estimation on the target object based on deep learning, which can effectively improve the prediction accuracy and prediction efficiency of the 6D pose estimation of the target object, and significantly reduces the difficulty of obtaining the complete point cloud data of the target object in the prior art, and reduces the production cost.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer vision technology, and more specifically, to a 6D pose estimation method and system based on deep learning. Background Art

[0002] With the accelerated development of China's industrialization process, industrial manufacturing has entered the intelligent era. Due to its high efficiency, good stability, and strong adaptability to harsh environments, robotic arms are important productive forces in many fields such as industry, 3C inspection, agriculture, and aerospace. In industrial production, in order to better improve the autonomy and environmental adaptability of robotic arms, many current robotic arms have been visualised. The visual grasping system can plan and control the subsequent movements of the robotic arm. The 6D pose estimation of the target object is a prerequisite for the smooth progress of the grasping task. 6D pose estimation refers to the pose relationship between the camera and the object. From the relative pose relationship between the camera and the robotic arm, the pose relationship between the base coordinate system or the tool coordinate system of the robotic arm and the object coordinate system can be obtained.

[0003] In industrial production, point cloud and RGB image information are generally used to estimate the 6D pose of the target object. The point cloud has the position and pose information of the surface points of the object and can provide real physical data; while the RGB image has rich colors and textures and can provide data for object recognition.

[0004] The current prior art discloses a 6D pose estimation method based on deep learning, including an image feature extraction module, a point cloud feature extraction module, an image and point cloud feature fusion module, and a 6D pose estimation module. First, the object is segmented from the scene to obtain the region and the corresponding depth map, and image features are extracted from the segmented image region; then, according to the segmented region and the corresponding depth map, the 3D coordinate information of the surface of the object that is observed is determined to obtain the point cloud data of the object, which is sent to the point cloud feature extraction module to extract the geometric features of the point cloud; then, the extracted image features and point cloud features are sent to the feature fusion network to extract the fusion features; finally, the 6D pose information of the object is estimated according to the extracted fusion features. Although the method in the prior art can solve the problems of low extraction efficiency of dense 2D-3D features and object occlusion to a certain extent, in the prior art, when generating data for the point cloud collected by the sensor or estimated by deep learning, the target object needs to be placed on the platform, and the point cloud information of the part where the object contacts the platform cannot be collected. After collecting the local point cloud data of the object, manual fine-tuning of the data is required, which invisibly increases the difficulty and time consumption of data collection; in addition, during the process of using RGB data to locate and detect the target object, it is limited by interference such as illumination, blur, and ghosting, and the robustness and generalization will drop significantly in complex scenes. In addition, the price of RGB-D cameras with depth information is generally high, resulting in increased production costs.

[0005] Therefore, due to many inconveniences such as difficult data acquisition and high equipment costs, the current existing technologies have problems of low prediction accuracy and prediction efficiency when performing 6D pose estimation of the target object. Summary of the Invention

[0006] In order to overcome the defects of low prediction accuracy and prediction efficiency when the above-mentioned existing technologies estimate the 6D pose of the target object, the present invention provides a 6D pose estimation method and system based on deep learning, which can effectively improve the prediction accuracy and prediction efficiency of the 6D pose estimation of the target object.

[0007] In order to solve the above technical problems, the technical solution of the present invention is as follows:

[0008] A 6D pose estimation method based on deep learning includes the following steps:

[0009] S1: Obtain the unmarked code image and the marked code image of the target object respectively;

[0010] S2: Segment the background and noise of the marked code image of the target object to obtain the surface sparse point cloud data of the target object;

[0011] S3: Process the surface sparse point cloud data of the target object to obtain the surface dense point cloud data of the target object and the 3D model of the target object;

[0012] S4: Input the unmarked code image and the surface dense point cloud data of the target object into a preset point cloud completion neural network model to obtain the surface complete point cloud data of the target object;

[0013] S5: Input the 3D model, the surface complete point cloud data and the unmarked code image of the target object into a preset 6D pose estimation neural network model for training to obtain an optimized 6D pose estimation neural network model;

[0014] S6: Obtain the target object image to be predicted, input the target object image to be predicted into the optimized 6D pose estimation neural network model, and obtain the 6D pose of the target object to be predicted.

[0015] Preferably, in the step S1, the specific method for respectively obtaining the unmarked code image and the marked code image of the target object is:

[0016] Use a monocular camera to take pictures of the target object from multiple directions, obtain a group of unmarked code images of the target object, and record the camera pose corresponding to each unmarked code image;

[0017] Place a number of ARUCO markers around the target object, and use the monocular camera to take pictures of the target object from multiple directions to obtain a set of marked images of the target object.

[0018] Preferably, at least 3 unobstructed and clear ARUCO markers are included in the marked images of the target object.

[0019] Preferably, in the step S3, the specific method for processing the sparse point cloud data on the surface of the target object to obtain the dense point cloud data on the surface of the target object and the 3D model of the target object is as follows:

[0020] Use PMVS2 to process the sparse point cloud data on the surface of the target object to obtain the dense point cloud data on the surface of the target object, the surface triangular meshing data, and the texture mapping data;

[0021] Generate the 3D model of the target object according to the surface triangular meshing data and the texture mapping data.

[0022] Preferably, in the step S4, the specific method for inputting the unmarked image of the target object and the dense point cloud data on the surface into the preset point cloud completion neural network model to obtain the complete point cloud data on the surface of the target object is as follows:

[0023] S4.1: Input any unmarked image of the target object and the dense point cloud data on the surface into the preset point cloud completion neural network model, perform a rotation operation on the dense point cloud data on the surface, and adjust it to match the camera pose corresponding to the unmarked image of the target object to obtain the dense point cloud rotation data P of the target object 1 ;

[0024] S4.2: Map the unmarked image of the target object to obtain the global dense point cloud data P of the target object 2 ;

[0025] S4.3: Concatenate the dense point cloud rotation data P of the target object 1 and the global dense point cloud data P 2 and perform uniform downsampling after concatenation;

[0026] S4.4: Use the downsampled global dense point cloud data P 2 to subtract the downsampled dense point cloud rotation data P of the target object 1 , to obtain the input partial point cloud data P f and the input missing partial point cloud data P c ;

[0027] S4.5: Input the partial point cloud data P f , the input missing partial point cloud data P c, the rotated data P of the dense point cloud of the target object after downsampling 1 and the unmarked code image of the target object are stitched to obtain the global feature vector V t ;

[0028] S4.6: Using the global feature vector V t obtain the surface point cloud offset vector of the target object, and stitch the input missing partial point cloud data P c with the surface point cloud offset vector of the target object to obtain the complete surface point cloud data P of the target object 3 .

[0029] Preferably, in the step S4.6, using the global feature vector V t obtain the surface point cloud offset vector of the target object, and stitch the input missing partial point cloud data P c with the surface point cloud offset vector of the target object to obtain the complete surface point cloud data P of the target object 3 The specific method is:

[0030] Input the global feature vector Vt into the preset 1D convolutional layer to obtain the N-dimensional embedding fpoint, then replicate and flatten the N-dimensional embedding fpoint and perform upsampling, and then input it into the preset 2D convolutional layer to obtain the offset point vector. After that, input the offset point vector into the 1D convolutional layer to obtain the surface point cloud offset vector of the target object and the mask of the incomplete point cloud. Stitch the input missing partial point cloud data P c selected by the mask of the incomplete point cloud with the surface point cloud offset vector of the target object to obtain the complete surface point cloud data P of the target object 3 .

[0031] Preferably, after the step S4, it further includes:

[0032] Optimize the preset point cloud completion neural network model using the distance optimization loss function based on the optimal transport theory. The loss function is:

[0033]

[0034]

[0035] L all =αL CD +βL EMD

[0036] where L CD is the first loss function, L EMD is the second loss function, L all is the total loss function, and P 3is the complete point cloud data of the surface of the target object output by the point cloud completion neural network model, P 2 is the global dense point cloud data of the target object, denotes P 3 and P 2 the correspondence between the point cloud data, α and β are the first and second hyperparameters, p 3 is a point in the complete point cloud data P of the surface of the target object 3 p 2 is a point in the global dense point cloud data of the target object.

[0037] Preferably, in the step S5, the 3D model of the target object, the complete point cloud data of the surface, and the unmarked code image of the target object are input into a preset 6D pose estimation neural network model for training to obtain an optimized 6D pose estimation neural network model. The specific method is as follows:

[0038] The 3D model of the target object, the complete point cloud data of the surface, and the unmarked code image of the target object are input into a preset 6D pose estimation neural network model. The preset 6D pose estimation neural network model includes an encoding module, a decoding module, and a data integration module;

[0039] The unmarked code image of the target object is input into the encoding module for encoding to obtain an encoded 2D image;

[0040] After that, the encoded 2D image is sent to the decoding module for decoding. The decoding module includes a rotation prediction head and a translation prediction head;

[0041] The rotation prediction head outputs the 2D point cloud data and the confidence map of the target object according to the encoded 2D image; using the confidence map of the target object and the RANSAC PnP algorithm, the complete point cloud data of the surface of the target object is matched with the 2D point cloud data of the target object to obtain the predicted rotation matrix of the target object;

[0042] The translation prediction head outputs the heat map data, the initial coordinate data, and the depth map data of the target object according to the encoded 2D image. The final coordinate data of the target object is obtained according to the heat map data and the initial coordinate data of the target object. The final coordinate data of the target object and the depth map data of the target object are spliced to obtain the predicted translation vector of the target object;

[0043] After that, the predicted rotation matrix of the target object and the predicted translation vector of the target object are input into the data integration module for data integration to obtain the 6D pose estimation of the target object;

[0044] Generate a new training dataset using the 3D model of the target object, and set the total prediction loss function in the decoding module to optimize the preset 6D pose estimation neural network model, and finally obtain the optimized 6D pose estimation neural network model.

[0045] Preferably, the total prediction loss function set in the decoding module includes the loss function of the rotation prediction head and the loss function of the translation prediction head:

[0046] The loss function of the rotation prediction head is:

[0047]

[0048] where L Map represents the loss function value of the rotation prediction head, n c represents n channels, represents the Hadamard product operation; M conf is the confidence map data of the target object, M coor is the surface complete point cloud coordinate map data of the target object, l 1 represents the first distance, and α and β are the first and second hyperparameters;

[0049] The loss function of the translation prediction head is:

[0050]

[0051] where L reg represents the loss function value of the translation prediction head, Coord gt represents the final coordinate data of the target object, Coord init represents the initial coordinate data of the target object, represents the normalized Gaussian heatmap corresponding to the final coordinate data of the target object, JS(*) represents the JS divergence, l 2 represents the second distance, and λ is the third hyperparameter;

[0052] The total prediction loss function of the decoding module is:

[0053] L trans = α 1 L reg + α 2 L map

[0054] where L trans represents the total prediction loss function value of the decoding module, α 1 and α 2 represent the first and second empirical parameters.

[0055] A 6D pose estimation system based on deep learning, applying the above-mentioned 6D pose estimation method based on deep learning, includes:

[0056] Image acquisition unit: used to acquire the unmarked code image and the marked code image of the target object;

[0057] Point cloud segmentation unit: used to segment the background and noise of the marked code image of the target object to obtain the surface sparse point cloud data of the target object;

[0058] Point cloud processing unit: used to process the surface sparse point cloud data of the target object to obtain the surface dense point cloud data of the target object and the 3D model of the target object;

[0059] Point cloud completion unit: used to input the unmarked code image and the surface dense point cloud data of the target object into a preset point cloud completion neural network model to obtain the surface complete point cloud data of the target object;

[0060] Model optimization unit: used to input the 3D model of the target object, the surface complete point cloud data and the unmarked code image of the target object into a preset 6D pose estimation neural network model for training to obtain an optimized 6D pose estimation neural network model;

[0061] 6D pose estimation unit: used to acquire the target object image to be predicted, and input the target object image to be predicted into the optimized 6D pose estimation neural network model to obtain the 6D pose of the target object to be predicted.

[0062] Compared with the prior art, the beneficial effects of the technical solution of the present invention are:

[0063] The present invention provides a 6D pose estimation method and system based on deep learning. The method first acquires the unmarked code image and the marked code image of the target object; then processes the marked code image of the target object to obtain the surface sparse point cloud data of the target object; then further processes the surface sparse point cloud data of the target object to obtain the surface dense point cloud data and the 3D model; then uses a preset point cloud completion neural network model to process the surface dense point cloud data of the target object to obtain the surface complete point cloud data of the target object; then inputs the 3D model of the target object, the surface complete point cloud data and the unmarked code image into a preset 6D pose estimation neural network model for training and iterative optimization; finally, inputs the target object image to be predicted into the optimized 6D pose estimation neural network model to obtain the 6D pose estimation of the target object to be predicted;

[0064] This method is based on deep learning for 6D pose estimation of target objects, which can effectively improve the prediction accuracy and efficiency of 6D pose estimation of target objects, and has important significance for the intelligent transformation of industrial production. In addition, this method significantly reduces the difficulty of obtaining complete point cloud data of target objects in the prior art and reduces the production cost. BRIEF DESCRIPTION OF THE DRAWINGS

[0065] Figure 1 It is a flowchart of a 6D pose estimation method based on deep learning described in Embodiment 1.

[0066] Figure 2 It is a structural diagram of a point cloud completion neural network model described in Embodiment 2.

[0067] Figure 3 It is a schematic diagram of obtaining complete point cloud data using a global feature vector described in Embodiment 2.

[0068] Figure 4 It is a structural diagram of a 6D pose estimation neural network model described in Embodiment 2.

[0069] Figure 5 It is a structural diagram of a 6D pose estimation system based on deep learning described in Embodiment 3.

[0070] 301 - Image acquisition unit, 302 - Point cloud segmentation unit, 303 - Point cloud processing unit, 304 - Point cloud completion unit, 305 - Model optimization unit, 306 - 6D pose estimation unit. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0071] The drawings are only for illustrative purposes and should not be construed as limitations of this patent;

[0072] To better illustrate this embodiment, some components in the drawings are omitted, enlarged or reduced, and do not represent the dimensions of the actual product;

[0073] For those skilled in the art, it is understandable that some well-known structures and their descriptions in the drawings may be omitted.

[0074] The technical solutions of the present invention will be further described below with reference to the drawings and embodiments.

[0075] Embodiment 1

[0076] As Figure 1 shown, this embodiment provides a 6D pose estimation method based on deep learning, including the following steps:

[0077] S1: Obtain the unmarked code image and the marked code image of the target object respectively;

[0078] S2: Segment the background and noise of the marked-code image of the target object to obtain the surface sparse point cloud data of the target object;

[0079] S3: Process the surface sparse point cloud data of the target object to obtain the surface dense point cloud data of the target object and the 3D model of the target object;

[0080] S4: Input the unmarked-code image and the surface dense point cloud data of the target object into a preset point cloud completion neural network model to obtain the surface complete point cloud data of the target object;

[0081] S5: Input the 3D model, the surface complete point cloud data, and the unmarked-code image of the target object into a preset 6D pose estimation neural network model for training to obtain an optimized 6D pose estimation neural network model;

[0082] S6: Obtain the target object image to be predicted, input the target object image to be predicted into the optimized 6D pose estimation neural network model, and obtain the 6D pose of the target object to be predicted.

[0083] In the specific implementation process, first place the target object on a plane, use a monocular camera to take pictures of the target object from multiple directions to obtain a set of unmarked-code images of the target object, and then place several marked codes around the target object, and use the monocular camera to take pictures of the target object from multiple directions to obtain a set of marked-code images of the target object;

[0084] After that, perform sparse point cloud reconstruction processing on the marked-code image of the target object, use the marked codes in the image to determine the background plane, and segment the background and noise to obtain the surface sparse point cloud data of the target object;

[0085] After that, process the surface sparse point cloud data of the target object to obtain the surface dense point cloud data, surface triangular meshing data, and texture mapping data of the target object, and generate the 3D model of the target object according to the surface triangular meshing data and texture mapping data;

[0086] After that, input the unmarked-code image and the surface dense point cloud data of the target object into a preset point cloud completion neural network model to obtain the surface complete point cloud data of the target object;

[0087] Afterwards, the 3D model of the target object, the surface complete point cloud data, and the markerless code image of the target object are input into a preset 6D pose estimation neural network model for training to obtain an optimized 6D pose estimation neural network model. The optimized 6D pose estimation neural network model can not only output the 2D point cloud data of the target object and match it with the surface complete point cloud data of the target object to obtain a rotation matrix, but also obtain the final coordinate data of the target object based on the heat map data and initial coordinate data of the target object to obtain a translation vector, and finally output the rotation matrix and the translation vector to obtain the 6D pose estimation of the target object;

[0088] Finally, the image of the target object to be predicted is input into the optimized 6D pose estimation neural network model to obtain the 6D pose estimation of the target object to be predicted;

[0089] This method performs 6D pose estimation on the target object based on deep learning, which can effectively improve the prediction accuracy and prediction efficiency of the 6D pose estimation of the target object; in addition, this method significantly reduces the difficulty of obtaining the complete point cloud data of the target object in the prior art and reduces the production cost.

[0090] Embodiment 2

[0091] This embodiment provides a 6D pose estimation method based on deep learning, including the following steps:

[0092] S1: Obtain the markerless code image and the marked code image of the target object respectively;

[0093] S2: Segment the background and noise of the marked code image of the target object to obtain the surface sparse point cloud data of the target object;

[0094] S3: Process the surface sparse point cloud data of the target object to obtain the surface dense point cloud data of the target object and the 3D model of the target object;

[0095] S4: Input the markerless code image and the surface dense point cloud data of the target object into a preset point cloud completion neural network model to obtain the surface complete point cloud data of the target object;

[0096] S5: Input the 3D model of the target object, the surface complete point cloud data, and the markerless code image of the target object into a preset 6D pose estimation neural network model for training to obtain an optimized 6D pose estimation neural network model;

[0097] S6: Obtain the image of the target object to be predicted, and input the image of the target object to be predicted into the optimized 6D pose estimation neural network model to obtain the 6D pose of the target object to be predicted.

[0098] In the step S1, the specific method for separately obtaining the unmarked code image and the marked code image of the target object is as follows:

[0099] Use a monocular camera to take pictures of the target object from multiple directions, obtain a set of unmarked code images of the target object, and record the camera pose corresponding to each unmarked code image;

[0100] Place several ARUCO markers around the target object, and use the monocular camera to take pictures of the target object from multiple directions to obtain a set of marked code images of the target object.

[0101] The marked code image of the target object includes at least 3 unobstructed and clear ARUCO markers.

[0102] In the step S3, the specific method for processing the surface sparse point cloud data of the target object to obtain the surface dense point cloud data of the target object and the 3D model of the target object is as follows:

[0103] Use PMVS2 to process the surface sparse point cloud data of the target object to obtain the surface dense point cloud data, surface triangular meshing data, and texture mapping data of the target object;

[0104] Generate the 3D model of the target object according to the surface triangular meshing data and texture mapping data.

[0105] In the step S4, the specific method for inputting the unmarked code image and the surface dense point cloud data of the target object into a preset point cloud completion neural network model to obtain the surface complete point cloud data of the target object is as follows:

[0106] S4.1: Input any unmarked code image of the target object and the surface dense point cloud data into a preset point cloud completion neural network model, perform a rotation operation on the surface dense point cloud data, and adjust it to match the camera pose corresponding to the unmarked code image of the target object to obtain the dense point cloud rotation data P of the target object 1 ;

[0107] S4.2: Map the unmarked code image of the target object to obtain the global dense point cloud data P of the target object 2 ;

[0108] S4.3: Concatenate the dense point cloud rotation data P of the target object 1 and the global dense point cloud data P 2 and perform uniform downsampling after concatenation;

[0109] S4.4: Use the downsampled global dense point cloud data P 2 to subtract the downsampled dense point cloud rotation data P of the target object 1, obtain the input partial point cloud data P f and the input missing partial point cloud data P c ;

[0110] S4.5: Combine the input partial point cloud data P f , the input missing partial point cloud data P c , the downsampled dense point cloud rotation data P of the target object 1 and the unlabeled code image of the target object to splice and obtain the global feature vector V t ;

[0111] S4.6: Use the global feature vector V t to obtain the surface point cloud offset vector of the target object, and splice the input missing partial point cloud data P c with the surface point cloud offset vector of the target object to obtain the complete surface point cloud data P of the target object 3 .

[0112] In the step S4.6, using the global feature vector V t to obtain the surface point cloud offset vector of the target object, and splice the input missing partial point cloud data P c with the surface point cloud offset vector of the target object to obtain the complete surface point cloud data P of the target object 3 The specific method is as follows:

[0113] Input the global feature vector Vt into the preset 1D convolutional layer to obtain the N-dimensional embedding fpoint, then copy and flatten the N-dimensional embedding fpoint and perform upsampling, and then input it into the preset 2D convolutional layer to obtain the offset point vector. After that, input the offset point vector into the 1D convolutional layer to obtain the surface point cloud offset vector of the target object and the mask of the corresponding incomplete point cloud. Splice the input missing partial point cloud data P c selected by the mask of the incomplete point cloud with the surface point cloud offset vector of the target object to obtain the complete surface point cloud data P of the target object 3 .

[0114] After the step S4, it further includes:

[0115] Optimize the preset point cloud completion neural network model using the distance optimization loss function based on the optimal transport theory. The loss function is:

[0116]

[0117]

[0118] L all = αL CD + βL EMD

[0119] Among them, L CD is the first loss function, L EMD is the second loss function, L all is the total loss function, P 3 is the complete point cloud data of the target object surface output by the point cloud completion neural network model, P 2 is the global dense point cloud data of the target object, represents P 3 and P 2 the correspondence between the point cloud data, α and β are the first and second hyperparameters, p 3 is the complete point cloud data P of the target object surface 3 in the points, p 2 is the point in the global dense point cloud data of the target object.

[0120] In the step S5, the 3D model of the target object, the complete point cloud data of the surface, and the unmarked code image of the target object are input into a preset 6D pose estimation neural network model for training to obtain an optimized 6D pose estimation neural network model. The specific method is as follows:

[0121] The 3D model of the target object, the complete point cloud data of the surface, and the unmarked code image of the target object are input into a preset 6D pose estimation neural network model. The preset 6D pose estimation neural network model includes an encoding module, a decoding module, and a data integration module;

[0122] The unmarked code image of the target object is input into the encoding module for encoding to obtain an encoded 2D image;

[0123] After that, the encoded 2D image is sent to the decoding module for decoding. The decoding module includes a rotation prediction head and a translation prediction head;

[0124] The rotation prediction head outputs the 2D point cloud data and confidence map of the target object according to the encoded 2D image; using the confidence map of the target object and the RANSAC PnP algorithm, the complete point cloud data of the target object surface is matched with the 2D point cloud data of the target object to obtain the predicted rotation matrix of the target object;

[0125] The translation prediction head outputs the heat map data, initial coordinate data, and depth map data of the target object according to the encoded 2D image, obtains the final coordinate data of the target object according to the heat map data and initial coordinate data of the target object, and splices the final coordinate data of the target object and the depth map data of the target object to obtain the predicted translation vector of the target object;

[0126] After that, the predicted target object rotation matrix and the predicted target object translation vector are input into the data integration module for data integration to obtain the 6D pose estimation of the target object;

[0127] A new training data set is generated using the 3D model of the target object, and the total prediction loss function is set in the decoding module to optimize the preset 6D pose estimation neural network model, and finally the optimized 6D pose estimation neural network model is obtained.

[0128] The total prediction loss function set in the decoding module includes the loss function of the rotation prediction head and the loss function of the translation prediction head:

[0129] The loss function of the rotation prediction head is:

[0130]

[0131] where, L Map represents the loss function value of the rotation prediction head, n c represents n channels, represents the Hadamard product operation; M conf is the confidence map data of the target object, M coor is the surface complete point cloud coordinate map data of the target object, l 1 represents the first distance, and α and β are the first and second hyperparameters;

[0132] The loss function of the translation prediction head is:

[0133]

[0134] where, L reg represents the loss function value of the translation prediction head, Coord gt represents the final coordinate data of the target object, Coord init represents the initial coordinate data of the target object, represents the normalized Gaussian heat map corresponding to the final coordinate data of the target object, JS(*) represents the JS divergence, l 2 represents the second distance, and λ is the third hyperparameter;

[0135] The total prediction loss function of the decoding module is:

[0136] L trans =α 1 L reg +α 2 L map

[0137] where, L trans represents the total prediction loss function value of the decoding module, α 1 and α 2Represent the first and second empirical parameters.

[0138] In a specific implementation process, first place the target object on a plane, and use a monocular camera to take pictures of the target object from multiple directions to obtain a set of markerless code images of the target object, and record the camera pose corresponding to each markerless code image; then place several ARUCO marker codes around the target object, and use the monocular camera to take pictures of the target object from multiple directions to obtain a set of marked code images of the target object.

[0139] It should be noted that at least 3 unobstructed and clear ARUCO marker codes are included in the marked code images of the target object; in this embodiment, the included angle between the viewing plane of the monocular camera and the background plane of the target object is not greater than 30°; the number of markerless code images and marked code images of the target object taken is at least 40.

[0140] In this embodiment, an Intel 11th generation i5 processor and an NVIDIA GTX3060 graphics card are used to process the images and point cloud data of the target object.

[0141] Use OpenMVG to perform sparse point cloud reconstruction processing on the marked code images of the target object. The program will use the ARUCO code to determine the background plane, segment the background and noise to obtain the surface sparse point cloud data of the target object.

[0142] After that, use PMVS2 to process the surface sparse point cloud data of the target object to obtain the surface dense point cloud data, surface triangular meshing data and texture mapping data of the target object; generate a 3D model of the target object according to the surface triangular meshing data and texture mapping data.

[0143] Such as Figure 2 and Figure 3 As shown, then input any markerless code image of the target object and the surface dense point cloud data into a preset point cloud completion neural network model. First, perform a rotation operation on the surface dense point cloud data to adjust it to match the camera pose corresponding to the markerless code image of the target object, and obtain the dense point cloud rotation data P of the target object. 1 ; Then map the markerless code image of the target object to the NOCS space through a neural network with an AEE structure, and this neural network outputs the global dense point cloud data P of the target object with N×3. 2 ; Then the dense point cloud rotation data P of the target object 1 and the global dense point cloud data P 2 After splicing, perform uniform downsampling to N c ×3, and then use the downsampled global dense point cloud data P 2 Subtract the downsampled dense point cloud rotation data P of the target object 1, the input partial point cloud data P is obtained f and the input missing partial point cloud data P c ; then the input partial point cloud data P f , the input missing partial point cloud data P c , the downsampled dense point cloud rotation data P of the target object 1 and the unmarked code image of the target object are stitched to obtain the global feature vector V t ; then the global feature vector V t is input into a 1D convolutional layer to obtain an N-dimensional embedding fpoint. After the N-dimensional embedding fpoint is copied and flattened, it is upsampled, and then input into a 2D convolutional layer of 1×R dimensions to obtain R×N c ×3-dimensional offset point vector. The offset point vector is input into the 1D convolutional layer, and the output is R×N c ×3-dimensional surface point cloud offset vector of the target object and the corresponding (N c -N m )×3 masks; the input missing partial point cloud data P c is raised to R dimensions. After being selected by the mask of the incomplete point cloud, it is stitched with the R×N c ×3-dimensional surface point cloud offset vector of the target object to obtain the surface complete point cloud data P of the target object 3 ;

[0144] In order to make the surface complete point cloud data of the target object closer to the real point cloud and improve the generation quality of the surface complete point cloud of the target object, in this embodiment, a distance optimization loss function based on the optimal transport theory is used to optimize the preset point cloud completion neural network model. The loss function is:

[0145]

[0146]

[0147] L all =αL CD +βL EMD

[0148] where L CD is the first loss function, L EMD is the second loss function, L all is the total loss function, P 3 is the surface complete point cloud data of the target object output by the point cloud completion neural network model, P 2 is the global dense point cloud data of the target object, represents the correspondence between P 3 and P 2 point cloud data, and α and β are the first and second hyperparameters, p3 is the complete point cloud data P of the target object surface 3 The point in the 2 is a point in the global dense point cloud data of the target object;

[0149] like Figure 4 As shown, after obtaining the complete surface point cloud data of the target object, the 3D model of the target object, the complete surface point cloud data and the unmarked code image of the target object are input into a preset 6D posture estimation neural network model for training, and the preset 6D posture estimation neural network model includes an encoding module, a decoding module and a data integration module;

[0150] Inputting the unmarked code image of the target object into the encoding module for encoding to obtain an encoded 2D image;

[0151] Then, the encoded 2D image is sent to a decoding module for decoding, wherein the decoding module includes a rotation prediction head and a translation prediction head;

[0152] The rotation prediction head outputs 2D point cloud data and a confidence map of the target object according to the encoded 2D image;

[0153] In this embodiment, hierarchical feature learning is first used to process the complete point cloud data of the target object surface. The specific method of the hierarchical feature learning is as follows: using the FPS (Farthest Point Sampling) algorithm to randomly select M points radiating the entire point cloud from the cluster center, clustering the N points within the area with a radius of R from the sampling center to obtain M clusters, calculating the number of cluster points and sorting them, and eliminating the point cloud in the cluster with the smallest number; expanding the radius R by 1 times and increasing the number of point clouds in the cluster by 1 times, and then continuing to iterate until the number of clusters is 1;

[0154] Then, the confidence map of the target object and the RANSAC PnP algorithm are used to match the hierarchically processed surface complete point cloud data of the target object with the 2D point cloud data of the target object:

[0155]

[0156] The {} in the formula represents the rounding operation, (c u ,c v ) is the center coordinate of the target object, is the size of the unmarked code image of the target object, (c i ,c j ) is the center of the target object in the coordinate space, is the coordinate space size, (i, j) is the key point coordinate in the coordinate space, is the coordinates of the corresponding points of the key points in the coordinate space on the unmarked code image of the target object;

[0157] After that, the predicted rotation matrix of the target object is obtained;

[0158] The translation prediction head outputs the heatmap data, initial coordinate data, and depth map data of the target object based on the encoded 2D image:

[0159]

[0160] where Coord = (Coord x , Coord y ) is the initial coordinate data of the target object, (δ x , δ y ) represents the scale factor of the 3D coordinate point, w and h represent the dimensions of the border of the encoded 2D image, and Size inp represents the size of the image of the target object without the marker code;

[0161] The final coordinate data of the target object is obtained based on the heatmap data and initial coordinate data of the target object. After that, the final coordinate data of the target object and the depth map data of the target object are concatenated to obtain the predicted translation vector of the target object;

[0162] After that, the predicted rotation matrix of the target object and the predicted translation vector of the target object are input into the data integration module for data integration to obtain the 6D pose estimation of the target object;

[0163] A new training data set is generated using the 3D model of the target object, and the total prediction loss function is set in the decoding module to optimize the preset 6D pose estimation neural network model, and finally the optimized 6D pose estimation neural network model is obtained;

[0164] The total prediction loss function set in the decoding module includes the loss function of the rotation prediction head and the loss function of the translation prediction head:

[0165] The loss function of the rotation prediction head is:

[0166]

[0167] where represents the loss function value of the rotation prediction head, n c represents n channels, represents the Hadamard product operation; M conf is the confidence map data of the target object, M coor is the surface complete point cloud coordinate map data of the target object, l 1 represents the first distance, and α and β are the first and second hyperparameters;

[0168] The loss function of the translation prediction head is:

[0169]

[0170] Among them, L reg represents the loss function value of the translation prediction head, Coord gt represents the final coordinate data of the target object, Coord init represents the initial coordinate data of the target object, represents the normalized Gaussian heatmap corresponding to the final coordinate data of the target object, JS(*) represents the JS divergence, l 2 represents the second distance, and λ is the third hyperparameter;

[0171] The total prediction loss function of the decoding module is:

[0172] L trans = α 1 L reg + α 2 L map

[0173] Among them, L trans represents the total prediction loss function value of the decoding module, α 1 and α 2 represent the first and second empirical parameters;

[0174] Finally, the image of the target object to be predicted is input into the optimized 6D pose estimation neural network model to obtain the 6D pose estimation of the target object to be predicted;

[0175] This method performs 6D pose estimation on the target object based on deep learning, which can effectively improve the prediction accuracy and efficiency of the 6D pose estimation of the target object, and has important significance for the intelligent transformation of industrial production; this method adopts the framework of deep learning, and its robustness and generalization are enhanced compared with existing methods such as template matching; in addition, this method significantly reduces the difficulty of obtaining the complete point cloud data of the target object in the prior art and reduces the production cost.

[0176] Embodiment 3

[0177] As Figure 5 shown, this embodiment provides a 6D pose estimation system based on deep learning, which applies the above-mentioned 6D pose estimation method based on deep learning, including:

[0178] Image acquisition unit 301: used to acquire the unmarked code image and the marked code image of the target object;

[0179] Point cloud segmentation unit 302: used to segment the background and noise of the marked code image of the target object to obtain the surface sparse point cloud data of the target object;

[0180] Point cloud processing unit 303: used to process the sparse point cloud data on the surface of the target object to obtain the dense point cloud data on the surface of the target object and the 3D model of the target object;

[0181] Point cloud completion unit 304: used to input the unmarked code image and the dense point cloud data on the surface of the target object into a preset point cloud completion neural network model to obtain the complete point cloud data on the surface of the target object;

[0182] Model optimization unit 305: used to input the 3D model of the target object, the complete point cloud data on the surface, and the unmarked code image of the target object into a preset 6D pose estimation neural network model for training to obtain an optimized 6D pose estimation neural network model;

[0183] 6D pose estimation unit 306: used to obtain the target object image to be predicted and input the target object image to be predicted into the optimized 6D pose estimation neural network model to obtain the 6D pose of the target object to be predicted.

[0184] In the specific implementation process, first, the image acquisition unit 301 is used to obtain the unmarked code image and the marked code image of the target object; then, the point cloud segmentation unit 302 is used to segment the background and noise of the marked code image of the target object to obtain the sparse point cloud data on the surface of the target object; then, the point cloud processing unit 303 is used to process the sparse point cloud data on the surface of the target object to obtain the dense point cloud data on the surface of the target object and the 3D model of the target object; then, the unmarked code image and the dense point cloud data on the surface of the target object are input into the point cloud completion unit 304, and a preset point cloud completion neural network model is used to perform point cloud reconstruction on the dense point cloud data on the surface of the target object to obtain the complete point cloud data on the surface of the target object; then, the 3D model of the target object, the complete point cloud data on the surface, and the unmarked code image of the target object are input into the model optimization unit 305 to train the preset 6D pose estimation neural network model to obtain an optimized 6D pose estimation neural network model; finally, the target object image to be predicted is input into the 6D pose estimation unit 306, and the optimized 6D pose estimation neural network model is used to obtain the 6D pose of the target object to be predicted;

[0185] This system performs 6D pose estimation on the target object based on deep learning, which can effectively improve the prediction accuracy and prediction efficiency of the 6D pose estimation of the target object. In addition, this system can significantly reduce the difficulty of obtaining the complete point cloud data of the target object in the prior art and reduce the production cost.

[0186] The same or similar reference numerals correspond to the same or similar components;

[0187] The terms used to describe the positional relationship in the drawings are for illustrative purposes only and should not be construed as limiting the present patent;

[0188] Obviously, the above embodiments of the present invention are merely examples given for clearly illustrating the present invention, and are not intended to limit the implementation manners of the present invention. For those of ordinary skill in the art, other different forms of changes or modifications can be made based on the above description. It is not necessary and impossible to list all implementation manners here. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the claims of the present invention.

Claims

1. A 6D pose estimation method based on deep learning, characterized in that, it includes the following steps: S1: Obtain the unmarked code image and the marked code image of the target object respectively; S2: Segment the background and noise of the marked code image of the target object to obtain the surface sparse point cloud data of the target object; S3: Process the surface sparse point cloud data of the target object to obtain the surface dense point cloud data of the target object and the 3D model of the target object; S4: Input the unmarked code image and the surface dense point cloud data of the target object into a preset point cloud completion neural network model to obtain the surface complete point cloud data of the target object; specifically as follows: S4.1: Input the unmarked code image of any target object and the surface dense point cloud data into a preset point cloud completion neural network model, perform a rotation operation on the surface dense point cloud data, and adjust it to match the camera pose corresponding to the unmarked code image of the target object, to obtain the dense point cloud rotation data P of the target object 1 ; S4.2: Map the unmarked code image of the target object to obtain the global dense point cloud data P of the target object 2 ; S4.3: Rotate the dense point cloud data P of the target object 1 and the global dense point cloud data P 2 After splicing, perform uniform downsampling; S4.4: Utilize the globally dense point cloud data P after downsampling 2 to subtract the rotated dense point cloud data of the target object after downsampling P 1 , obtaining the input partial point cloud data P f and the input missing partial point cloud data P c ; S4.5: Combine the input partial point cloud data P f , the input missing partial point cloud data P c , the downsampled dense point cloud rotation data P of the target object 1 and the unmarked code image of the target object to splice and obtain the global feature vector V t ; S4.6: Utilize the global feature vector V t Obtain the surface point cloud offset vector of the target object, and splice the input partial point cloud data P c with the surface point cloud offset vector of the target object to obtain the complete surface point cloud data P of the target object 3 ; S5: Input the 3D model of the target object, the surface complete point cloud data and the unmarked code image of the target object into a preset 6D pose estimation neural network model for training to obtain an optimized 6D pose estimation neural network model; S6: Obtain the target object image to be predicted, input the target object image to be predicted into the optimized 6D pose estimation neural network model, and obtain the 6D pose of the target object to be predicted.

2. The 6D pose estimation method based on deep learning according to claim 1, characterized in that, in the step S1, the specific method for obtaining the unmarked code image and the marked code image of the target object respectively is: Use a monocular camera to take pictures of the target object from multiple directions to obtain a set of unmarked code images of the target object, and record the camera pose corresponding to each unmarked code image; Place several ARUCO markers around the target object, and use the monocular camera to take pictures of the target object from multiple directions to obtain a set of marked code images of the target object.

3. The 6D pose estimation method based on deep learning according to claim 2, characterized in that, the marked code image of the target object includes at least 3 unobstructed and clear ARUCO markers.

4. The 6D pose estimation method based on deep learning according to claim 3, characterized in that, in the step S3, the specific method for processing the surface sparse point cloud data of the target object to obtain the surface dense point cloud data of the target object and the 3D model of the target object is: Use PMVS2 to process the surface sparse point cloud data of the target object to obtain the surface dense point cloud data, the surface triangular meshing data and the texture mapping data of the target object; Generate the 3D model of the target object according to the surface triangular meshing data and the texture mapping data.

5. The 6D pose estimation method based on deep learning according to claim 4, characterized in that, In the step S4.6, using the global feature vector V t to obtain the surface point cloud offset vector of the target object, and splicing the input missing partial point cloud data P c with the surface point cloud offset vector of the target object to obtain the complete surface point cloud data P 3 of the target object, the specific method is as follows: Input the global feature vector Vt into a preset 1D convolutional layer to obtain an N-dimensional embedding fpoint. Then, after copying and flattening the N-dimensional embedding fpoint and performing upsampling, input it into a preset 2D convolutional layer to obtain an offset point vector. After that, input the offset point vector into the 1D convolutional layer to obtain the surface point cloud offset vector of the target object and the mask of the corresponding incomplete point cloud. Select the input P with missing partial point cloud data through the mask of the incomplete point cloud c Concatenate it with the surface point cloud offset vector of the target object to obtain the complete surface point cloud data P of the target object 3 .

6. The 6D pose estimation method based on deep learning according to claim 5, characterized in that, after the step S4, it further includes: Optimize the preset point cloud completion neural network model by using a distance optimization loss function based on the optimal transport theory, and the loss function is: Among them, is the first loss function, is the second loss function, is the total loss function, P 3 is the complete point cloud data of the target object surface output by the point cloud completion neural network model, P 2 is the global dense point cloud data of the target object, represents P 3 and P 2 the corresponding relationship between the point cloud data, and are the first and second hyperparameters, p 3 is the point in the complete point cloud data P of the target object surface 3 among them, p 2 is the point in the global dense point cloud data of the target object.

7. The 6D pose estimation method based on deep learning according to claim 6, characterized in that, In step S5, the 3D model of the target object, the surface complete point cloud data, and the unmarked code image of the target object are input into a preset 6D pose estimation neural network model for training to obtain an optimized 6D pose estimation neural network model. The specific method is as follows: The 3D model of the target object, the surface complete point cloud data, and the unmarked code image of the target object are input into a preset 6D pose estimation neural network model. The preset 6D pose estimation neural network model includes an encoding module, a decoding module, and a data integration module; The unmarked code image of the target object is input into the encoding module for encoding to obtain an encoded 2D image; After that, the encoded 2D image is sent to the decoding module for decoding. The decoding module includes a rotation prediction head and a translation prediction head; The rotation prediction head outputs the 2D point cloud data and the confidence map of the target object according to the encoded 2D image; using the confidence map of the target object and the RANSAC PnP algorithm, the surface complete point cloud data of the target object is matched with the 2D point cloud data of the target object to obtain the predicted rotation matrix of the target object; The translation prediction head outputs the heat map data, the initial coordinate data, and the depth map data of the target object according to the encoded 2D image, obtains the final coordinate data of the target object according to the heat map data and the initial coordinate data of the target object, and splices the final coordinate data of the target object and the depth map data of the target object to obtain the predicted translation vector of the target object; After that, the predicted rotation matrix of the target object and the predicted translation vector of the target object are input into the data integration module for data integration to obtain the 6D pose estimation of the target object; A new training data set is generated using the 3D model of the target object, and a total prediction loss function is set in the decoding module to optimize the preset 6D pose estimation neural network model, and finally an optimized 6D pose estimation neural network model is obtained.

8. A 6D pose estimation method based on deep learning according to claim 7, wherein, The total prediction loss function set in the decoding module includes the loss function of the rotation prediction head and the loss function of the translation prediction head: The loss function of the rotation prediction head is: Among them, represents the loss function value of the rotation prediction head, represents channels, represents the Hadamard product operation; M conf is the confidence map data of the target object, M coor is the surface complete point cloud coordinate map data of the target object, represents the first distance, and are the first and second hyperparameters; The loss function of the translation prediction head is: Among them, represents the loss function value of the translation prediction head, represents the final coordinate data of the target object, represents the initial coordinate data of the target object, represents the normalized Gaussian heatmap corresponding to the final coordinate data of the target object, and JS(*) represents the JS divergence, represents the second distance, is the third hyperparameter; The total prediction loss function of the decoding module is: Among them, represents the total prediction loss function value of the decoding module, and represent the first and second empirical parameters.

9. A 6D pose estimation system based on deep learning, applying the 6D pose estimation method based on deep learning according to any one of claims 1-8, wherein, It includes: Image acquisition unit: used to acquire the unmarked code image and the marked code image of the target object; Point cloud segmentation unit: used to segment the background and noise of the marked code image of the target object to obtain the surface sparse point cloud data of the target object; Point cloud processing unit: used to process the surface sparse point cloud data of the target object to obtain the surface dense point cloud data of the target object and the 3D model of the target object; Point cloud completion unit: used to input the unmarked code image and the surface dense point cloud data of the target object into a preset point cloud completion neural network model to obtain the surface complete point cloud data of the target object; Model optimization unit: used to input the 3D model of the target object, the surface complete point cloud data, and the unmarked code image of the target object into a preset 6D pose estimation neural network model for training to obtain an optimized 6D pose estimation neural network model; 6D pose estimation unit: used to obtain the target object image to be predicted and input the target object image to be predicted into the optimized 6D pose estimation neural network model to obtain the 6D pose of the target object to be predicted.

Citation Information

Patent Citations

  • Ship pose estimation method based on three-dimensional point cloud features

    CN111915677A

  • Class-level 6D pose and size estimation method and device

    CN113012122A