Class-level object 6D attitude estimation method and system based on double-projection feature fusion

Through the dual projection feature fusion method, the RGBD camera and feature pyramid network combined with the cross attention algorithm are used to solve the accuracy and stability of class-level objects 6D pose estimation in complex environments, and efficient pose estimation is achieved.

CN120339382APending Publication Date: 2025-07-18JIANGSU UNIV OF SCI & TECH
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202510194226.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-21
Publication Date
2025-07-18

AI Technical Summary

Technical Problem

The existing 6D posture estimation method of class-level objects has insufficient accuracy and stability when dealing with severe occlusion or variable object categories. It is especially difficult to obtain all information for complex objects, and the inference efficiency needs to be improved.

Method used

Using a method based on dual projection feature fusion, the 3D point cloud data of the object is obtained through an RGBD camera, combined with plane and spherical projection feature extraction, and feature fusion is performed using feature pyramid network and cross attention algorithm to generate a rotation matrix to estimate the 6D pose of the object.

Benefits of technology

It improves the robustness and generalization of pose estimation of 6D in class-level objects, can provide accurate pose estimation in complex environments and diverse object categories, and improves the estimation accuracy and processing speed of the system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120339382A_ABST
    Figure CN120339382A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of computer vision and deep learning, and discloses a category-level object 6D attitude estimation method and system based on double projection feature fusion, and the method comprises the steps: S1, obtaining an RGB image and a depth image through a camera, and processing the RGB image and the depth image to obtain 3D point cloud data of an object; s2, obtaining an edge image of the object through an edge detection algorithm; s3, obtaining a point cloud plane projection code according to the 3D point cloud data; s4, according to the 3D point cloud data and the edge image of the object, obtaining a point cloud spherical feature and generating a code; s5, fusing the point cloud spherical feature code and the point cloud plane projection code to obtain an azimuth angle and an inclination angle of the object; s6, obtaining a rotation matrix according to the azimuth angle and the inclination angle, and performing feature conversion on the point cloud spherical features to obtain a plane rotation matrix; and S7, obtaining a three-dimensional rotation matrix of the object. According to the method, accurate 6D attitude estimation can be carried out on various types of objects without depending on a specific object model, and the precision is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical fields of computer vision and deep learning, and particularly to a method and system for 6D pose estimation of category-level objects based on dual-projection feature fusion. Background Art

[0002] 6D pose estimation of category-level objects plays a crucial role in application scenarios such as robot grasping, scene understanding, or autonomous driving. Traditional 6D pose estimation methods usually rely on 3D model information of specific objects and achieve pose recovery of objects in space through feature matching or geometric projection methods. With the development of deep learning technology, pose estimation methods based on deep features have gradually been proposed. For category-level objects, it is no longer necessary to use the CAD model of the object, but instead, the information and models of the same type of objects are used to construct an unknown object model and estimate its pose for generalization on unseen categories, and its application scope is more extensive. Precise object pose information can not only help the system understand the three-dimensional structure of the scene, but also improve the intelligent level of human-computer interaction, thereby enhancing the automation ability of the system. Therefore, 6D pose estimation of category-level objects has very important application value.

[0003] Although current deep learning methods perform well under the condition of sufficient data, there are still problems of insufficient accuracy and stability when dealing with severely occluded or variable object categories. For relatively complex objects (such as cameras), only partial geometric features or image features can be obtained, resulting in feature loss and difficulty in obtaining all information. Moreover, for objects like cameras, the intra-class differences are relatively large, and there is little intra-class prior information that can be utilized. As a result, for complex objects with large intra-class differences, there are problems of insufficient accuracy and stability. In addition, the existing methods need to improve in terms of inference efficiency to better meet the real-time requirements.

[0004] Therefore, to improve the robustness and generalization of 6D pose estimation of category-level objects, the present invention proposes a 6D pose estimation method based on dual-projection feature fusion. Through an improved network structure, an attention module (Attention) is introduced, and combined with a decoupling method of the rotation matrix (R), the estimation accuracy and processing speed of the system are effectively improved, enabling it to provide accurate pose estimation in complex environments and diverse object categories.

[0005] Chinese Invention Patent: Publication No. "CN116630394A", titled "A Multimodal Target Object Pose Estimation Method and System with 3D Modeling Constraints", discloses a multimodal target object pose estimation method with 3D modeling constraints, including obtaining a 2D image of the target object and lidar point cloud data of the target object; extracting image semantic information and image depth information of the 2D image of the target object, fusing the image semantic information and image depth information for 3D expression to obtain 3D semantic features; extracting point cloud features of the lidar point cloud data of the target object; fusing the 3D semantic features and point cloud features to obtain multimodal fusion features; performing key point detection processing on the multimodal fusion features to obtain 3D key points; and matching the 3D key points with the CAD model of the target object to solve the pose of the target object. This technical solution can effectively identify the pose of the target object, help complete subsequent tasks such as path planning or visual positioning according to its pose, and meet the high-precision work requirements of pose estimation. However, this technical solution depends on a specific object model, which may cause feature loss on complex objects and it is difficult to obtain all information. Summary of the Invention

[0006] In order to solve the problems in the above-mentioned existing deep learning methods for category-level object 6D pose estimation, which still have insufficient accuracy and stability when dealing with severely occluded or object category-variable situations, and at the same time, for relatively complex objects, feature loss occurs and it is difficult to obtain all information, the present invention proposes a method and system for category-level object 6D pose estimation based on dual-projection feature fusion, which improves the robustness and generalization of category-level object 6D pose estimation.

[0007] The present invention is realized through the following technical solutions: including the following steps:

[0008] S1. Obtain an RGB image and a depth image through an RGBD camera, obtain a mask of the object in the RGB image through an object detection and segmentation model, and combine the masked image and the depth image with the camera internal parameters to obtain 3D point cloud data of the object;

[0009] S2. Use an edge detection algorithm on the RGB image to obtain an edge image of the object;

[0010] S3. According to the 3D point cloud data of the object obtained in step S1, through a plane projection algorithm and a plane feature extraction network, obtain the point cloud plane projection encoding Emb p ;

[0011] S4. According to the 3D point cloud data of the object and the edge image, perform spherical projection feature extraction and feature aggregation through a Feature Pyramid Network (FPN) to obtain the point cloud spherical feature X 256×(H×W), where 256 is the input channel, and H and W are the spatial dimensions in the vertical and horizontal directions respectively; horizontal and vertical feature extraction and encoding are respectively performed to obtain the horizontal feature encoding Emb H and the vertical feature encoding Emb V ;

[0012] S5. Respectively combine the horizontal feature encoding Emb H and the vertical feature encoding Emb V with the point cloud plane projection encoding Emb p , perform feature fusion through the cross-attention algorithm, and obtain the azimuth angle and the tilt angle θ of the object through the regression network;

[0013] S6. According to the azimuth angle and the tilt angle θ, obtain the rotation matrix R v , perform feature transformation on the point cloud spherical surface features obtained in step S4, and regress to obtain the plane rotation matrix R β ;

[0014] S7. According to the formula R = R v R β , obtain the three-dimensional rotation matrix R of the object.

[0015] As a further preference, the specific steps of step S3 are as follows:

[0016] S31. Perform plane projection on the 3D point cloud data of the object obtained in step S1 on three orthogonal planes to obtain the plane projection data Fp 3×64×64 ;

[0017] S32. Perform feature extraction on the plane projection data obtained in step S31 through the ResNet18 network, and then obtain the projection encoding Emb1 256 through a multi-layer perceptron MLP; encode the projection direction to obtain the projection direction encoding Emb2 256 ;

[0018] S33. Add the projection encoding Emb1 256 and the projection direction encoding Emb2 256 to obtain the point cloud plane projection encoding Emb p .

[0019] As a further preference, the specific steps of step S4 are as follows:

[0020] S41. According to the 3D point cloud data of the object obtained in step S1 and the edge image of the object obtained in step S2, perform spherical projection feature extraction and feature aggregation through the Feature Pyramid Network FPN to obtain the point cloud spherical surface feature X 256×(H×W) ;

[0021] S42. Point cloud spherical feature X 256×(H×W) The transformed point cloud spherical feature Y is obtained through transformation by a multi - layer perceptron MLP 1024×(H×W) ;

[0022] S43. Perform max - pooling processing in the horizontal direction and encode to obtain the horizontal feature encoding Emb H , and the formula is as follows:

[0023] max_pooling θ = max(Y 1024×H×W , horizontal)

[0024] Emb H = MLP θ (max_pooling θ )

[0025] where max_pooling θ is the horizontal direction feature with a size of 1024×W; MLP θ represents a multi - layer perceptron for one - dimensional convolution;

[0026] S44. Perform max - pooling processing in the vertical direction and encode to obtain the vertical feature encoding Emb V , and the formula is as follows:

[0027]

[0028] where is the vertical direction feature with a size of 1024×H; represents a multi - layer perceptron for one - dimensional convolution.

[0029] As a further preference, the specific steps of step S6 are as follows:

[0030] S61. Obtain the rotation matrix R based on the azimuth angle v and the tilt angle θ obtained in step S5, and the formula is as follows:

[0031]

[0032] where and R θ are matrices about the azimuth angle and the tilt angle θ respectively;

[0033] S62. Perform feature transformation on the point cloud spherical feature obtained in step S4 through the rotation matrix R v , and the formula is as follows:

[0034] FTransformed = Transformed(X, R v )

[0035] where Transformed means to perform feature transformation on feature X on the spherical surface;

[0036] S63. Obtain the planar rotation matrix R from the features transformed in step S62 through a convolutional module Conv and a linear module Linear β , and the formula is as follows:

[0037] F c = Conv(F Transformed )

[0038] R β = Linear(Fc)

[0039] where Conv is a 5-layer two-dimensional convolutional layer, and ReLU activation function is used for non-linearity in each layer; Linear is a 3-layer linear layer, and ReLU function is used as the activation function in the first two layers.

[0040] As a further preference, the specific steps of converting the point cloud spherical surface feature X 256×(H×W) to the converted point cloud spherical surface feature Y 1024×(H×W) through the multi-layer perceptron MLP in step S42 are as follows:

[0041] S421. The point cloud spherical surface feature X 256×(H×W) undergoes a channel transformation through the first convolutional layer of the multi-layer perceptron MLP;

[0042] S422. Through the batch normalization layer, perform normalization;

[0043] S423. Perform non-linear processing through the ReLU activation function;

[0044] S424. Through the second convolution, expand the number of channels from 256 to 1024;

[0045] S425. Sequentially pass through the batch normalization layer and the ReLU activation function to obtain the converted point cloud spherical surface feature Y 1024 ×(H×W) .

[0046] As a further preference, the specific steps of the cross-attention algorithm in step S5 are as follows:

[0047] S51. Calculate the similarity between the query and the key, and the formula is as follows:

[0048]

[0049] K = Embp W K

[0050] Q = Emb s W Q

[0051] V = Emb s W V

[0052] Wherein, A is the attention value; Q is the query tensor; K is the key tensor; V is the value tensor; d k is the dimension of the key; W K 、W Q and W V are all trainable weight coefficients; Emb p is the point cloud plane projection encoding; Emb s is the point cloud spherical feature encoding, including the horizontal feature encoding Emb H and the vertical feature encoding Emb V , that is

[0053]

[0054] Step 52: Combine the attention value A obtained in step 51 with the value tensor V to obtain the output OUT. The formula is as follows:

[0055] OUT = A × V.

[0056] Step 53: Obtain the angle by linearly regressing the output OUT obtained in step 52.

[0057] As a further preference, the loss function adopted by the method is the L1 smooth loss function, and the formula is as follows:

[0058]

[0059] L = λ1(L1 + L2) + λ2(L3)

[0060] And:

[0061]

[0062] Where L is the total loss; L1 is the azimuth angle loss; L2 is the tilt angle θ loss; L3 is the rotation matrix R loss; λ1 and λ2 are hyperparameters.

[0063] The present invention also provides a dual-projection feature fusion-based category-level object 6D pose estimation system, which is applied to the fields of robot grasping, scene understanding, and autonomous driving, and specifically includes:

[0064] An image input module for acquiring RGB images and depth images and generating 3D point cloud data and edge images of an object;

[0065] A feature extraction module for extracting planar features and spherical features of an object in two ways respectively. The specific process is as follows:

[0066] Method for extracting planar features of an object: Project the point cloud data onto three orthogonal planes to obtain planar projection data, extract planar projection features through ResNet18, obtain a projection encoding through a multi-layer perceptron, then encode the projection direction to obtain a projection direction encoding, and finally add the projection encoding and the projection direction encoding to obtain a point cloud planar projection encoding;

[0067] Method for extracting spherical features: Project the point cloud data onto a sphere, then flatten the sphere into a plane, perform multi-scale feature fusion using a Feature Pyramid Network (FPN), then perform maximum pooling processing in the horizontal direction and maximum pooling processing in the vertical direction respectively to extract the longitude and latitude information of the point cloud spherical features, and finally perform encoding to obtain a horizontal feature encoding and a vertical feature encoding;

[0068] An attention cross-fusion module for fusing the spherical projection features and planar projection features of the point cloud data to obtain the azimuth angle and tilt angle of the object;

[0069] An output module for obtaining a rotation matrix based on the azimuth angle and tilt angle of the object, converting the spherical projection features of the point cloud through the rotation matrix, obtaining a planar rotation matrix through convolution and linear processing of the converted features, and finally combining the rotation matrix and the planar rotation matrix to obtain a three-dimensional rotation matrix of the object.

[0070] The present invention also provides a computer-readable storage medium on which a computer program is stored. When the program is executed by a processor, the steps in the method for class-level object 6D pose estimation based on dual-projection feature fusion described in the present invention are implemented.

[0071] The present invention also provides a computer device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, the steps in the method for class-level object 6D pose estimation based on dual-projection feature fusion described in the present invention are implemented.

[0072] Compared with the prior art, the beneficial effects of the present invention are as follows:

[0073] 1. The DPFF-Net network in the category-level object 6D pose estimation method designed by the present invention can accurately estimate the 6D poses of various categories of objects without relying on a specific object model, thereby improving the category generalization ability and pose estimation accuracy of the system. It can be widely applied in fields such as robot grasping and autonomous driving.

[0074] 2. The category-level object 6D pose estimation method designed by the present invention adopts multi-scale feature extraction, cross-attention feature fusion, and rotation matrix decoupling methods, and still maintains high accuracy and stability when dealing with occlusion, illumination changes, and complex backgrounds.

[0075] 3. The DPFF-Net network in the category-level object 6D pose estimation method designed by the present invention extracts the structural features of the point cloud from multiple angles by combining point cloud spherical projection feature extraction and planar projection feature extraction, and introduces cross-attention for multi-feature fusion to form an efficient feature representation, and then obtains the pose matrix of the object, improving the accuracy and stability of the system. BRIEF DESCRIPTION OF THE DRAWINGS

[0076] Figure 1 is the overall flowchart of the method of the present invention.

[0077] Figure 2 is the overall architecture diagram of the DPFF-Net network in the method of the present invention.

[0078] Figure 3 is the flowchart of the image preprocessing part in the method of the present invention.

[0079] Figure 4 is the architecture diagram of the CA cross-attention module in the method of the present invention.

[0080] Figure 5 is the schematic diagram of the real data in the embodiment of the present invention.

[0081] Figure 6 is the schematic diagram of the predicted value in the embodiment of the present invention.

[0082] Figure 7 is the 3D IoU accuracy curve graph of various objects in the experiment of the present invention.

[0083] Figure 8 is the rotation accuracy curve graph of various objects in the experiment of the present invention.

[0084] Figure 9 is the translation accuracy curve graph of various objects in the experiment of the present invention. DETAILED DESCRIPTION OF THE INVENTION

[0085] The advantages and features of the present invention will be illustrated and explained through the non-restrictive description of the following preferred embodiments, which are given only as examples with reference to the accompanying drawings.

[0086] Embodiment 1:

[0087] As Figure 1 shown, the present invention provides a method for 6D pose estimation of category-level objects based on dual-projection feature fusion. The present invention constructs a DPFF-Net network for 6D pose estimation of category-level objects based on dual-projection feature fusion. The DPFF-Net network includes an image input module, a feature extraction module, an attention cross-fusion module, and an output module. And a decoupling method of the rotation matrix is adopted to decouple the rotation in the three-dimensional space into three planar rotations to reduce the dimension, and three channels are used to estimate three rotation angles respectively, and finally the 6D pose parameters of the target object are estimated. The input data of the DPFF-Net network are 3D point cloud data and image edge data respectively. The 3D point cloud data obtains the planar projection data F p through the planar projection algorithm, and then obtains the point cloud planar projection encoding Emb p . The 3D point cloud data and the image edge data obtain the spherical projection data F fpn through the spherical projection feature extraction and feature aggregation network FPN; the spherical projection data F fpn obtains the horizontal feature encoding Emb H through the horizontal direction (Horizontal) feature extraction and MLP feature encoding; at the same time, the spherical projection data F fpn obtains the vertical feature encoding Emb V through the vertical direction (Vertital) feature extraction and MLP feature encoding. The horizontal feature encoding Emb H and the vertical feature encoding Emb V respectively obtain the azimuth angle and the tilt angle θ through the cross-attention feature fusion algorithm with the planar projection feature encoding Emb p , and then obtain the rotation matrix R V . Multiply the spherical projection data F fpn by the rotation matrix R V to obtain the projection feature, and obtain the planar rotation matrix R β through the convolutional neural network CNN and the linear layer MLP feature encoding, and then obtain the three-dimensional rotation matrix R of the object, that is, R is the 6D pose estimation result of the object.

[0088] As Figure 2 shown, a method for 6D pose estimation of category-level objects based on dual-projection feature fusion according to the present invention is specifically as follows:

[0089] Step 1: Obtain the RGB image and the depth image through an RGBD camera. Obtain the mask of the object in the RGB image through an object detection and segmentation model. Based on the mask image and the depth image, and combined with the camera intrinsic parameters, obtain the 3D point cloud data of the object.

[0090] Step 2: Apply an edge detection algorithm to the RGB image to obtain the edge image of the object.

[0091] As Figure 3 shown, in the present invention, the RGB image and the depth image of the object are obtained through an RGBD camera or read from a hard disk, but the RGB image and the depth image are not directly used as the input of the method of the present invention. Therefore, the present invention first obtains the mask of the object in the RGB image through an object detection and segmentation model (such as Mask-RCNN). Based on the mask image and the depth image, and combined with the camera intrinsic parameters, the 3D point cloud data of the object is restored. The RGB image undergoes an object edge detection algorithm to obtain the edge image of the object. The edge image of the object and the 3D point cloud data are used as the input data of the DPFF-Net network.

[0092] Step 3: Based on the 3D point cloud data of the object obtained in Step 1, through a plane projection algorithm and a plane feature extraction network, obtain the point cloud plane projection encoding Emb p .

[0093] Step 31: Project the 3D point cloud data of the object obtained in Step 1 onto three orthogonal planes to obtain the plane projection data Fp 3×64×64 .

[0094] Step 32: Apply the ResNet18 network to extract features from the plane projection data obtained in Step 31, and then through a multi-layer perceptron MLP, obtain the projection encoding Emb1 256 ; Encode the projection direction to obtain the projection direction encoding Emb2 256 .

[0095] Step 33: Add the projection encoding Emb1 256 and the projection direction encoding Emb2 256 to obtain the point cloud plane projection encoding Emb p .

[0096] Step 4: Based on the 3D point cloud data of the object obtained in Step 1 and the edge image of the object obtained in Step 2, through a feature pyramid network FPN, perform spherical projection feature extraction and feature aggregation to obtain the spherical projection data F fpn , that is, the point cloud spherical feature X 256×(H×W) , where 256 is the input channel, and H and W are the spatial dimensions in the vertical and horizontal directions respectively; Through horizontal and vertical feature extraction and encoding respectively, obtain the horizontal feature encoding EmbH and vertical feature encoding Emb V 。

[0097] Step 41: According to the 3D point cloud data of the object obtained in Step 1 and the edge image of the object obtained in Step 2, perform spherical projection feature extraction and feature aggregation through the Feature Pyramid Network (FPN) to obtain the spherical projection data F fpn ; The spherical projection data F fpn can also be written in the form of point cloud spherical features X 256×(H×W) , where 256 is the input channel, and H and W are the spatial dimensions in the vertical and horizontal directions respectively.

[0098] Step 42: The point cloud spherical features X 256×(H×W) are converted through a Multi-Layer Perceptron (MLP) to obtain the converted point cloud spherical features Y 1024×(H×W) 。

[0099] The specific process of converting the point cloud spherical features X 256×(H×W) through the MLP into the converted point cloud spherical features Y 1024 ×(H×W) is as follows:

[0100] Step 421: The point cloud spherical features X 256×(H×W) undergo a transformation in the number of channels through the first convolutional layer of the MLP, with the spatial dimensions remaining unchanged;

[0101] Step 422: Normalization is performed through a batch normalization layer; this step can improve the stability of training.

[0102] Step 423: Non-linear processing is performed through a ReLU activation function;

[0103] Step 424: Through the second convolution, the number of channels is expanded from 256 to 1024; the number of input channels of the second convolution is 256, and the output channels are 1024.

[0104] Step 425: Sequentially pass through a batch normalization layer and a ReLU activation function to obtain the converted point cloud spherical features Y 1024×(H×W) 。

[0105] Step 43: Perform max pooling processing along the horizontal direction and encode to obtain the horizontal feature encoding Emb H , and the formula is as follows:

[0106] max_pooling θ =max(Y 1024×H×W ,horizontal)

[0107] Emb H =MLPθ (max_pooling θ )

[0108] where max_pooling θ is a horizontal-direction feature with a size of 1024×W; MLP θ represents a multi-layer perceptron for one-dimensional convolution.

[0109] Step 44: Perform max pooling processing along the vertical direction and encode to obtain the vertical feature encoding Emb V , and the formula is as follows:

[0110]

[0111]

[0112] where is a vertical-direction feature with a size of 1024×H; represents a multi-layer perceptron for one-dimensional convolution.

[0113] The max pooling processing in the horizontal direction and the max pooling processing in the vertical direction are respectively used to extract the information of the longitude and latitude of the point cloud spherical features. The max pooling in the vertical direction performs pooling operations along the latitude direction to extract features related to the latitude. Through the max pooling operation, the maximum feature value in the local area is taken out, so as to retain the most significant feature in the latitude direction of this area. The max pooling in the horizontal direction performs pooling operations along the longitude direction to extract features related to the longitude. The horizontal pooling extracts the most significant feature in this direction by performing max pooling on each interval in the horizontal direction.

[0114] Feature extraction mainly captures 3D point cloud features. Currently, the main methods for processing 3D point clouds are PointNet++ or 3D spherical convolution. However, due to the incomplete input point cloud data, these methods cannot extract features well. The present invention adopts two methods to convert 3D point cloud data into planar data, and then processes it by the method of extracting image features.

[0115] The first method is as shown in Step 3. The point cloud data is projected onto planes in multiple planes, and the point cloud features in multiple directions can be captured. The planar projections of the point cloud in multiple directions obtained are encoded through a convolutional network, and the direction information of the projection plane is added at the same time. Finally, the size of the encoding obtained is N×512, where N is the number of projection planes. The present invention projects the point cloud data P 2048×3 onto three orthogonal planes to obtain the planar projection data Fp 3×64×64 . First, the planar projection feature Fp 256×16×16 is extracted through ResNet18, and then a projection encoding Emb1 is obtained through a multi-layer perceptron MLP256 , and then encode the projection direction to obtain the projection direction encoding Emb2 256 , finally, the projection encoding Emb1 256 and the projection direction encoding Emb2 256 are added together to obtain the point cloud plane projection encoding Emb p .

[0116] The second method is shown in Step 4. The point cloud data is projected onto the sphere, and then the sphere is flattened into a plane. The Feature Pyramid Network (FPN) structure is used for multi-scale feature fusion, and the backbone network can use the classic convolutional network ResNet. Two bottom-up feature extraction network paths respectively represent the point cloud image information and the edge contour information of the pixels corresponding to the point cloud. The resolution of these feature maps decreases layer by layer, and the semantic information increases layer by layer. A top-down feature fusion path passes the high-level features layer by layer to the low-level features and makes lateral connections at each layer to fuse the upsampled high-level features with the features of the current layer. Finally, the last layer is selected as the result of feature extraction for subsequent conversion.

[0117] The point cloud plane projection encoding Emb p , the horizontal feature encoding Emb H and the vertical feature encoding Emb V respectively represent the information of each point cloud in different dimensions.

[0118] Step 5: The horizontal feature encoding Emb H and the vertical feature encoding Emb V obtained in Step 4 are respectively combined with the point cloud plane projection encoding Emb p obtained in Step 3, and feature fusion is performed through the cross-attention algorithm, and the azimuth angle and the tilt angle θ of the object are obtained through the regression network.

[0119] The specific process of the cross-attention algorithm is as follows:

[0120] Step 51: Calculate the similarity between the query and the key, and the formula is as follows:

[0121]

[0122] K = Emb p W K

[0123] Q = Emb s W Q

[0124] V = Emb s W V

[0125] Among them, A is the attention value; Q is the query tensor; K is the key tensor; V is the value tensor; d k is the dimension of the key; W K , W Q and W V are all trainable weight coefficients; Emb p is the point cloud plane projection encoding; Emb s is the point cloud spherical feature encoding, including the horizontal feature encoding Emb H and the vertical feature encoding Emb V , that is

[0126]

[0127] Step 52: Combine the attention value A obtained in step 51 with the value tensor V to obtain the output OUT, and the formula is as follows:

[0128] OUT = A × V.

[0129] Step 53: Regress the output OUT obtained in step 52 through a linear layer to obtain the angle.

[0130] In this step, using the point cloud plane projection encoding Emb p as the key K and the horizontal feature encoding Emb H as the query Q, the obtained is the azimuth angle of the object Using the point cloud plane projection encoding Emb p as the key K and the vertical feature encoding Emb V as the query Q, the obtained is the tilt angle θ of the object.

[0131] As Figure 4 shown, the core operation of the cross-attention algorithm is to calculate the similarity between the query and the key, and perform weighted summation on the value according to this similarity.

[0132] Step 6: Obtain the rotation matrix R and the tilt angle θ obtained in step 5, and perform feature transformation on the point cloud spherical feature obtained in step 4, and regress through the convolutional neural network CNN and the linear layer MLP to obtain the plane rotation matrix R v . β .

[0133] Step 61: Obtain the rotation matrix R and the tilt angle θ obtained in step 5, and the formula is as follows:

[0134]

[0135] Among them and R θ are matrices with respect to the azimuth angle φ and the tilt angle θ, respectively.

[0136] Step 62: Perform feature transformation on the spherical feature of the point cloud obtained in Step 4 through the rotation matrix R v , and the formula is as follows:

[0137] F Transformed = Transformed(X, R v )

[0138] where Transformed represents performing feature transformation on the feature X on the sphere.

[0139] Step 63: Obtain the planar rotation matrix R β from the feature transformed in Step 62 through a convolutional module Conv and a linear module Linear, and the formula is as follows:

[0140] F c = Conv(F Transformed )

[0141] R β = Linear(Fc)

[0142] where Conv is a 5-layer two-dimensional convolutional layer, and the ReLU activation function is used for non-linearity in each layer; Linear is a 3-layer linear layer, and the ReLU function is used as the activation function in the first two layers.

[0143] Step 7: Obtain the three-dimensional rotation matrix R of the object according to the rotation matrix R v and the planar rotation matrix R β , and the formula is as follows:

[0144] R = R v R β

[0145] The three-dimensional rotation matrix R of the object is the pose estimation result of the object.

[0146] For any orthogonal three-dimensional rotation matrix R, it can be decoupled and decomposed into R θ and R p , which respectively represent first rotating the object counterclockwise around the z-axis to obtain the azimuth angle then rotating counterclockwise around the y-axis to obtain the tilt angle θ, and finally rotating counterclockwise around the z-axis by β. Finally, the rotation matrix R can be expressed as:

[0147]

[0148] where R θ and Rβ It can be expressed as:

[0149]

[0150] Rv can be expressed as:

[0151] The attitude of the object, i.e., the rotation matrix R, can be expressed as: R = R v R β .

[0152] Example 2:

[0153] The loss function adopted during the training of the method for category-level object 6D pose estimation based on dual-projection feature fusion described in the present invention is the Smooth L1 Loss function. The Smooth L1 Loss function combines the advantages of the L1 loss and the L2 loss. It uses the L2 loss when the error is small, and uses the L1 loss when the error is large. The benefit of this is to avoid the problem of non-smooth gradients of the L1 loss when the error is small, and at the same time avoid the sensitivity of the L2 loss to outliers. The formula is as follows:

[0154]

[0155] The training loss in the present invention consists of three parts, namely the azimuth loss L1, the tilt angle θ loss L2, and the rotation matrix R loss L3

[0156]

[0157] The total loss L can be expressed as:

[0158] L = λ1(L1 + L2) + λ2(L3)

[0159] Where λ1 and λ2 are hyperparameters.

[0160] Example 3:

[0161] This embodiment provides a category-level object 6D pose estimation system based on dual-projection feature fusion, which is applied to the fields of robot grasping, scene understanding, and autonomous driving. Specifically, it includes an image input module, a feature extraction module, an attention cross-fusion module, and an output module.

[0162] The image input module is used to obtain RGB images and depth images, and generate 3D point cloud data and edge images of objects. The RGB images and depth images are obtained through an RGBD camera or read from a hard disk, but the RGB images and depth images are not directly used as the input of this model. First, the mask of the object in the RGB image is obtained through an object detection and segmentation model (such as Mask-RCNN). Based on the mask image and the depth image, combined with the camera internal parameters, the point cloud data of the object is restored. The RGB image undergoes an object edge detection algorithm to obtain the edge image of the object. The edge image and the point cloud data are used as the input data of the DPFF-Net network designed in the present invention.

[0163] The feature extraction module is used to extract the planar features and spherical features of the object in two ways respectively. The specific process is as follows:

[0164] Method for extracting planar features of the object: The point cloud data is projected onto three orthogonal planes to obtain planar projection data. The planar projection features are extracted through ResNet18, and then a projection encoding is obtained through a multi-layer perceptron. Secondly, the projection direction is encoded to obtain the projection direction encoding. Finally, the projection encoding and the projection direction encoding are added together to obtain the point cloud planar projection encoding.

[0165] Method for extracting spherical features: The point cloud data is projected onto a sphere, and then the sphere is flattened into a plane. The feature pyramid network FPN is used for multi-scale feature fusion. Then, maximum pooling processing is performed in the horizontal direction and the vertical direction respectively to extract the longitude and latitude information of the point cloud spherical features. Finally, encoding is performed to obtain the horizontal feature encoding and the vertical feature encoding.

[0166] The attention cross-fusion module is used to fuse the spherical projection features and planar projection features of the point cloud data to obtain the azimuth angle and tilt angle of the object.

[0167] The output module is used to obtain the rotation matrix based on the azimuth angle and tilt angle of the object, and transform the spherical projection features of the point cloud through the rotation matrix. The transformed features are processed through convolution and linear processing to obtain the planar rotation matrix. Finally, the rotation matrix and the planar rotation matrix are combined to obtain the three-dimensional rotation matrix of the object.

[0168] The image input module, the feature extraction module, the attention cross-fusion module, and the output module constitute the DPFF-Net network.

[0169] Example 4:

[0170] This embodiment provides a computer-readable storage medium, on which a computer program is stored. When the program is executed by a processor, it implements the steps in the method for 6D pose estimation of category-level objects based on dual-projection feature fusion as described in the first embodiment above.

[0171] Embodiment Five:

[0172] This embodiment provides a computer device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, it implements the steps in the method for 6D pose estimation of category-level objects based on dual-projection feature fusion as described in the first embodiment above.

[0173] In the embodiments provided by the present invention, it should be understood that the disclosed method and system can be implemented in other ways. The system embodiments described above are only illustrative. For example, the division of modules is only a logical function division. In actual implementation, there may be other division methods, such as: multiple modules or components can be combined, or can be integrated into another system, or some features can be ignored, or not executed. In addition, the coupling, direct coupling, or communication connection between the components shown or discussed with each other can be indirect coupling or communication connection through some interfaces, devices, or modules, and can be electrical, mechanical, or other forms.

[0174] In addition, in the present invention, each functional module in each embodiment can be all integrated in one processor, or each module can be separately used as a device, or two or more modules can be integrated in one device; each functional module in each embodiment of the present invention can be implemented in the form of hardware, or in the form of a combination of hardware and software functional units.

[0175] Those of ordinary skill in the art can understand that all or part of the steps of implementing the above method embodiments can be completed through program instructions and related hardware. The foregoing program instructions can be stored in a computer-readable storage medium. When the program instructions are executed, they execute the steps including the above method embodiments; and the foregoing storage medium includes: removable storage devices, read-only memory (ROM), magnetic disks, or optical disks, etc., which can store program codes.

[0176] The method for 6D pose estimation of category-level objects based on dual-projection feature fusion of the present invention is trained using the Adam optimizer, the learning rate is set to 0.001, the sample size (batch size) is 128, and the training iterates 50 times. At the end of each iteration, the model is evaluated using the validation set, and the accuracy and loss function values are calculated. The loss hyperparameters λ1 is 5 and λ2 is 1. The verification results are as Figures 5 to 9 and Table 1 shows.

[0177] Table 1 Experimental Results

[0178]

[0179] Figure 5 is the real object box, Figure 6 is the predicted object box. By comparison, it can be seen that the predicted object box is close to the real object box, with high accuracy.

[0180] Figure 7 is the 3D IoU index curve graph. It can be seen that the curve starts to decline near the threshold of 0.8, indicating that the predicted object box is similar to the real object box.

[0181] Figure 8 is the rotation accuracy curve graph. It can be seen that the rotation error of most objects can be controlled within 15°.

[0182] Figure 9 is the translation accuracy curve graph. It can be seen that the translation error of most objects can be controlled within 5 cm.

[0183] It can be seen from the table that the correct rate of 3D IoU being 0.5 is 81.6%; the correct rate of the rotation error being 10° and the translation error being 5 cm is 80.1%.

[0184] Through experimental verification, the verification results show that the method described in the present invention has high efficiency and accuracy in multiple scenarios. The present invention significantly reduces the error of 6D pose estimation of objects and has a wide range of application scenarios.

[0185] In addition to the above embodiments, the present invention can also have other implementation manners. All technical solutions formed by equivalent replacement or equivalent transformation fall within the protection scope required by the present invention.

Claims

1. A method for category-level object 6D pose estimation based on dual-projection feature fusion, characterized in that It includes the following steps: S1. Obtain an RGB image and a depth image through an RGBD camera, obtain the mask of the object in the RGB image through an object detection and segmentation model, and obtain the 3D point cloud data of the object by combining the mask image and the depth image with the camera internal parameters; S2. Use an edge detection algorithm on the RGB image to obtain the edge image of the object; S3. The 3D point cloud data of the object obtained in step S1 is processed through a plane projection algorithm and a plane feature extraction network to obtain the point cloud plane projection encoding Emb p ; S4. Based on the 3D point cloud data and edge image of the object, spherical projection feature extraction and feature aggregation are performed through the Feature Pyramid Network (FPN) to obtain the point cloud spherical feature X 256×(H×W) , where 256 is the input channel, and H and W are the spatial dimensions in the vertical and horizontal directions respectively; horizontal and vertical feature extraction and encoding are respectively performed to obtain the horizontal feature encoding Emb H and the vertical feature encoding Emb V ; S5. Encode the horizontal feature Emb H and the vertical feature Emb V respectively with the point cloud plane projection encoding Emb p , perform feature fusion through the cross-attention algorithm, and obtain the azimuth angle and the tilt angle θ of the object; S6. Obtain the rotation matrix R according to the azimuth angle v and the tilt angle θ, perform feature transformation on the point cloud spherical features obtained in step S4, and regress to obtain the plane rotation matrix R β ; S7. According to the formula R = R v R β , the three-dimensional rotation matrix R of the object is obtained.

2. The method for category-level object 6D pose estimation based on dual-projection feature fusion according to claim 1, characterized in that: The specific steps of step S3 are as follows: S31. Project the 3D point cloud data of the object obtained in step S1 onto three orthogonal planes to obtain plane projection data Fp 3×64×64 ; S32. Feature extraction is performed on the planar projection data obtained in step S31 through the ResNet18 network, and then a projection code Emb1 is obtained through a multi-layer perceptron MLP 256 ; The projection direction is encoded to obtain a projection direction code Emb2 256 ; S33. Add the projection code Emb1 256 and the projection direction code Emb2 256 to obtain the point cloud plane projection code Emb p .

3. The method for category-level object 6D pose estimation based on dual-projection feature fusion according to claim 2, wherein: The specific steps of step S4 are as follows: S41. Based on the 3D point cloud data of the object obtained in step S1 and the edge image of the object obtained in step S2, spherical projection feature extraction and feature aggregation are performed through the Feature Pyramid Network (FPN) to obtain the point cloud spherical feature X 256×(H×W) ; S42. Point cloud spherical feature X 256×(H×W) The transformed point cloud spherical feature Y is obtained through transformation by a multi-layer perceptron MLP 1024 ×(H×W) ; S43. Perform max pooling processing in the horizontal direction and encode to obtain the horizontal feature encoding Emb H , and the formula is as follows: max_pooling θ = max(Y 1024×H×W , horizontal) Emb H = MLP θ (max_pooling θ ) Among them, max_pooling θ is a horizontal feature with a size of 1024×W; MLP θ represents a multi-layer perceptron for one-dimensional convolution; S44. Perform maximum pooling processing in the vertical direction and encode to obtain the vertical feature encoding Emb V , and the formula is as follows: Among them is a vertical direction feature with a size of 1024×H; represents a multi-layer perceptron for one-dimensional convolution.

4. The method for category-level object 6D pose estimation based on dual-projection feature fusion according to claim 3, wherein: The specific steps of step S6 are as follows: S61. Obtain the rotation matrix R based on the azimuth angle and the tilt angle θ obtained in step S5, v The formula is as follows: wherein and R θ are matrices with respect to the azimuth angle and the tilt angle θ, respectively; S62. Through the rotation matrix R v Perform feature transformation on the point cloud spherical feature obtained in step S4, and the formula is as follows: F Transformed = Transformed(X, R v ) Where Transformed means performing feature transformation on feature X on the spherical surface; S63. Obtain the planar rotation matrix R from the features after conversion in step S62 through a convolutional module Conv and a linear module Linear. The formula is as follows: β , as follows: F c = Conv(F Transformed ) R β = Linear(Fc) Where Conv is a 5-layer two-dimensional convolutional layer, and a ReLU activation function is used for non-linearity in each layer; Linear is a 3-layer linear layer, and the ReLU function is used as the activation function for the first two layers.

5. The method for category-level object 6D pose estimation based on dual-projection feature fusion according to claim 4, wherein: In step S42, the point cloud spherical feature X 256×(H×W) is converted into the converted point cloud spherical feature Y through a multi-layer perceptron MLP 1024×(H×W) The specific steps are as follows: S421, Point cloud spherical feature X 256×(H×W) Through the first convolutional layer of the multi-layer perceptron MLP, perform channel transformation; S422. Perform normalization through a batch normalization layer; S423. Perform non-linear processing through a ReLU activation function; S424. Expand the number of channels from 256 to 1024 through the second convolution; S425. Sequentially pass through the batch normalization layer and the ReLU activation function to obtain the transformed point cloud spherical feature Y 1024×(H×W) .

6. The method for category-level object 6D pose estimation based on dual-projection feature fusion according to claim 4, wherein: The specific steps of the cross-attention algorithm in step S5 are as follows: S51. Calculate the similarity between the query and the key, and the formula is as follows: K = Emb p W K Q = Emb s W Q V = Emb s W V Among them, A is the attention value; Q is the query tensor; K is the key tensor; V is the value tensor; d k is the dimension of the key; W K , W Q and W V are all trainable weight coefficients; Emb p is the point cloud plane projection encoding; Emb s is the point cloud spherical feature encoding, including the horizontal feature encoding Emb H and the vertical feature encoding Emb V . S52. Combine the attention value A obtained in step S51 with the value tensor V to obtain the output OUT, and the formula is as follows: OUT = A × V S53. Regress the output OUT obtained in step S52 through a linear layer to obtain the angle.

7. The method for category-level object 6D pose estimation based on dual-projection feature fusion according to claim 1, wherein: The loss function adopted by the method is the L1 smooth loss function, and the formula is as follows: L = λ1(L1 + L2) + λ2(L3) Where L is the total loss; L1 is the azimuth angle φ loss; L2 is the tilt angle θ loss; L3 is the rotation matrix R loss; λ1 and λ2 are hyperparameters.

8. A 6D pose estimation system for category-level objects based on dual-projection feature fusion, characterized in that: It is applied to the fields of robot grasping, scene understanding, and autonomous driving, specifically including: An image input module for obtaining an RGB image and a depth image, and generating the 3D point cloud data and edge image of the object; A feature extraction module for extracting the planar features and spherical features of the object in two ways respectively, and the specific process is as follows: Method for extracting planar features of the object: Project the point cloud data onto three orthogonal planes to obtain planar projection data, extract the planar projection features through ResNet18, then obtain the projection encoding through a multi-layer perceptron, and then encode the projection direction to obtain the projection direction encoding. Finally, add the projection encoding and the projection direction encoding to obtain the point cloud planar projection encoding; Method for extracting spherical features: Project the point cloud data onto the spherical surface, then flatten the spherical surface into a plane, use a feature pyramid network FPN for multi-scale feature fusion, then perform maximum pooling processing in the horizontal direction and maximum pooling processing in the vertical direction respectively, extract the longitude and latitude information of the point cloud spherical features, and finally perform encoding to obtain the horizontal feature encoding and the vertical feature encoding; An attention cross-fusion module for fusing the spherical projection features and planar projection features of the point cloud data to obtain the azimuth angle and tilt angle of the object; An output module, configured to obtain a rotation matrix based on the azimuth angle and tilt angle of an object, convert the spherical projection features of the point cloud through the rotation matrix, obtain a planar rotation matrix through convolution and linear processing of the converted features, and finally combine the rotation matrix and the planar rotation matrix to obtain a three-dimensional rotation matrix of the object.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, it implements the steps in the category-level object 6D pose estimation method based on dual-projection feature fusion according to any one of claims 4 to 7.

10. A computer device, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that: When the processor executes the program, it implements the steps in the category-level object 6D pose estimation method based on dual-projection feature fusion according to any one of claims 4 to 7.

Citation Information

Patent Citations

  • Three-dimensional modeling constrained multi-modal target object attitude estimation method and system

    CN116630394A