Class-level object 6D pose estimation method based on lightweight feature fusion network
By constructing a lightweight feature fusion network and utilizing the PointNet++ network to extract volume point clouds and shape prior features for dense fusion and reconstruction, the problems of network complexity and low accuracy in existing technologies are solved, and efficient 6D pose estimation is achieved.
Patent Information
- Application Number
- CN202310943921.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-07-31
- Publication Date
- 2025-11-21
- Estimated Expiration
- 2043-07-31
AI Technical Summary
Existing 6D pose estimation methods for objects suffer from poor real-time performance due to complex network models, and fail to fully exploit shape priors and geometric relationships between objects, resulting in low estimation accuracy and failing to meet high-precision requirements.
We construct an end-to-end lightweight feature fusion network, extract object point clouds and shape prior features through PointNet++ network, use cosine similarity to measure correlation for dense fusion, and reconstruct object model through multilayer perceptron and deformable network, and obtain 6D pose by combining pose regression.
It significantly improves the accuracy and speed of 6D pose estimation and achieves good generalization effect on different instances within the category.
Smart Images

Figure CN116805383B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The application belongs to the technical field of robot environment perception, and particularly relates to a category-level object 6D pose estimation method based on a lightweight feature fusion network. BACKGROUND
[0002] With the continuous development of robot technology, robot environment perception technology based on three-dimensional vision has penetrated into various fields such as intelligent manufacturing and intelligent logistics due to its wide range of applications. The task of object 6D pose estimation is to estimate the rigid transformation from the object coordinate system to the camera coordinate system, that is, the rotation and translation transformation of the object coordinate system to the camera coordinate system. 6D object pose estimation has strong practical significance, and estimating the pose of 6D objects has attracted more attention from researchers, and it has been increasingly widely used in many practical applications, such as robot operation, autonomous driving and augmented reality.
[0003] Existing pose estimation methods can be divided into instance-level object 6D pose estimation and category-level object 6D pose estimation according to their generalization ability to objects. Instance-level 6D pose estimation refers to predicting the pose of an object instance based on the three-dimensional model of the object. Instance-level 6D pose estimation can only estimate the object instances that appear in the training process. The mainstream methods can be divided into three categories: correspondence-based methods, template-based methods and voting-based methods. The correspondence-based method is to find the correspondence between the observed point cloud and the object model point cloud, and then obtain the object 6D pose through the PnP algorithm. The template-based method is to select the template closest to the current observed object from the templates with labeled 6D poses, and use its 6D pose as the predicted 6D pose. The voting-based method is that each 3D point cloud on the object point cloud contributes to the final 6D pose estimation through voting. Category-level pose estimation can estimate all instances within the object category that appear in the training process. Existing methods convert different instances of the same category to the NOCS space by extracting the RGB features and point cloud features of the object instances to realize the generalization of different instances within the category, and finally obtain the object 6D pose through the point cloud registration algorithm.
[0004] Most existing patents use the RGB-D information of the object to extract object features, which leads to more complex network models and reduces real-time performance. Moreover, the existing 6D pose estimation methods based on shape priors fail to fully exploit the geometric relationship between the shape priors and the objects, resulting in low 6D pose estimation accuracy and inability to meet high-precision requirements.
[0005] Glossary:
[0006] Shape prior: The data form is a point cloud, which can be regarded as the average shape of a certain category of objects, providing geometric prior information of a certain category of objects.
[0007] Fully connected layer: a common layer type in neural networks, also known as dense connection layer. The fully connected layer can perform matrix multiplication and bias addition operation between the connection weight of input features and each neuron, so as to obtain the output result. SUMMARY
[0008] The application discloses a category-level object 6D pose estimation method based on a lightweight feature fusion network, constructs an end-to-end lightweight feature fusion network, effectively fuses object point cloud features and shape prior features, and significantly improves 6D pose estimation accuracy, so that at least one technical problem involved in the background art can be effectively solved.
[0009] To achieve the above-mentioned purpose, the technical scheme of the present application is:
[0010] A category-level object 6D pose estimation method based on a lightweight feature fusion network, comprising the following steps:
[0011] Step S1, obtaining a depth image of an object through a three-dimensional camera, and performing back projection on the depth image according to the internal parameters of the camera to obtain an object point cloud;
[0012] Step S2, inputting the object point cloud and the shape prior into a PointNet++ network to extract object point cloud geometric features and shape prior geometric features;
[0013] Step S3, inputting the object point cloud geometric features and the shape prior geometric features into a feature fusion network for dense fusion;
[0014] Step S4, inputting the fusion features obtained after dense fusion into a multi-layer perceptron and performing average pooling to obtain object point cloud global features and shape prior global features, and reconstructing a reconstructed object model through a deformation network;
[0015] Step S5, obtaining the 6D pose of the object by pose regression through the object point cloud geometric features, the object point cloud global features, the shape prior geometric features and the shape prior global features.
[0016] As a preferred improvement of the present application, in step S2, 1024 points are first sampled from the object point cloud and the shape prior by uniform down-sampling, and then multi-scale features are adaptively combined based on the PointNet++ network using the metric space distance to obtain the geometric features of the object point cloud and the shape prior.
[0017] As a preferred improvement of the present application, step S3 specifically comprises the following steps:
[0018] Step S31, measuring the object point cloud geometric features F o ={Y1,Y2,Y3…,Yn} and shape prior geometric features F r = {X1, X2, X3…, X m} i , Y i} are represented by the following formula:
[0019]
[0020] Wherein, n represents the total number of object point cloud sampling points, m represents the total number of shape prior sampling points, Y i represents the feature vector of the i-th sampling point in the object point cloud, X i represents the feature vector of the i-th sampling point in the shape prior;
[0021] Step S32, the correlation feature F d = {D1, D2, D3, …, D k} between the object point cloud and the shape prior is calculated by constructing a multi-layer perception machine i , Y i );
[0022] Step S33, the object point cloud geometric feature F o , the shape prior geometric feature F r and the correlation feature F d are point-by-point densely fused, which is represented by the following formula:
[0023] F r→o = {(X1, Y1, d1), (X2, Y2, d2), …, (X n , Y n , d n}
[0024] F o→r = {(Y1, X1, d1), (Y2, X2, d2), …, (Y n , X n , d n}
[0025] Wherein, F r→o represents the object point cloud feature fused with F r and F d , F o→r represents the shape prior feature fused with F o and F d .
[0026] As a preferred improvement of the present application, step S4 specifically comprises the following steps:
[0027] Step S41, the object point cloud feature F r→o, shape prior feature F o→r respectively input into the multi-layer perception and average pooling to obtain the global feature G of the object point cloud o and the shape prior global feature G r ;
[0028] Step S42, the shape prior feature F r , the shape prior global feature G r , the object point cloud feature F o and the object point cloud global feature G o are input into the deformation network as input;
[0029] Step S43, the shape prior feature F r , the shape prior global feature G r and the object point cloud global feature G o constitute the shape prior fusion feature, the object point cloud feature F o , the object point cloud global feature G o and the shape prior global feature G r constitute the object instance fusion feature, and the reconstructed object model is obtained by the deformation module composed of the convolutional neural network;
[0030] Step S44, a reconstruction loss is set to reduce the gap between the reconstructed object model and the real object model, and the reconstruction loss is the chamfer distance between the reconstructed object model and the real object model, and the expression is as follows:
[0031]
[0032] Wherein, L CD represents the reconstruction loss, d CD represents the chamfer distance, represents the real object model, represents the reconstructed object model.
[0033] As a preferred improvement of the present application, step S5 specifically comprises the following steps:
[0034] Step S51, the shape prior feature F r , the shape prior global feature G r , the object point cloud feature F o and the object point cloud global feature G o are spliced and input into the multi-layer perception network and average pooling;
[0035] Step S52, the features obtained after pooling are sent into the fully connected layer to obtain the 6D pose of the object.
[0036] The beneficial effects of the present application: the present application constructs an end-to-end lightweight feature fusion network, effectively fuses object point cloud features and shape prior features, introduces shape prior geometric features into instances, makes the shape information more reliable, and fuses the geometric information of instances into shape prior, so that it has better generalization effect on different instances in the category, and the performance of 6D pose estimation accuracy, speed and the like is significantly improved. BRIEF DESCRIPTION OF DRAWINGS
[0037] Fig. 1 The flow framework diagram provided for the embodiments of the present application is provided;
[0038] Fig. 2 The PointNet++ network structure diagram provided for the embodiments of the present application is provided;
[0039] Fig. 3 The multi-layer perceptron network structure diagram provided for the embodiments of the present application is provided;
[0040] Fig. 4 The feature fusion network structure diagram provided for the embodiments of the present application is provided. DETAILED DESCRIPTION
[0041] The technical solutions in the embodiments of the present application will be clearly and completely described below in conjunction with the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, not all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative labor fall within the scope of protection of the present application.
[0042] It should be noted that all directional indications (such as up, down, left, right, front, back, etc.) in the embodiments of the present application are only used to explain the relative positional relationship, motion condition, etc. between components in a certain posture (as shown in the drawings), and if the certain posture changes, the directional indications also change accordingly.
[0043] In addition, the description such as "first", "second" and the like in the present application is only for the purpose of description, and cannot be understood as indicating or implying the relative importance of the indicated technical features or implicitly indicating the number of the indicated technical features. Therefore, the features limited by "first", "second" can explicitly or implicitly include at least one of the features. In the description of the present application, the meaning of "multiple" is at least two, such as two, three, etc., unless otherwise specifically limited.
[0044] In the present application, unless otherwise explicitly specified and limited, the terms "connection", "fixation" and the like should be understood broadly, for example, "fixation" can be fixed connection, or detachable connection, or integral; can be mechanical connection, or electrical connection; can be direct connection, or indirect connection through intermediate medium, can be internal connection of two elements or interaction relationship between two elements, unless otherwise explicitly limited. For those skilled in the art, the specific meaning of the above terms in the present application can be understood according to the specific circumstances.
[0045] In addition, the technical solutions among various embodiments of the present application can be combined with each other, but it must be based on the fact that a person skilled in the art can realize it, when the combination of technical solutions appears contradictory or unachievable, it should be considered that the combination of technical solutions does not exist, and is not within the protection scope required by the present application.
[0046] Reference Figs. 1 to 4 As shown in the drawings, the present application proposes a class-level object 6D pose estimation method based on a lightweight feature fusion network, comprising the following steps:
[0047] Step S1, obtaining the depth image of the object through the three-dimensional camera, and performing back projection on the depth image according to the known camera internal parameter to obtain the object point cloud;
[0048] Step S2, inputting the object point cloud and shape priori into the PointNet++ network to extract the object point cloud geometric feature and shape priori geometric feature;
[0049] In this step, first, 1024 points are sampled from the object point cloud and shape priori by uniform down-sampling, and then input into the PointNet++ network. PointNet++ adaptively combines multi-scale features by using the distance in the metric space to obtain the geometric features of the object point cloud and shape priori.
[0050] Step S3, inputting the object point cloud geometric feature and shape priori geometric feature into the feature fusion network for dense fusion, specifically comprising the following steps:
[0051] Step S31, measuring the correlation between the object point cloud geometric feature F o ={Y1,Y2,Y3…,Y n} and the shape priori geometric feature F r ={X1,X2,X3…,X m} by cosine similarity, which is represented by the following formula:
[0052]
[0053] Wherein, n represents the total number of points sampled from the object point cloud, Y nY i represents the feature vector of the i-th sampling point in the n-th object point cloud, m represents the total number of shape prior sampling points, X m represents the feature vector of the m-th sampling point in the shape prior, X i represents the feature vector of the i-th sampling point in the m shape priors; the total number of object point cloud sampling points is the same as the total number of shape prior sampling points, that is, n = m.
[0054] Step S32, calculate the correlation feature F d between the object point cloud and the shape prior by constructing a multi-layer perception machine k} = MLP(f(X i , Y i ));
[0055] Step S33, point-by-point dense fusion of the object point cloud geometric feature F o , the shape prior geometric feature F r and the correlation feature F d , which is represented by the following formula:
[0056] F r→o = {(X1,Y1,d1),(X2,Y2,d2),…,(X n ,Y n ,d n )}
[0057] F o→r = {(Y1,X1,d1),(Y2,X2,d2),…,(Y n ,X n ,d n )}
[0058] Where F r→o represents the object point cloud feature fused with F r and F d , F o→r represents the shape prior feature fused with F o and F d .
[0059] Step S4, input the fused feature obtained after dense fusion into a multi-layer perception machine and perform average pooling to obtain object point cloud global feature and shape prior global feature, and obtain a reconstructed object model through a deformation network, which specifically includes the following steps:
[0060] Step S41, input the object point cloud feature F r→o and the shape prior feature F o→rThe object point cloud global feature G is obtained after inputting the multi-layer perception and average pooling respectively o and the shape prior global feature G r ;
[0061] In step S42, the shape prior P r ∈R m×3 , the shape prior geometric feature F r , the shape prior global feature G r , the object point cloud geometric feature F o and the object point cloud global feature G o are input into the deformation network;
[0062] In step S43, the shape prior geometric feature F r , the shape prior global feature G r and the object point cloud global feature G o constitute the shape prior fusion feature, the object point cloud geometric feature F o , the object point cloud global feature G o and the shape prior global feature G r constitute the object instance fusion feature, and the reconstructed object model is obtained by the deformation module composed of the convolutional neural network;
[0063] In step S44, the reconstruction loss is set to reduce the gap between the reconstructed object model and the object real model, the reconstruction loss is the chamfer distance between the reconstructed object model and the object real model, and the expression is as follows:
[0064]
[0065] Wherein, L CD represents the reconstruction loss, d CD represents the chamfer distance, represents the object real model, represents the reconstructed object model.
[0066] In step S5, the object point cloud geometric feature F o , the object point cloud global feature G o , the shape prior geometric feature F r and the shape prior global feature G r are obtained by the pose regression to obtain the 6D pose of the object;
[0067] In step S51, the shape prior geometric feature F r , the shape prior global feature G r , the object point cloud geometric feature F o and the object point cloud global feature G o are spliced and input into the multi-layer perception network and average pooling;
[0068] Step S52, the features obtained after the pooling are sent to a fully connected layer, and the 6D pose of the object is calculated.
[0069] The application has the beneficial effects that the application constructs an end-to-end lightweight feature fusion network, effectively fuses the object point cloud features and shape prior features, introduces the shape prior geometric features into the instance, makes the shape information more reliable, fuses the geometric information of the instance into the shape prior, makes the shape prior have a better generalization effect on different instances in the same category, and significantly improves the 6D pose estimation accuracy, speed and other performances.
[0070] The embodiments of the application are described above in combination with the drawings, but the application is not limited to the specific implementation described above, and the specific implementation described above is only illustrative but not restrictive, and those skilled in the art can make many forms under the inspiration of the application without departing from the scope of the application and the protection scope of the claims.
Claims
1. A category-level 6D object pose estimation method based on a lightweight feature fusion network, characterized in that, Includes the following steps: Step S1: Obtain the depth image of the object using a 3D camera, and back-project the depth image according to the camera's internal parameters to obtain the object's point cloud. Step S2 involves inputting the object point cloud and shape prior into the PointNet++ network to extract the geometric features of the object point cloud and the geometric features of the shape prior, specifically including: First, 1024 points are sampled from the object point cloud and shape prior using uniform downsampling. Then, based on the PointNet++ network, multi-scale features are adaptively combined using metric space distance to obtain the geometric features of the object point cloud and shape prior. Step S3 involves inputting the geometric features of the object's point cloud and the prior geometric features of its shape into a feature fusion network for dense fusion. This specifically includes the following steps: Step S31: Measure the geometric features of the object's point cloud using cosine similarity. and shape prior geometric features The correlation f(X) between them i, Y i ), expressed by the following formula: Where n represents the total number of points sampled from the object's point cloud, m represents the total number of points sampled from the shape prior, and Y i X represents the feature vector of the i-th sampling point in the object point cloud. i This represents the feature vector of the i-th sampling point in the shape prior; Step S32: Calculate the correlation features between the object point cloud and shape prior by constructing a multilayer perceptron. ; Step S33, extract the geometric features F of the object point cloud. o Prior geometric features of shape F r and correlation feature F d Point-by-point dense fusion is represented by the following formula: in, It indicates that it has been integrated. and Point cloud features of objects, It indicates that it has been integrated. and Prior shape features; Step S4 involves inputting the fused features obtained after dense fusion into a multilayer perceptron and performing average pooling to obtain global features of the object point cloud and prior global features of its shape. The reconstructed object model is then obtained through a deformable network. This process includes the following steps: Step S41, extract the point cloud features of the object. Shape prior features The global features of the object point cloud are obtained by inputting the data into a multilayer perceptron and performing average pooling. and shape prior global features ; Step S42, combine the shape prior and the shape prior geometric features F r Shape prior global features Geometric features of object point clouds F o Global features of object point clouds As input to the deformable network; Step S43, the shape prior geometric features Shape prior global features Global features of object point clouds The shape prior fusion features constitute the geometric features of the object's point cloud. Global features of object point clouds and shape prior global features The fusion features of the object instances are used to obtain the reconstructed object model through the deformation module composed of a convolutional neural network; Step S44: Set a reconstruction loss to reduce the gap between the reconstructed object model and the real object model. The reconstruction loss is the chamfer distance between the reconstructed object model and the real object model, and the expression is as follows: Among them, L CD Indicates reconstruction loss, d CD Indicates the chamfer distance. Represents a real model of an object. This indicates the reconstruction of the object model; Step S5: Obtain the 6D pose of the object by performing pose regression on the geometric features of the object point cloud, the global features of the object point cloud, the prior geometric features of the shape, and the prior global features of the shape.
2. The method according to claim 1, characterized in that, Step S5 specifically includes the following steps: Step S51, shape prior geometric features Shape prior global features Geometric features of object point clouds and global features of object point clouds After splicing, the data are fed into a multilayer perceptron network and subjected to average pooling. Step S52: The features obtained after pooling are fed into the fully connected layer to calculate the 6D pose of the object.