Category-Level 6D Object Pose Estimation Method Based on Point Cloud Map Attention Network
Through the point cloud diagram attention network, the multi-scale object structural features are extracted and combined with the shape prior adaptation mechanism is solved, and the problem of not being able to fully utilize the structural features of individual instances is achieved in the prior art, and the position estimation of objects with higher accuracy is achieved.
Patent Information
- Application Number
- CN202311083936.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-08-25
- Publication Date
- 2025-07-22
- Estimated Expiration
- 2043-08-25
AI Technical Summary
The prior art cannot fully utilize the unique structural characteristics of individual instances in class-level 6D object position estimation, resulting in limited prediction capabilities.
Using a point cloud diagram attention network method, multi-scale local to global object structural features are extracted from the observed point cloud through the encoder-decoder architecture, and the posture and size of the object are calculated by combining the shape prior adaptation mechanism and the Umeyama algorithm.
It improves the accuracy of estimation of class-level object positions and significantly improves the estimation accuracy of object posture and size, especially in complex structural objects.
Smart Images

Figure CN117132650B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of computer vision and object pose estimation, and particularly to a category-level 6D object pose estimation method based on a point cloud graph attention network. Background Art
[0002] Category-level six-degree-of-freedom (6D) object pose estimation is a fundamental problem in the field of computer vision, which involves predicting the 3D rotation, 3D translation from the object coordinate system to the camera coordinate system, and the 3D size of the object. This technology is widely used in applications such as robotics, augmented reality, and autonomous driving. Due to the large variation in the shapes of objects within different categories, this problem is extremely challenging.
[0003] To address the above problems, the pose and size of the object are recovered by calculating the similarity transformation between the observed point cloud and the reconstructed NOCS coordinates. Therefore, the quality of the reconstructed NOCS coordinates implicitly indicates the accuracy of the subsequent pose estimation. To improve the reconstruction quality of the NOCS coordinates, some methods such as SPD and SGPA reconstruct the 3D model of the object by using the category shape prior point cloud representing the average shape of objects within the same category, and establish a 3D-3D correspondence between the observed point cloud model and the reconstructed point cloud model, thereby realizing 6D pose and size estimation. Although these correspondence-based methods have made great progress, they cannot fully utilize the unique structural features in individual instances, thus limiting their prediction ability. Therefore, how to provide a category-level 6D object pose estimation method based on a point cloud graph attention network is an urgent problem to be solved by those skilled in the art. Summary of the Invention
[0004] An object of the present invention is to propose a category-level 6D object pose estimation method based on a point cloud graph attention network. Experiments conducted on the NOCS-REAL dataset prove that the solution proposed by the present invention is superior to the prior art and achieves better results.
[0005] A category-level 6D object pose estimation method based on a point cloud graph attention network according to an embodiment of the present invention includes:
[0006] S1. Preprocess the input RGB-D image data and extract the observed point cloud of the object under the depth camera;
[0007] S2. Use the point cloud graph attention network to extract multi-scale local-to-global object structure features from the observed point cloud;
[0008] S3. Utilize the shape prior adaptation mechanism and the category shape prior point cloud to reconstruct the 3D point cloud model of the object and regress the normalized NOCS coordinates of the object;
[0009] S4. Calculate the similarity transformation between the reconstructed NOCS coordinates and the observed point cloud through the Umeyama algorithm to obtain the pose and size information of the object.
[0010] Optionally, the S1 specifically includes:
[0011] S11. Use Mask R-CNN to segment and detect the objects in the RGB-D image data to obtain the object mask regions;
[0012] S12. Map the object mask regions to the depth image of the object to obtain the object depth regions;
[0013] S13. Use the camera parameters to convert the depth information of the object into the three-dimensional point cloud of the object, and generate the observed point cloud data observed by the camera.
[0014] Optionally, the point cloud graph attention network is an encoder-decoder architecture, and the encoder-decoder architecture includes:
[0015] A graph attention encoder for extracting multi-scale local-to-global object features from the observed point cloud;
[0016] The graph attention encoder takes the observed point cloud P o as the input:
[0017]
[0018] where represents the set of real numbers, N o represents the number of points, and 3 represents the XYZ three-dimensional coordinates of the points;
[0019] Use a position embedding module to convert the original three-dimensional coordinates of the observed point cloud into high-dimensional feature embeddings;
[0020] Through the graph attention module, hierarchically extract local-to-global instance geometric features from the input feature embeddings;
[0021] An iterative non-parametric decoder for aggregating multi-scale geometric features.
[0022] Optionally, the position embedding module uses a 3D graph convolutional layer to encode the position information in the observed point cloud. For each observed point in the observed point cloud:
[0023] Use the nearest neighbor search algorithm to search for the coordinate set of its M nearest neighbor points as the receptive field of the convolutional kernel of the 3D graph convolution:
[0024]
[0025] Among them, M represents the number of nearest points, and m represents one of the points. represents the three-dimensional coordinates of one of the points;
[0026] Calculate the direction vector d in the receptive field obtained by the nearest neighbor search algorithm m,n :
[0027] d m,n = p m - p n ;
[0028]
[0029] And initialize the support point kernel vector k through a uniform distribution s :
[0030]
[0031] Among them, S represents the number of support points, and each support point k s is a three-dimensional coordinate;
[0032] Embed the position information of p n into a C0-dimensional feature vector, and the obtained feature vector passes through the ReLU activation function to generate a position embedding
[0033]
[0034] Among them, max represents the maximum operation, <> represents the inner product of vectors, and |||| represents the norm of the vector.
[0035] Optionally, the graph attention module performs multi-stage operations on the position embedding:
[0036]
[0037] Among them, G e (P o ) is denoted as G e , N o represents the number of points, and C0 represents the feature dimension of each point;
[0038] Each stage has a different hidden dimension C i , and in each stage i, three different pointwise feature extraction layers are respectively applied to convert the input point features into the corresponding dimension C i :
[0039] iii) Graph convolutional layer GCL;
[0040] iv) Pointwise self-attention layer PSAL;
[0041] iii) Feedforward layer FFN.
[0042] Optionally, the graph convolutional layer GCL extracts the local geometric features F of the object from the input point feature embeddings by utilizing the graph structure defined by the neighboring points of the points i :
[0043]
[0044] where N i represents the number of points in the i-th stage, and C i represents the feature dimension of each point;
[0045] The graph convolutional layer GCL includes a 3D graph convolutional layer and a ReLU function.
[0046] The pointwise self-attention layer PSAL uses the point cloud self-attention mechanism to extract the global geometric feature G from the local geometric features i :
[0047]
[0048] The pointwise self-attention layer applies a shared multi-layer perceptron network to project the local geometric features onto query vectors, key vectors, and value vectors, denoted as Q i , K i and V i :
[0049] Q i = F i W i Q ;
[0050] K i = F i W i K ;
[0051] V i = F i W i V ;
[0052] where W i Q , W i K and W i V are matrices of dimension C i × C i ;
[0053] By calculating the point-to-point attention weights between the query vector and the value vector to capture the global geometric relationships between different points;
[0054] The global geometric features are obtained through the dot product operation of the attention weights and the value vectors.
[0055]
[0056] The feed-forward layer FFN generates the final output for each stage of the graph attention module by using a shared multi-layer perceptron network and the ReLU activation function.
[0057] A residual connection is added before and after the local geometric feature F i :
[0058]
[0059] The obtained geometric features are used as the input point features for the next stage;
[0060] The geometric features are subjected to global max pooling operation and repetition operation to obtain the global shape features of the object where C4 = C5, which describes the global shape information of the instance.
[0061] Optionally, the graph attention module includes two graph max pooling layers, and the output of the graph attention module includes five parts:
[0062]
[0063] At each stage, the graph convolutional layer GCL, the point-wise self-attention layer PSAL, and the feed-forward layer FFL are stacked in sequence to extract the object geometric features from the input point features in a point-wise manner.
[0064] Optionally, the iterative non-parametric decoder includes:
[0065] Suppose is the i-th stage of the graph attention module, where i ∈ {2, 3, 4}, and the geometric feature is the downsampled point set;
[0066] In each iteration of the i-th stage, the nearest neighbor search algorithm is used to search for the nearest neighbor point q i-1 of each point p n,i-1 in the i-th stage in the previous stage point set P n,i ;
[0067] After determining the nearest neighbor point q n,i-1 of each point p n,i , the feature of the point q n,i is propagated to the point p n,i-1 ;
[0068] Update the features of the points in the (i-2)-th stage using the updated features of the points in the (i-1)-th stage;
[0069] The whole process is carried out in an iterative manner until the features of all points in the first stage are determined and applied to the output of each stage of the graph attention module;
[0070] When all the multi-scale geometric features are aligned to the same number of points Through a concatenation operation With the position embedding (G e ) and the global shape feature Are aggregated to generate the final geometric feature
[0071]
[0072] Wherein, Its feature dimension is the sum of the dimensions of all features
[0073] Optionally, the S3 includes:
[0074] Given the class shape prior point cloud corresponding to the observed point cloud Wherein, N And N o And N r Respectively represent the number of points, and each point is the XYZ three-dimensional coordinate;
[0075] Extract the point-wise prior feature G r from the class shape prior, and the point-wise prior feature G r is the concatenation of the local prior feature and the global prior feature ;
[0076]
[0077] Use a three-layer multi-layer perceptron network to generate the local prior feature Use another two-layer multi-layer perceptron network to generate the global prior feature based on the local prior feature Wherein, D1 and D2 are the dimensions of the features respectively;
[0078] Use a ReLU activation function after each multi-layer perceptron layer, and use an adaptive max pooling operation after the last ReLU activation function to generate the global prior feature;
[0079] After obtaining the prior feature G r and the geometric feature After that, a shape prior adaptation mechanism is adopted to generate a deformation field and a correspondence matrix from the features respectively. and the correspondence matrix where N r and N o are both the number of points, which are used for NOCS coordinate regression;
[0080] Each row d of D i represents the deformation from each point of the prior point cloud P r to the reconstructed point cloud :
[0081]
[0082] Each row A of A i has the sum of elements equal to 1, representing the soft correspondence between each point in the observed point cloud P o and all points in its reconstructed point cloud ;
[0083] In the shape prior adaptation stage, two parallel networks are used, and each network consists of a three-layer multi-layer perceptron network, which respectively regresses the deformation field D and the correspondence matrix A;
[0084] Match A with to obtain the NOCS coordinates of the object:
[0085]
[0086] Optionally, the observed point cloud P o and its reconstructed NOCS coordinates use the Umeyama algorithm combined with the RANSAC algorithm to calculate the 6D object pose and 3D size.
[0087] The beneficial effects of the present invention are:
[0088] (1) The present invention proposes a novel point cloud graph attention network, which adopts a network model with an encoder-decoder architecture to extract unique structural features of individual instances from the object point cloud to improve the accuracy of category-level object pose estimation.
[0089] (2) The present invention proposes a graph attention encoder, which first uses 3D graph convolution to extract multi-scale local structural features of the point cloud, and then adopts a self-attention mechanism to extract multi-scale global structural features from these local structural features.
[0090] (3) The present invention proposes an iterative non-parametric decoder, which is used to propagate multi-scale global structural features from fine-grained to coarse-grained, while retaining multi-scale structural features and avoiding information loss during feature propagation. Brief Description of the Drawings
[0091] The drawings are used to provide a further understanding of the present invention, and constitute a part of the specification. Together with the embodiments of the present invention, they are used to explain the present invention, and do not constitute a limitation to the present invention. In the drawings:
[0092] Figure 1 is a flowchart of a category-level 6D object pose estimation method based on a point cloud graph attention network proposed by the present invention;
[0093] Figure 2 is a structural diagram of the point cloud graph attention network in a category-level 6D object pose estimation method based on a point cloud graph attention network proposed by the present invention;
[0094] Figure 3 is a visualization result graph of 6D pose and 3D size estimation of the present invention and an advanced method on the REAL275 dataset in a category-level 6D object pose estimation method based on a point cloud graph attention network proposed by the present invention;
[0095] Figure 4 is a visualization result graph of 3D shape reconstruction of the present invention and an advanced method on the REAL275 dataset in a category-level 6D object pose estimation method based on a point cloud graph attention network proposed by the present invention. Detailed Description of the Preferred Embodiments
[0096] Now, the present invention will be further described in detail with reference to the accompanying drawings. These drawings are all simplified schematic diagrams, only illustrating the basic structure of the present invention in a schematic way, so they only show the components related to the present invention.
[0097] Refer to Figure 1 , a category-level 6D object pose estimation method based on a point cloud graph attention network, includes:
[0098] S1. Preprocess the input RGB-D image data, and extract the observed point cloud of the object under the depth camera;
[0099] In this embodiment, S1 specifically includes:
[0100] S11. Use Mask R-CNN to segment and detect the object in the RGB-D image data, and obtain the object mask region;
[0101] S12. Map the object mask region to the depth image of the object, and obtain the object depth region;
[0102] S13. Use the camera parameters to convert the depth information of the object into the three-dimensional point cloud of the object, and generate the observed point cloud data observed by the camera.
[0103] S2. Use the point cloud graph attention network to extract multi-scale local-to-global object structure features from the observed point cloud;
[0104] Reference Figure 2 , in this embodiment, the point cloud graph attention network is an encoder-decoder architecture, and the encoder-decoder architecture includes:
[0105] A graph attention encoder for extracting multi-scale local-to-global object features from the observed point cloud;
[0106] The graph attention encoder takes the observed point cloud P o as input:
[0107]
[0108] where represents the set of real numbers, N o represents the number of points, and 3 represents the three-dimensional XYZ coordinates of the points;
[0109] Use a position embedding module to transform the original three-dimensional coordinates of the observed point cloud into high-dimensional feature embeddings;
[0110] Through the graph attention module, hierarchically extract local-to-global instance geometric features from the input feature embeddings;
[0111] In this embodiment, the position embedding module uses a 3D graph convolutional layer to encode the position information in the observed point cloud. For each observed point in the observed point cloud:
[0112] Use the nearest neighbor search algorithm to search for the coordinate set of its M nearest neighbor points as the receptive field of the convolutional kernel of the 3D graph convolution:
[0113]
[0114] where M represents the number of nearest points, m represents one of the points, represents the three-dimensional coordinates of one of the points;
[0115] Calculate the direction vector d in the receptive field obtained by the nearest neighbor search algorithm m,n :
[0116] d m,n = p m - p n ;
[0117]
[0118] And initialize the support point kernel vector k through a uniform distribution s :
[0119]
[0120] Among them, S represents the number of support points, and each support point k s is a three-dimensional coordinate;
[0121] Embed the position information of p n into a C0-dimensional feature vector, and the obtained feature vector passes through the ReLU activation function to generate a position embedding
[0122]
[0123] Among them, max represents the operation of finding the maximum value, <> represents the inner product of vectors, and |||| represents the norm of the vector.
[0124] In this embodiment, the graph attention module performs multi-stage operations on the position embedding:
[0125]
[0126] Among them, G e (P o ) is denoted as G e , N o represents the number of points, and C0 represents the feature dimension of each point;
[0127] Each stage has a different hidden dimension C i , and in each stage i, three different pointwise feature extraction layers are respectively applied to convert the input point features into the corresponding dimension C i :
[0128] v) Graph convolutional layer GCL;
[0129] vi) Pointwise self-attention layer PSAL;
[0130] iii) Feed-forward layer FFN.
[0131] In this embodiment, the graph convolutional layer GCL extracts the local geometric features F of the object from the input point feature embedding by using the graph structure defined by the neighbor points of the points i :
[0132]
[0133] Among them, N i represents the number of points in the i-th stage, and C i represents the feature dimension of each point;
[0134] The graph convolutional layer GCL includes a 3D graph convolutional layer and a ReLU function.
[0135] The point-wise self-attention layer PSAL uses the point cloud self-attention mechanism to extract the global geometric feature G from the local geometric features i :
[0136]
[0137] The point-wise self-attention layer applies a shared multi-layer perceptron network to project the local geometric features onto query vectors, key vectors, and value vectors, denoted as Q i , K i and V i :
[0138] Q i = F i W i Q ;
[0139] K i = F i W i K ;
[0140] V i = F i W i V ;
[0141] where W i Q , W i K and W i V are matrices of dimension C i ×C i ;
[0142] By calculating the point-to-point attention weights between the query vector and the value vector to capture the global geometric relationships between different points;
[0143] Through the dot product operation of the attention weights and the value vector, the global geometric feature
[0144]
[0145] The feed-forward layer FFN generates the final output for each stage of the graph attention module by using a shared multi-layer perceptron network and the ReLU activation function;
[0146] A residual connection is added before and after the local geometric feature F i :
[0147]
[0148] The obtained geometric features are used as the input point features for the next stage;
[0149] Geometric features After global max pooling operation and repetition operation, the global shape features of the object are obtained where C4 = C5, describing the global shape information of the instance.
[0150] In this embodiment, the graph attention module includes two graph max pooling layers, and the output of the graph attention module includes five parts:
[0151]
[0152] In each stage, the graph convolutional layer GCL, the point-wise self-attention layer PSAL, and the feed-forward layer FFL are stacked in sequence to extract the object geometric features from the input point features in a point-wise manner, realizing the expression from local to global, thus effectively describing the complex object geometric shape.
[0153] An iterative non-parametric decoder for aggregating multi-scale geometric features.
[0154] In this embodiment, the iterative non-parametric decoder includes:
[0155] Suppose is the i-th stage of the graph attention module, where i ∈ {2, 3, 4}, and the geometric feature is the downsampled point set;
[0156] In each iteration of the i-th stage, the nearest neighbor search algorithm is used to search for the nearest neighbor point q of each point p i-1 in the previous stage point set P n,i-1 in the i-th stage; n,i ;
[0157] After determining the nearest neighbor point q of each point p n,i-1 the feature of point q n,i is propagated to point p n,i ; n,i-1 ;
[0158] The updated features of the points in the i-1-th stage are used to update the features of the points in the i-2-th stage;
[0159] The whole process is carried out in an iterative manner until the features of all points in the first stage are determined and applied to the output of each stage of the graph attention module;
[0160] When all the multi-scale geometric features are aligned to the same number of points will be through the concatenation operation with the position embedding (Ge ) and global shape features are aggregated to generate the final geometric features
[0161]
[0162] Among them, its feature dimension is the sum of the dimensions of all features
[0163] The iterative non-parametric decoder enables the network to propagate point-wise geometric features from fine-grained to coarse-grained in a progressive manner. It retains multi-scale geometric features, avoids information loss during feature propagation between different scales, and does not require any additional learnable parameters.
[0164] S3. Use the shape prior adaptation mechanism and the category shape prior to reconstruct the 3D point cloud model of the object and regress the normalized NOCS coordinates of the object;
[0165] In this embodiment, S3 includes:
[0166] Given the observed point cloud corresponding category shape prior point cloud where N o and N r respectively represent the number of points, and each point is an XYZ three-dimensional coordinate;
[0167] Extract the point-wise prior feature G r from the category shape prior. The point-wise prior feature G r is a local prior feature and a global prior feature cascade
[0168]
[0169] Use a three-layer multi-layer perceptron network to generate local prior features Use another two-layer multi-layer perceptron network to generate global prior features based on local prior features where D1 and D2 are the dimensions of the features respectively;
[0170] Use a ReLU activation function after each multi-layer perceptron layer, and use an adaptive max pooling operation after the last ReLU activation function to generate global prior features;
[0171] After obtaining the prior feature G r and the geometric feature respectively generate the deformation field from the features using the shape prior adaptation mechanism and the correspondence matrix wherein, N r and N o are both the number of points and are used for NOCS coordinate regression;
[0172] Each row d of D i represents the deformation of each point from the prior point cloud P r to the reconstructed point cloud :
[0173]
[0174] The sum of the elements of each row A of A i is 1, indicating the soft correspondence between each point in the observed point cloud P o and all points in its reconstructed point cloud ;
[0175] In the shape prior adaptation stage, two parallel networks are used, and each network consists of a three-layer multi-layer perceptron network, which respectively regresses the deformation field D and the correspondence matrix A;
[0176] Match A with to obtain the NOCS coordinates of the object:
[0177]
[0178] S4. Calculate the similarity transformation between the reconstructed NOCS coordinates and the observed point cloud through the Umeyama algorithm to obtain the pose and size information of the object.
[0179] In this embodiment, the observed point cloud P o and its reconstructed NOCS coordinates use the Umeyama algorithm combined with the RANSAC algorithm to calculate the 6D object pose and 3D size. The Umeyama algorithm is used to estimate the optimal similarity transformation parameters, namely rotation, translation, and scale, where the rotation and translation parameters correspond to the 6D object pose, and the scale parameter corresponds to the object size. The RANSAC algorithm is used to remove outliers and achieve robust estimation.
[0180] Example 1:
[0181] This Example 1 is implemented using the PyTorch framework and experiments are carried out on a desktop computer equipped with an NVIDIA GeForce RTX3090 GPU. The batch size is 64. First, the depth image is cropped using the instance segmentation mask generated by Mask-RCNN, and the cropped depth image is resized to 256×256 pixels. Then, N is randomly sampled from the point cloud converted from the depth image o= 1024 points, forming an observation point cloud. Next, sample N from the class shape prior point cloud pre-trained by the SPD technique r = 1024 points to obtain the prior point cloud. In step two, the hidden layer dimensions of the point cloud graph attention network are set to C0 = 128, C1 = 128, C2 = 256, C3 = 256, C4 = 512, C5 = 512. The hyperparameters in the 3D graph convolutional layer all use the default settings, that is, the number of nearest neighbor points is set to M = 50, and the number of support point kernel vectors is S = 1. In step three, the hidden layer dimension of the multi-layer perceptron network for local prior feature extraction is [64, 64, 64], and the hidden layer dimension of the multi-layer perceptron network for global prior feature extraction is [128, 1024], that is, D1 = 64 and D2 = 1024. For deformation field regression, the hidden layer dimension of the multi-layer perceptron network is set to [512, 256, No × 3], and for correspondence regression, the hidden layer dimension is set to [512, 256, N o ×N r . During the training process, the Adam optimizer is used to optimize the network, the initial learning rate is 1e-4, and the model is trained for a total of 100 rounds. Every 20 rounds, the learning rate is decayed at a ratio of 0.6, 0.3, 0.1, and 0.01. The same loss function as in the SPD technique is used to train the network, and all classes are trained using a single model.
[0182] This embodiment reports the average precision of the 3D intersection over union (IoU) at 50% and 75% thresholds respectively to comprehensively evaluate the accuracy of rotation, translation, and size estimation. To directly compare the rotation and translation errors, the metrics of 5° 2cm, 5° 5cm, 10° 2cm, and 10° 5cm are also adopted. If the rotation and translation errors are lower than the given threshold, the pose is considered correct. In addition, the Chamfer distance is used to evaluate the accuracy of the 3D model reconstruction result.
[0183] Table 1 Quantitative analysis and comparison results of the 6D pose and 3D size estimation of the present invention with advanced methods on the REAL275 dataset
[0184]
[0185] According to the results in Table 1, it can be clearly seen that the method proposed by the present invention is significantly superior to the prior art in object pose and size estimation, achieving the best performance. In comprehensively evaluating the accuracy of rotation, translation, and size estimation, compared with the NOCS technique that only uses RGB features, the proposed solution of the present invention improves by 4.0% on the 3D 50 metric and by 40.3% on the 3D 75 metric. Compared with the SGPA technique that uses RGB-D features at the same time, the proposed solution of the present invention is on the 3D50 The index has increased by 1.9%, and the 3D75 index has increased by 8.5%. In directly evaluating the accuracy of rotation and translation estimation, compared with the NOCS technology that only uses RGB features, the proposed solution of the present invention has increased by 38.7% in the 5°2cm index, 43.8% in the 5°5cm index, 49.3% in the 10°2cm index, and 52.5% in the 10°5cm index. Compared with the SGPA technology that uses RGB-D features simultaneously, the proposed solution of the present invention has increased by 10.0% in the 5°2cm index, 14.2% in the 5°5cm index, 1.8% in the 10°2cm index, and 7.0% in the 10°5cm index. These results clearly show that the method proposed in the present invention has achieved significant improvements compared with the prior art on the REAL275 dataset. Whether compared with the method that only uses RGB features or the method that uses RGB-D features simultaneously, the proposed solution of the present invention shows the best results under multiple evaluation metrics.
[0186] Table 2 Quantitative analysis and comparison results of 3D shape reconstruction of the present invention with advanced methods on the REAL275 dataset
[0187]
[0188] According to the results in Table 2, it can be clearly seen that the method of the present invention has achieved the lowest shape reconstruction error in the three object categories of bottles, cans, and laptops in the REAL275 dataset. For the two categories of bowls and cameras, the error of the method of the present invention is only 0.05 worse than the best SGPA technology, and in the category of cups, the error is only 0.14 worse than the best SPD technology. The average error of the six categories is lower than that of all other methods. Compared with the best SGPA technology, the error is reduced by 0.44, thus achieving the best 3D shape reconstruction result. These results fully prove the superiority of the method of the present invention in category-level object pose estimation, especially its excellent performance in shape reconstruction in categories such as bottles, cans, and laptops. This further proves the effectiveness of the method in dealing with the pose estimation task of complex structure objects.
[0189] Reference Figure 3 , it can be clearly observed that the method proposed in the present invention is closer to the true label (white bounding box) than the SGPA method in terms of object pose and size estimation. This means that the method of the present invention can better capture the geometric features of the object, thereby achieving more accurate pose and size estimation, and effectively reducing the error between the true label.
[0190] Reference Figure 4, it can be clearly observed that the 3D shape reconstructed by the method proposed in the present invention is very close to the true shape of the object. This proves that the method of the present invention has achieved excellent performance in restoring the three-dimensional shape of the object from point cloud data. As can be seen from the figure, the method of the present invention can accurately capture the geometric structure and details of the object, realizing high-quality three-dimensional shape reconstruction.
[0191] In summary, the present invention proposes a category-level object pose estimation method based on graph attention network. By using a point cloud graph attention network composed of a graph attention encoder and an iterative non-parametric decoder, unique geometric features are extracted from the observed object point cloud, and the structural information of the object is gradually perceived from local to global. Subsequently, a shape prior adaptation mechanism is used to regress the normalized coordinates of the object, and finally, the six-degree-of-freedom pose and size information of the object are obtained through the Umeyama algorithm.
[0192] The method proposed in the present invention has achieved state-of-the-art performance in the category-level object pose estimation task through experiments on the REAL275 dataset. Its innovation lies in the use of graph attention network to realize the learning and aggregation of multi-scale geometric features, and the introduction of shape prior adaptation mechanism, thus significantly improving the accuracy of pose estimation for objects with complex structures. This method brings new ideas and breakthroughs to the field of category-level object pose estimation, and has important research and application value in the fields of machine vision and three-dimensional perception.
[0193] The above is only a preferred specific embodiment of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention, according to the technical solution and inventive concept of the present invention, makes equivalent substitutions or changes, and should be covered by the protection scope of the present invention.
Claims
1. A category-level 6D object pose estimation method based on a point cloud graph attention network, characterized in that, Including: S1. Preprocess the input RGB-D image data to extract the observed point cloud of the object under the depth camera; S2. Use the point cloud graph attention network to extract multi-scale local-to-global object structure features from the observed point cloud; S3. Utilize the shape prior adaptation mechanism and the category shape prior point cloud to reconstruct the 3D point cloud model of the object and regress the normalized NOCS coordinates of the object; S4. Calculate the similarity transformation between the reconstructed NOCS coordinates and the observed point cloud through the Umeyama algorithm to obtain the pose and size information of the object; The point cloud graph attention network is an encoder-decoder architecture, and the encoder-decoder architecture includes: A graph attention encoder for extracting multi-scale local-to-global object features from the observed point cloud; The graph attention encoder takes the observed point cloud P o as input: Among them, represents the set of real numbers, N o represents the number of points, and 3 represents the three-dimensional XYZ coordinates of the points; Use a position embedding module to transform the original three-dimensional coordinates of the observed point cloud into high-dimensional feature embeddings; Through the graph attention module, hierarchically extract local-to-global instance geometric features from the input feature embeddings; An iterative non-parametric decoder for aggregating multi-scale geometric features; The position embedding module uses a 3D graph convolutional layer to encode the position information in the observed point cloud. For each observed point in the observed point cloud: Search for the coordinate set of its M nearest neighbor points using the nearest neighbor search algorithm Receptive field of the convolutional kernel for 3D graph convolution: Among them, M represents the number of nearest points, and m represents one of the points, represents the three-dimensional coordinates of one of the points; Calculate the direction vector d in the receptive field obtained by the nearest neighbor search algorithm m,n : d m,n = p m -p n ; And initialize the support point kernel vector k through uniform distribution s : where S represents the number of support points, and each support point k s is a three-dimensional coordinate; Embed the position information of p n into the C0-dimensional feature vector, and the obtained feature vector passes through the ReLU activation function to generate a position embedding Where max represents the maximum operation, < > represents the inner product of vectors, and |||| represents the norm of the vector.
2. The method for category-level 6D object pose estimation based on a point cloud map attention network according to claim 1, wherein The specific steps of S1 include: S11. Use Mask R-CNN to segment and detect the object in the RGB-D image data to obtain the object mask region; S12. Map the object mask region to the depth image of the object to obtain the object depth region; S13. Use the camera parameters to convert the depth information of the object into the three-dimensional point cloud of the object to generate the observed point cloud data observed by the camera.
3. A category-level 6D object pose estimation method based on a point cloud map attention network according to claim 2, characterized in that, The graph attention module performs multi-stage operations on the position embedding: Among them, denoted as N o represents the number of points, and C0 represents the feature dimension of each point; Each stage has a different hidden dimension C i At each stage i, three different point-wise feature extraction layers are applied respectively to convert the input point features into the corresponding dimension C i : i) Graph convolutional layer GCL; ii) Pointwise self-attention layer PSAL; iii) Feed-forward layer FFN.
4. A category-level 6D object pose estimation method based on a point cloud map attention network according to claim 3, characterized in that The graph convolutional layer GCL extracts the local geometric features F of an object from the input point feature embeddings by leveraging the graph structure defined by the neighboring points of the points i : Among them, N i represents the number of points in the i-th stage, and C i represents the feature dimension of each point; The graph convolutional layer GCL includes a 3D graph convolutional layer and a ReLU function; The pointwise self-attention layer PSAL uses a point cloud self-attention mechanism to extract global geometric features from local geometric features The pointwise self-attention layer projects local geometric features using a shared multi-layer perceptron network into query vectors, key vectors, and value vectors, denoted as Q i , K i , and V i : Among them, W i Q , W i K and W i V are matrices of dimension C i ×C i ; By calculating the pointwise attention weights between the query vector and the value vector to capture the global geometric relationships between different points; The global geometric features are obtained through the dot product operation of the attention weights and the value vectors The feed-forward layer FFN uses a shared multi-layer perceptron network and a ReLU activation function to generate the final output for each stage of the graph attention module; At the local geometric feature F i a residual connection is added before and after: The obtained geometric features are used as the input point features for the next stage; Geometric features After global maximum pooling operation and repetition operation, the global shape features of the object are obtained Among them, C4 = C5, which describes the global shape information of the instance.
5. A category-level 6D object pose estimation method based on a point cloud map attention network according to claim 4, characterized in that The graph attention module includes two graph max pooling layers, and the output of the graph attention module includes five parts: In each stage, the graph convolutional layer GCL, the pointwise self-attention layer PSAL, and the feed-forward layer FFL are stacked in sequence to extract object geometric features from the input point features in a pointwise manner.
6. The category-level 6D object pose estimation method based on a point cloud map attention network according to claim 5, characterized in that The iterative non-parametric decoder includes: Hypothesis is the i-th stage of the graph attention module, where the geometric feature in i ∈ {2, 3, 4} is the downsampled point set of In each iteration of the $i$-th stage, the nearest neighbor search algorithm is used to search for each point $p$ in the point set of the previous stage in n,i-1 the nearest neighbor point $q$ in the $i$-th stage n,i ; The nearest neighbor point q of each point p n,i-1 is determined, and then the features of point q n,i are propagated to point p n,i ; n,i-1 Use the updated features of the points in the (i - 1)-th stage to update the features of the points in the (i - 2)-th stage; The whole process is carried out iteratively until the features of all points in the first stage are determined and applied to the output of each stage of the graph attention module; When all the multi-scale geometric features are aligned to the same number of points They will be pass through a concatenation operation With the positional embedding And the global shape feature To be aggregated to generate the final geometric feature Among them, its characteristic dimension is the sum of the dimensions of all features 7. A category-level 6D object pose estimation method based on a point cloud map attention network according to claim 6, characterized in that The specific steps of S3 include: Given the observed point cloud and the corresponding class shape prior point cloud where N o and N r represent the number of points respectively, and each point is the three-dimensional coordinates of XYZ; Extract point-wise prior features from the class shape prior Point-wise prior features are local prior features and global prior features cascaded Generate local prior features using a three-layer multi-layer perceptron network Generate global prior features based on the local prior features using another two-layer multi-layer perceptron network where D1 and D2 are the dimensions of the features respectively; Use a ReLU activation function after each multi-layer perceptron layer, and use an adaptive max pooling operation after the last ReLU activation function to generate global prior features; After obtaining the prior features and geometric features a shape prior adaptation mechanism is adopted to generate a deformation field and a correspondence matrix from the features respectively, where N r and N o are both the number of points and are used for NOCS coordinate regression; each line d i represents the deformation from the prior point cloud to the reconstructed point cloud of each point: Each line of The sum of elements is 1, representing the observed point cloud Between each point in And all points in its reconstructed point cloud, representing the soft correspondence relationship The shape prior adaptation stage uses two parallel networks, each consisting of a three-layer multi-layer perceptron network, to respectively regress the deformation field and the correspondence matrix Match with to obtain the NOCS coordinates of the object:
8. A method for category-level 6D object pose estimation based on a point cloud map attention network according to claim 7, characterized in that, The observed point cloud and its reconstructed NOCS coordinates Use the Umeyama algorithm combined with the RANSAC algorithm to calculate the 6D object pose and 3D dimensions.
Citation Information
Patent Citations
Mobile robot map construction method based on graph neural network feature extraction and matching, storage medium and equipment
CN114707611A