Class-level 6D pose estimation method and device based on double-feature fusion and channel attention, and electronic equipment
By combining the semantic features of DINOv2 and the geometric features of PointNet++, and using the channel attention mechanism and instance adaptive key point detection module, the problem of insufficient accuracy and robustness of pose estimation in the prior art is solved, and more efficient pose prediction is achieved.
Patent Information
- Application Number
- CN202510332610.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-20
- Publication Date
- 2025-06-17
AI Technical Summary
The existing class-level pose estimation methods have limitations in shape priors, the singularity of feature extraction and the lack of local features, resulting in insufficient accuracy and robustness of pose estimation.
The dual-feature fusion and channel attention method is adopted, combined with the semantic features of DINOv2 and the geometric features of PointNet++, feature enhancement and fusion are performed through the SE channel attention mechanism, and pose estimation is performed using the instance adaptive key point detection module IAKD-ECA and multi-layer perceptron MLP.
It improves the generalization ability and robustness of pose estimation, enhances feature expression ability, and ensures the accuracy and stability of pose prediction.
Smart Images

Figure CN120163875A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of computer vision and deep learning, and particularly relates to a category-level 6D pose estimation method, device and electronic device based on dual feature fusion and channel attention, which is applicable to application scenarios such as 3D object detection, robot vision perception, autonomous driving and augmented reality. Specifically, the present invention combines the geometric features of an object with the prior knowledge of semantic categories, and uses a variety of channel attention mechanisms to optimize the pose estimation process, so as to improve the accuracy and robustness of target pose prediction. Background Art
[0002] Object pose estimation is a key technology in the field of computer vision. Its core task is to predict the rotation and displacement of an object in three-dimensional space, and it is widely used in many fields such as robot operation, augmented reality, and autonomous driving. For example, in autonomous driving, object pose estimation algorithms can identify and track dynamic targets, providing the vehicle with more accurate environmental perception capabilities; in robot operation, accurate object pose information helps the robotic arm to perform precise grasping and manipulation. Currently, pose estimation methods based on deep learning have achieved rapid development, and among them, instance-level pose estimation methods rely on CAD models of specific objects for training. However, this method has weak generalization ability when encountering new category objects, which limits its practical application. To solve this problem, researchers have begun to study category-level pose estimation methods, that is, predicting the 6D pose of an object without relying on a CAD model. This method needs to handle the shape variations within the same category, and how to effectively model the intra-class shape variations has become the focus of current research.
[0003] The existing category-level pose estimation methods mainly model the intra-class object structure by constructing a category-average shape as prior information, and perform pose estimation based on this. However, these methods usually assume that the category-average shape can accurately represent the geometric structures of all objects, but in reality, objects of the same category may have large morphological differences, resulting in the invalidation of this assumption and thus affecting the accuracy of pose estimation. In addition, important progress has been made in recent large-scale self-supervised vision models, and DINOv2 has received extensive attention due to its excellent semantic feature extraction ability. Research shows that DINOv2 can capture high-level semantic information at the category level, and compared with traditional shape-prior-based methods, the feature representation ability of DINOv2 has better generalization at the object category level. Therefore, researchers have begun to explore combining the semantic information of DINOv2 with geometric features to enhance the accuracy of pose estimation.
[0004] However, the existing category-level pose estimation methods still have the following problems: (1) Limitations of shape priors: Traditional methods rely on the category-average shape as the shape prior, but it is difficult to adapt to intra-class morphological variations, which affects the robustness of pose estimation. (2) Singularity of feature extraction: Existing methods usually only use geometric features or semantic features, without fully combining the advantages of both. (3) Lack of local features: The pose estimation task requires high-precision local features, but existing methods may lead to the loss of local details during the feature extraction process. Therefore, current category-level pose estimation still faces great challenges and there is an urgent need for a method that can predict object poses more stably and robustly to effectively improve the adaptability to objects with different morphologies and the overall prediction accuracy of pose estimation. Summary of the Invention
[0005] The purpose of the present invention is to overcome the limitations of the existing category-level pose estimation methods, and provide a category-level 6D pose estimation method, device and electronic device based on dual feature fusion and channel attention, which can effectively solve the problems existing in the prior art: (1) Limitations of shape priors: Traditional methods rely on the category-average shape as the shape prior, but cannot accurately adapt to the morphological diversity of intra-class objects, thus affecting the accuracy of pose estimation. (2) Singularity of feature extraction: Existing methods often only use geometric features or semantic features, and fail to fully combine the advantages of both, resulting in insufficient generalization ability of pose estimation. (3) Lack of local features: Existing methods are prone to losing key local information during the feature extraction process, affecting the final pose prediction accuracy.
[0006] To solve the above technical problems, the present invention adopts the following technical solutions.
[0007] A category-level 6D pose estimation method based on dual feature fusion and channel attention of the present invention adopts a system framework composed of a feature extraction module, an instance adaptive key point detection module IAKD-ECA and a pose estimation module; the system also includes a training module, which is trained using the AdamW optimizer and combined with a linear learning rate warm-up strategy to improve the convergence stability of the model;
[0008] The method includes the following steps:
[0009] Step 1, Feature extraction: Obtain the input RGB-D image, use MaskRCNN for target instance segmentation to extract the target object region; respectively use DINOv2 to extract the semantic features of the object, and extract the geometric features through PointNet++; input the extracted semantic and geometric features into the SE channel attention mechanism for feature enhancement respectively, and fuse them by means of channel splicing to form the final feature vector;
[0010] Step 2, Adaptive Key Point Detection: Construct an Instance Adaptive Key Point Detection Module IAKD-ECA that integrates the ECA attention mechanism, initialize a set of class-shared learnable query vectors, and perform instance adaptive transformation using the fused features to generate a key point detector for a specific instance; calculate the similarity between the key point detector and the object features, generate a key point heat map, and obtain the final key point position through weighted summation; at the same time, use a diversity loss to constrain the uniform distribution of key points and introduce a chamfer distance loss to ensure that key points are close to the object surface, improving the accuracy of key point detection.
[0011] Step 3, Pose Estimation: Based on the detected key points and their features, use a multi-layer perceptron MLP and SE attention mechanism for feature mapping to predict the positions of key points in the normalized coordinate space NOCS; then input the key point features into a pose regression network to regress the rotation matrix R, translation vector T, and scale factor S respectively. Among them, the rotation parameter R is represented in 6D, and the translation parameter T is calculated by the residual between the predicted reference value and the point cloud mean to ensure the stability and accuracy of pose estimation.
[0012] The process of the above-mentioned Step 1 includes:
[0013] Given an RGB-D image, first use MaskRCNN to obtain the segmentation mask and class label of each object; for each segmented object, use the segmentation mask to obtain the cropped RGB image I obj and the corresponding point cloud P obj ; for I obj , use DINOv2 to extract semantic features F r ; for P obj , use PointNet++ to extract point cloud features F p ; then input the semantic feature F r and the point cloud feature F p into the channel attention layer for feature enhancement respectively, and obtain the enhanced features and Finally, connect and to form F obj as the input of the subsequent network.
[0014] The process of the above-mentioned Step 2 includes:
[0015] To establish a robust correspondence between the observed image points RGB, RGB-D, and the normalized object coordinate space NOCS, design an Instance Adaptive Key Point Detection Module IAKD-ECA that integrates an efficient attention mechanism, which can adaptively detect the key points of instances with different shapes, thus more effectively displaying the object and avoiding paying attention to abnormal points;
[0016] Initialize a detector that shares a set of learnable queries across categories where N kpt and C represent the number of key points and the feature dimension respectively, and each query represents a key point detector; input F obj into Q cat , and this process transforms Q cat into an instance-adaptive detector conditioned on F obj First, this transformation process injects the initialized Q cat into the ordinary attention layer Att-Layers for local enhancement, and then injects it into the ECA module. Relying on the self-adaptability and efficient feature representation of the ECA module, an instance-adaptive detector Q ins is generated; then, calculate the cosine similarity between Q ins and F obj to generate the key point heat map Finally, the key points in the camera space and their corresponding features
[0017] P kpt = softmax(H)×P obj (1)
[0018] F kpt = softmax(H)×F obj (2)
[0019] Use a set of sparse key points to represent the geometric information of the object. The detected key points will gather in small areas and concentrate on non-surface or abnormal points; to solve this problem, make the key points evenly distributed in different parts of the object, and further use the diversity loss L div to force the detected key points to disperse from each other; at the same time, to make the key points located on the surface of the object and exclude outliers at the same time, use a chamfer distance loss L ocd to constrain the distribution of P kpt ;
[0020]
[0021]
[0022] where P kpt i , P kpt j represent the i-th and j-th key points respectively, and P obj * Represent key points on the object surface and remove outliers; by constraining the key points to be close to P obj * , the IAKD-ECA module can automatically learn to filter out outliers during training.
[0023] The process of step 3 includes:
[0024] Use the GFGA module in AG-Pose to obtain the key point feature F with geometric information gfga , and then based on the pose estimation module in DPDN, design a more accurate pose estimation method; pair F gfga , and apply MLP to extract features, and then predict the NOCS coordinates according to F gfga to obtain
[0025]
[0026] Next, by connecting the key point P in the camera space kpt and its corresponding feature F kpt , apply a variant of MLP and SE model, relying on its self-adaptive fusion of local features and global information to obtain the feature vector f pose :
[0027]
[0028] Finally, apply three parallel MLPs to regress R, t, and s respectively:
[0029] R, t, s = [MLP R (f pose ), MLP t (f pose ), MLP s (f pose )] (7).
[0030] In the feature extraction step, the semantic features are extracted by DINOv2 and enhanced by SE channel attention, while the geometric features are extracted by PointNet++ to ensure the full fusion of semantic information and geometric information.
[0031] The instance adaptive key point detection module IAKD-ECA adopts the ECA attention mechanism to optimize key point detection, which can adapt to different morphological instances and improve the accuracy and stability of key point detection.
[0032] The key point heat map is generated by calculating the cosine similarity between the key point detector and the object features, and the final key point position is calculated by weighted summation.
[0033] A 6D pose estimation device for category level based on dual feature fusion and channel attention of the present invention includes:
[0034] A feature extraction module, which is used to obtain the input RGB-D image, perform target instance segmentation using MaskRCNN, and extract the target object area; respectively use DINOv2 to extract the semantic features of the object, and extract the geometric features through PointNet++; input the extracted semantic and geometric features into the SE channel attention mechanism for feature enhancement respectively, and fuse them by means of channel splicing to form the final feature vector; among them, given the RGB-D image, first use MaskRCNN to obtain the segmentation mask and category label of each object; for each segmented object, use the segmentation mask to obtain the cropped RGB image I obj and the corresponding point cloud P obj ; for I obj , use DINOv2 to extract the semantic feature F r ; for P obj , use PointNet++ to extract the point cloud feature F p ; then input the semantic feature F r and the point cloud feature F p into the channel attention layer for feature enhancement respectively, and obtain the enhanced features and Finally, connect and to form F obj as the input of the subsequent network;
[0035] A key point adaptive detection module, which is used to construct an instance adaptive key point detection module IAKD-ECA integrating the ECA attention mechanism, initialize a group of class-shared learnable query vectors, and perform instance adaptive transformation using the fused features to generate a key point detector for a specific instance; calculate the similarity between the key point detector and the object features, generate a key point heat map, and obtain the final key point position through weighted summation; at the same time, adopt a diversity loss to constrain the uniform distribution of key points, and introduce a chamfer distance loss to ensure that the key points are close to the object surface, improving the key point detection accuracy; among them, in order to establish a robust correspondence between the observed image points RGB, RGB-D and the normalized object coordinate space NOCS, design an instance adaptive key point detection module IAKD-ECA integrating an efficient attention mechanism, which can adaptively detect the key points of instances with different shapes, so as to display the object more effectively and avoid paying attention to abnormal points;
[0036] Initialize a group of detectors with class-shared learnable queries where N kptC and represent the number of key points and the feature dimension respectively, and each query represents a key point detector; Feed F obj into Q cat , and this process transforms Q cat into an instance-adaptive detector conditioned on F obj ; First, this transformation process injects the initialized Q cat into the ordinary attention layer Att-Layers for local enhancement, and then injects it into the ECA module. Relying on the self-adaptability and efficient feature representation of the ECA module, an instance-adaptive detector Q ins is generated; Then, calculate the cosine similarity between Q ins and F obj to generate the key point heat map Finally, the key points in the camera space and their corresponding features
[0037] P kpt = softmax(H) × P obj (1)
[0038] F kpt = softmax(H) × F obj (2)
[0039] Use a set of sparse key points to represent the geometric information of the object. The detected key points will gather in small areas and concentrate on non-surface or abnormal points; To solve this problem and make the key points evenly distributed in different parts of the object, further use the diversity loss L div to force the detected key points to disperse from each other; At the same time, to make the key points located on the object surface and exclude outliers, use a chamfer distance loss L ocd to constrain the distribution of P kpt ;
[0040]
[0041]
[0042] where, P kpt i , P kpt j represent the i-th and j-th key points respectively, and P obj * represents the key points on the object surface and removes the outliers; By constraining the key points to be close to P obj * , the IAKD-ECA module can automatically learn to filter out outliers during training.
[0043] The pose estimation module is used to perform feature mapping based on the detected key points and their features, using a multi-layer perceptron (MLP) and a SE attention mechanism to predict the positions of the key points in the normalized coordinate space (NOCS). Then, the key point features are input into a pose regression network to respectively regress the rotation matrix R, the translation vector T, and the scale factor S. Among them, the rotation parameter R is represented in 6D notation, and the translation parameter T is calculated by the residual between the predicted reference value and the point cloud mean to ensure the stability and accuracy of pose estimation. Among them, the GFGA module in AG-Pose is used to obtain the key point features F with geometric information. gfga , and then based on the pose estimation module in DPDN, a more accurate pose estimation method is designed; for F gfga pairing is performed, and MLP is applied to extract features, and then according to F gfga to predict the NOCS coordinates to obtain
[0044]
[0045] Next, by connecting the key point P in the camera space kpt and its corresponding feature F kpt , a variant of the MLP and SE model is applied, which can adaptively fuse local features and global information to obtain the feature vector f pose :
[0046]
[0047] Finally, three parallel MLPs are respectively applied to regress R, t, and s:
[0048] R, t, s = [MLP R (f pose ), MLP t (f pose ), MLP s (f pose )] (7).
[0049] An electronic device according to the present invention includes a memory, a processor, and a computer program stored on the memory and executable on the processor. The processor, when executing the computer program, implements the above-mentioned category-level 6D pose estimation method based on dual feature fusion and channel attention.
[0050] A non-transitory computer-readable storage medium according to the present invention stores a computer program thereon. When the computer program is executed by a processor, it implements the above-mentioned category-level 6D pose estimation method based on dual feature fusion and channel attention.
[0051] Compared with the prior art, the present invention has the following advantages and beneficial effects:
[0052] 1. A category-level 6D pose estimation method, device, and electronic device based on dual feature fusion and channel attention are proposed. By combining the semantic features of DINOv2 and the geometric features of PointNet++, the generalization ability of pose estimation is improved, and the influence of intra-class shape variations on the estimation accuracy is effectively overcome.
[0053] 2. An instance-adaptive keypoint detection method integrating the ECA attention mechanism is proposed, which avoids the limitations of the fixed keypoint method, makes keypoint detection more adaptable, and improves the stability of pose estimation.
[0054] 3. The multi-channel attention mechanism and the MLP-SE variant are adopted to effectively enhance the feature expression ability, enabling the pose estimation model to more accurately predict rotation, translation, and scale parameters.
[0055] 4. Comprehensive verification is carried out on multiple public datasets, including the CAMERA25, REAL275, and HouseCat6D datasets. The experimental results show that the method of the present invention achieves state-of-the-art performance with or without the average shape prior. BRIEF DESCRIPTION OF THE DRAWINGS
[0056] Figure 1 is a flowchart of the method of an embodiment of the present invention.
[0057] Figure 2 is an instance keypoint adaptive detector of an embodiment of the present invention.
[0058] Figure 3 is the performance curve (CAMERA25 + REAL275) of an embodiment of the present invention at different thresholds.
[0059] Figure 4 is the performance curve (HouseCat6D) of an embodiment of the present invention at different thresholds.
[0060] Figure 5 is the visualization diagram of the object pose estimation prediction of an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0061] The present invention proposes a category-level 6D pose estimation method, device, and electronic device based on dual-feature fusion and channel attention, constructing a complete framework including a feature extraction module, an instance-adaptive key point detection module, and a pose estimation module. Among them, the feature extraction module is responsible for extracting information from the input data, and fusing the semantic features extracted by DINOv2 and the geometric features extracted by PointNet++, ensuring the complementarity of global and local information to enhance the feature expression ability. The instance-adaptive key point detection module (IAKD-ECA) uses a key point detector integrating the ECA attention mechanism to achieve the adaptive detection of key points of instances with different shapes, improving the accuracy and stability of key point localization. The pose estimation module combines a multi-layer perceptron and the SE attention mechanism to fully explore the correlation between local geometric information and global semantic information, and finally accurately regresses the 6D pose parameters.
[0062] The following further elaborates on the present invention with reference to the accompanying drawings.
[0063] (1) Dataset
[0064] The present invention conducts experiments on three mainstream category-level object pose estimation datasets, namely NOCS-REAL275, NOCS-CAMERA25, and HouseCat6D, to verify the effectiveness and applicability of the proposed method. Among them, NOCS-CAMERA25 and NOCS-REAL275 are the most commonly used standard benchmark datasets for category-level pose estimation, while HouseCat6D provides more complex real-world object data, which helps to evaluate the generalization ability of the present invention.
[0065] NOCS-CAMERA25 is a synthetic RGB-D dataset, containing 1,085 instances from 6 different categories, with a total of 300K synthetic images. Among them, 25,000 images of 184 instances are used for evaluation, and the rest are used for model training. This dataset provides synthetic images of multiple categories, which helps to evaluate the generalization ability of the method on different object categories.
[0066] NOCS-REAL275 is a more challenging real-world dataset, having the same 6 categories as CAMERA25, but the images are all from real scenes. This dataset contains 7K images, covering 13 different scenes. Among them, 2,750 images of 6 scenes are used for verification, and each category contains 3 unseen instances, which are used to evaluate the robustness of the method in practical applications.
[0067] HouseCat6D is a comprehensive multi-modal real-world dataset that contains high-fidelity 3D models of 10 categories and 194 household items. This dataset includes objects from 41 scenes, covering transparent and reflective materials, with a wider range of perspectives and complex occlusions. At the same time, it does not provide category annotations, so it is closer to real application scenarios. This dataset is used to verify the adaptability and generalization ability of the present invention under different scenarios and object materials.
[0068] (2) Network architecture and experimental configuration
[0069] The method of the present invention is as shown in the attached drawings Figure 1 and includes core components such as a feature extraction module, an instance-adaptive key point detection module, and a pose estimation module. The feature extraction module uses DINOv2 to extract semantic features, combines PointNet++ to extract geometric features, and performs feature fusion through a channel attention mechanism to enhance the feature expression ability. The key point detection module uses the ECA attention mechanism to optimize key point detection, enabling it to adapt to different-shaped instances and improving the accuracy and stability of key point localization. The pose estimation module combines MLP and SE attention mechanisms to fully exploit the correlation between local geometric information and global semantic information, and finally regresses 6D pose parameters (rotation, translation, scale).
[0070] All experiments of the present invention were carried out on an NVIDIA RTX 3090Ti GPU. All experiments used a single GPU for training and inference, and the batch size was set to 24. In data preprocessing, to ensure a fair comparison, the present invention uses MaskRCNN to perform target instance segmentation on the input data to obtain object masks, and normalizes and resizes the cropped images to 224×224 before feature extraction. At the same time, the number of points in the point cloud is fixed at N = 1024. In the experiment, the number of key points was set to a fixed value, the local range of each key point in the GAFA (geometric feature enhancement) module was set to K = 16, and the feature dimension was set to C = 256 to ensure sufficient feature expression ability and computational efficiency. The hyperparameters in the loss function were set as follows: λ1 = 1.0, λ2 = 5.0, λ3 = 1.0, λ4 = 0.3 to optimize the key point matching, rotation error, and translation error losses.
[0071] The present invention is trained on synthetic data and real data on the REAL275 dataset, while only synthetic data is used for training on the CAMERA25 dataset to evaluate the generalization ability of the method in real scenarios. During the training process, the AdamW optimizer is used, combined with a linear learning rate warm-up strategy, followed by linear decay to ensure stable convergence of the model in the initial stage of training. Among them, the encoder learning rate is set to 1e-4, the decoder learning rate is set to 5e-4, the weight decay is set to 0.01, and 40K iterations of training are performed.
[0072] In order to improve the robustness of the model, cosine similarity loss is used to optimize key point detection during training to improve the stability of key point distribution. Rotation error loss and translation error loss are used to ensure the accuracy of object pose estimation, and instance matching loss is introduced to optimize key point matching. This paper compares the existing category-level 6D pose estimation method, evaluates performance under multiple threshold settings, and conducts experimental analysis for different categories and scenarios to verify the adaptability of the method in complex environments.
[0073] (3) Evaluation indicators
[0074] This paper adopts two commonly used evaluation metrics to comprehensively evaluate the performance of the model in the category-level 6D pose estimation task. 3DIoU: reports the average precision of the intersection-over-union (IoU) of 3D bounding boxes at thresholds of 50% and 75%, which incorporates the pose and size of the object. ο m: For direct evaluation of rotation and translation errors. Only when the rotation error is less than n ο Predictions with a translation error less than mcm are considered correct.
[0075] The present invention proposes a category-level 6D pose estimation method, device and electronic device based on dual feature fusion and channel attention, and constructs a complete framework including feature extraction, key point detection and pose estimation. First, the feature extraction module extracts key information from the input data, uses DINOv2 to obtain semantic features, and combines PointNet++ to extract the geometric structure information of the object to ensure the complementarity of global and local features, thereby improving the feature expression ability. Subsequently, the instance adaptive key point detection module (IAKD-ECA) uses the ECA attention mechanism to optimize key point detection, which can adapt to instances of different shapes and improve the positioning accuracy and stability of key points. Finally, the pose estimation module combines the multi-layer perceptron with the SE attention mechanism to deeply explore the relationship between local geometric features and global semantic information, and finally realizes the accurate prediction of the 6D pose of the object. The method described in the present invention comprises the following steps:
[0076] Step 1: Feature extraction
[0077] Given an RGB-D image, we first use the off-the-shelf MaskRCNN to obtain the segmentation mask and category label of each object. For each segmented object, we use the segmentation mask to get the cropped RGB image I obj and the corresponding point cloud P obj For I obj , using DINOv2 to extract semantic features F r For P obj, use PointNet++ to extract point cloud features F p Then the two features are input into the channel attention layer for feature enhancement, and the enhanced features are obtained respectively. and Finally, and Connect to form F obj As the input of the subsequent network.
[0078] Step 2: Instance-adaptive keypoint detection with ECA
[0079] In order to establish a robust correspondence between observed image points (RGB or RGB-D) and normalized object coordinate space (NOCS), geometric information is particularly important, and the direct approach to provide geometric information for each point is to use a common attention mechanism to aggregate the features of all other points, but this approach will cause the problem of excessive computational cost due to the large number of points. Therefore, it is envisaged to use a set of sparse key points to represent different instance shapes. However, different instance shapes are different, and the instance model cannot be accessed during the reasoning process, so some fixed key point detection methods are not applicable. Therefore, the present invention designs an instance adaptive key point detection module (IAKD-ECA) that incorporates an efficient attention mechanism, which can adaptively detect key points of instances with different shapes, which can more effectively display objects and avoid focusing on abnormal points.
[0080] The process of IAKD-ECA is shown in the attached figure Figure 2 Initialize a set of class-shared learnable query detectors Where N kpt and C represent the number of key points and feature dimensions respectively, and each query represents a key point detector. obj Input to Q cat , this process will Q cat Convert to F obj Conditional instance adaptive detector First, the conversion process is to initialize Q cat Injected into the common attention layer (Att-Layers) for local enhancement, and then injected into the ECA module. With the adaptability and efficient feature representation of the ECA module, the instance adaptive detector Q is generated. ins ; Then, calculate Q ins and F obj The cosine similarity between them generates a key point heat map Finally, the key points in the camera space are obtained by weighted summation And its corresponding characteristics
[0081] P kpt = softmax(H) × P obj (1)
[0082] F kpt = softmax(H) × F obj (2)
[0083] The idea of the present invention is to use a set of sparse key points to represent the geometric information of an object. Since the detected key points should be evenly distributed on the surface of the object, it is actually found that the key points will gather in small areas and often concentrate on non-surface or abnormal points. To solve this problem and make the key points evenly distributed in different parts of the object, a diversity loss L div is further used to force the detected key points to be dispersed from each other. At the same time, in order to make the key points located on the object surface and exclude outliers, a chamfer distance loss L ocd is used to constrain the distribution of P kpt .
[0084]
[0085] where P kpt i , P kpt j represent the i-th and j-th key points respectively, and P obj * represents the key points on the object surface with outliers removed. By constraining the key points to be close to P obj * , the IAKD-ECA module can automatically learn to filter out outliers during training.
[0086] Step 3, Pose Estimation
[0087] The present invention uses the GFGA module in AG-Pose to obtain the key point features F with geometric information gfga , and then designs a more accurate pose estimation method based on the pose estimation module in DPDN. Pair F gfga , apply MLP to extract features, and then predict the NOCS coordinates according to F gfga to obtain
[0088]
[0089] Next, by connecting the key point P kpt in the camera space and its corresponding feature F kpt , applying a variant of the MLP and SE model, and relying on its self-adaptive fusion of local features and global information, the feature vector f is obtainedpose :
[0090]
[0091] Finally, three parallel MLPs are applied to regress R, t, and s respectively:
[0092] R, t, s = [MLP R (f pose ), MLP t (f pose ), MLP s (f pose )] (7).
[0093] Based on the experimental results of the present invention:
[0094] As shown in Table 1, the experimental results of the present invention on the REAL275 dataset exceed the existing methods. It should be emphasized that the method proposed by the present invention does not use shape priors. The accuracy in the strictest metric scale of 5° 2cm reaches 60.4%, which is 5.7% and 4.2% higher than AG-Pose and SecondPose respectively. Compared with MH6D using shape priors, the method is 7.4% higher. In other metric scales of 5° 5cm, 10° 2cm, and 10° 5cm, the indicators are 66.6%, 78.4%, and 86.5% respectively, which are 5.5%, 6.4%, and 4.5% higher than MH6D respectively. Compared with AG-Pose and SecondPose, the method is 4.9%, 3.7%, 3.4% and 3.0%, 3.7%, 0.5% higher respectively.
[0095] Table 1 Quantitative comparison with existing methods on the REAL275 dataset
[0096]
[0097] Table 2 shows that the quantitative results on the CAMERA25 dataset are better than the existing methods and achieve the best performance under most metrics. Specifically, TC-Pose is 0.5%, 0.5%, 0.6%, and 0.9% higher than AG-Pose in the metric scales of IoU 50 , IoU 75 , 5° 2cm, and 10° 2cm in terms of accuracy respectively.
[0098] Table 2 Quantitative comparison with existing methods on the CAMERA25 dataset
[0099]
[0100]
[0101] The performance on the HouseCat6D dataset is shown in Table 3. The results indicate that the method of the present invention outperforms recent excellent methods in all metrics. The AG-Pose* in the table is the result reproduced on the HouseCat6D dataset based on the improved version of AG-Pose. Specifically, the present invention improves by 31%, 11.3%, 10.0%, 28.2%, 20.8% and 23.8%, 11.8%, 11.4%, 20.8%, 20.7% respectively compared with SecondPose and AG-Pose in terms of the precision of IoU 75 , 5°2cm, 5°5cm, 10°2cm and 10°5cm.
[0102] Table 3 Quantitative comparison with existing methods on the HouseCat6D dataset
[0103]
[0104] [1]He Wang,Srinath Sridhar,Jingwei Huang,Julien Valentin,Shuran Song,and Leonidas J.Guibas.Normalized object coordinate space for category-level6d object pose and size estimation.In The IEEE Conference on Computer Visionand Pattern Recognition(CVPR),2019.1,2,3,6,7
[0105] [2]Yan Di,Ruida Zhang,Zhiqiang Lou,Fabian Manhardt,Xiangyang Ji,Nassir Navab,and Federico Tombari.Gpv-pose:Category-level object poseestimation via geometry-guided point-wise voting.In Proceedings of the IEEE / CVF Conference on Computer Vision and Pattern Recognition,pages 6781–6791,2022.3,5,6,8,13
[0106] [3]Jianhui Liu,Yukang Chen,Xiaoqing Ye,and Xiaojuan Qi.Ist-net:Prior-free category-level pose estimation with implicit space transformation,2023.3,6
[0107] [4]Ruiqi Wang,Xinggang Wang,Te Li,Rong Yang,Minhong Wan,andWenyuLiu.Query6dof:Learning sparse queries as implicit shape prior for category-level 6dof pose estimation.In Proceedings of the IEEE / CVF InternationalConference on Computer Vision,pages 14055–14064,2023.6,7
[0108] [5]Jiehong Lin,Zewei Wei,Yabin Zhang,and Kui Jia.Vi-net:Boostingcategory-level 6dobject pose estimation via learning decoupled rotations onthe spherical representations.In Proceedings of the IEEE / CVF InternationalConference on Computer Vision,pages 14001–14011,2023.2,3,5,6,7,8,11,13
[0109] [6] Lin X, Yang W, Gao Y, et al. Instance-adaptive and geometric-aware keypoint learning for category-level 6d object pose estimation[C] / / Proceedings of the IEEE / CVF Conference on Computer Vision and Pattern Recognition. 2024:21040-21049.
[0110] [7] Chen Y, Di Y, Zhai G, et al. Secondpose: Se(3)-consistent dual-stream feature fusion for category-level pose estimation[C] / / Proceedings of the IEEE / CVF Conference on Computer Vision and Pattern Recognition. 2024:9959-9969.
[0111] [8] Meng Tian, Marcelo H Ang, and Gim Hee Lee. Shape prior deformation for categorical 6d object pose and size estimation. In European Conference on Computer Vision, pages 530–546. Springer, 2020. 1,3,6
[0112] [9] Kai Chen and Qi Dou. Sgpa: Structure-guided prior adaptation for category-level 6d object pose estimation. In Proceedings of the IEEE / CVF International Conference on Computer Vision, pages 2773–2782, 2021. 3,6,7
[0113]
[10] Haitao Lin,Zichang Liu,Chilam Cheang,Yanwei Fu,Guodong Guo,and Xiangyang Xue.Sar-net:Shape alignment and recovery network for category-level 6d object pose and size estimation.In Proceedings of the IEEE / CVF Conference on Computer Vision and Pattern Recognition,pages 6707–6717,2022.6,7
[0114]
[11] Jiehong Lin,Zewei Wei,Changxing Ding,and Kui Jia.Category-level 6d object pose and size estimation using selfsupervised deep priordeformation networks.In European Conference on Computer Vision,pages 19–34.Springer,2022.1,6
[0115]
[12] Liu J,Sun W,Liu C,et al.Mh6d:Multi-hypothesis consistencylearning for category-level 6-d object pose estimation[J].IEEE Transactions on Neural Networks and Learning Systems,2024.
[0116]
[13] Wei Chen,Xi Jia,Hyung Jin Chang,Jinming Duan,Linlin Shen,and AlesLeonardis.Fs-net:Fast shape-based network for category-level 6d object poseestimation with decoupled rotation mechanism.In Proceedings ofthe IEEE / CVFConference on Computer Vision and Pattern Recognition,pages 1581–1590,2021.3,6,8,13。
Claims
1. A category-level 6D pose estimation method based on dual feature fusion and channel attention, characterized in that: The system framework is composed of a feature extraction module, an instance adaptive key point detection module IAKD-ECA and a posture estimation module. The system also includes a training module, which uses the AdamW optimizer for training and combines a linear learning rate warm-up strategy to improve the convergence stability of the model. The method comprises the following steps: Step 1, feature extraction: Get the input RGB-D image, use MaskRCNN to perform target instance segmentation, and extract the target object area; use DINOv2 to extract the semantic features of the object, and use PointNet++ to extract the geometric features; input the extracted semantic and geometric features into the SE channel attention mechanism for feature enhancement, and fuse them through channel splicing to form the final feature vector; Step 2, key point adaptive detection: Construct an instance adaptive key point detection module IAKD-ECA that integrates the ECA attention mechanism, initialize a set of category-shared learnable query vectors, and use the fused features to perform instance adaptive transformation to generate a key point detector for a specific instance; calculate the similarity between the key point detector and the object features, generate a key point heat map, and obtain the final key point position through weighted summation; at the same time, use diversity loss to constrain the uniform distribution of key points, and introduce chamfer distance loss to ensure that the key points are close to the object surface, thereby improving the key point detection accuracy; Step 3, pose estimation: Based on the detected key points and their features, the multi-layer perceptron MLP and SE attention mechanism are used for feature mapping to predict the position of the key points in the normalized coordinate space NOCS; then the key point features are input into the pose regression network to regress the rotation matrix R, translation vector T and scale factor S respectively, where the rotation parameter R uses the 6D representation, and the translation parameter T is calculated by the residual between the predicted reference value and the point cloud mean to ensure the stability and accuracy of the pose estimation.
2. According to claim 1, a category-level 6D pose estimation method based on dual feature fusion and channel attention is characterized in that: The process of step 1 includes: Given an RGB-D image, we first use MaskRCNN to obtain the segmentation mask and category label of each object; for each segmented object, we use the segmentation mask to obtain the cropped RGB image I obj And the corresponding point cloud P obj ; For I obj , using DINOv2 to extract semantic features F r For P obj , use PointNet++ to extract point cloud features F p ; Then the semantic feature F r and point cloud features F p Input the channel attention layer respectively for feature enhancement, and obtain the enhanced features respectively and Finally, and Connect to form F obj As the input of the subsequent network.
3. According to claim 1, a category-level 6D pose estimation method based on dual feature fusion and channel attention is characterized in that: The process of step 2 includes: In order to establish a robust correspondence between the observed image points RGB, RGB-D and the normalized object coordinate space NOCS, an instance-adaptive key point detection module IAKD-ECA is designed that integrates an efficient attention mechanism to adaptively detect key points of instances with different shapes, thereby more effectively displaying objects and avoiding focusing on outliers. Initialize a set of class-shared learnable query detectors Where N kpt and C represent the number of key points and feature dimensions respectively, and each query represents a key point detector; obj Input to Q cat , this process will Q cat Convert to F obj Conditional instance adaptive detector First, the conversion process is to initialize Q cat Injected into the ordinary attention layer Att-Layers for local enhancement, and then injected into the ECA module. With the adaptability and efficient feature representation of the ECA module, the instance adaptive detector Q is generated. ins ; Then, calculate Q ins and F obj The cosine similarity between them generates a key point heat map Finally, the key points in the camera space are obtained by weighted summation And its corresponding characteristics P kpt =softmax(H)×P obj (1) F kpt =softmax(H)×F obj (2) Using a set of sparse key points to represent the geometric information of an object, the detected key points will be concentrated in a small area and concentrated on non-surface or abnormal points. To solve this problem, the key points can be evenly distributed in different parts of the object, and the diversity loss L is further used. div To force the detected key points to disperse from each other; at the same time, in order to make the key points located on the surface of the object and exclude outliers, a chamfer distance loss L is used ocd To constrain P kpt Distribution of Among them, P kpt i ,P kpt j Represent the i-th and j-th key points respectively, P obj * Represents the key points on the surface of the object and removes outliers; by constraining the key points to be close to P obj * , the IAKD-ECA module can automatically learn to filter out outliers during training.
4. According to claim 1, a category-level 6D pose estimation method based on dual feature fusion and channel attention is characterized in that: The process of step 3 includes: Use the GFGA module in AG-Pose to obtain key point features F with geometric information gfga Then, based on the pose estimation module in DPDN, a more accurate pose estimation method is designed; gfga Pair them and apply MLP to extract features, then calculate the features according to F gfga To predict NOCS coordinates we get Next, by connecting the key points P in the camera space kpt And its corresponding feature F kpt , applying the variants of MLP and SE models, relying on their adaptive fusion of local features and global information, to obtain the feature vector f pose : Finally, three parallel MLPs are applied to regress R, t, and s respectively: R,t,s=[MLP R (f pose ),MLP t (f pose ),MLP s (f pose )] (7)。 5. According to claim 2, a category-level 6D pose estimation method based on dual feature fusion and channel attention is characterized in that: In the feature extraction step, semantic features are extracted by DINOv2 and enhanced by SE channel attention, while geometric features are extracted by PointNet++ to ensure full integration of semantic information and geometric information.
6. According to claim 3, a category-level 6D pose estimation method based on dual feature fusion and channel attention is characterized in that: The instance adaptive key point detection module IAKD-ECA adopts the ECA attention mechanism to optimize key point detection, which can adapt to different morphological instances and improve the accuracy and stability of key point detection.
7. According to claim 3, a category-level 6D pose estimation method based on dual feature fusion and channel attention is characterized in that: The key point heat map is generated by calculating the cosine similarity between the key point detector and the object features, and the final key point position is calculated by weighted summation.
8. A category-level 6D pose estimation device based on dual feature fusion and channel attention, characterized in that: include: The feature extraction module is used to obtain the input RGB-D image, use MaskRCNN to perform target instance segmentation, and extract the target object area; DINOv2 is used to extract the semantic features of the object, and PointNet++ is used to extract the geometric features; The extracted semantic and geometric features are respectively input into the SE channel attention mechanism for feature enhancement, and fused by channel splicing to form the final feature vector. Given an RGB-D image, MaskRCNN is first used to obtain the segmentation mask and category label of each object. For each segmented object, the segmentation mask is used to obtain the cropped RGB image I obj And the corresponding point cloud P obj ; For I obj , using DINOv2 to extract semantic features F r For P obj , use PointNet++ to extract point cloud features F p ; Then the semantic feature F r and point cloud features F p Input the channel attention layer respectively for feature enhancement, and obtain the enhanced features respectively and Finally, and Connect to form F obj As input to subsequent networks; The key point adaptive detection module is used to construct an instance adaptive key point detection module IAKD-ECA that integrates the ECA attention mechanism, initialize a set of category-shared learnable query vectors, and use the fused features to perform instance adaptive transformation to generate a key point detector for a specific instance; calculate the similarity between the key point detector and the object features, generate a key point heat map, and obtain the final key point position through weighted summation; at the same time, the diversity loss is used to constrain the key points to be evenly distributed, and the chamfer distance loss is introduced to ensure that the key points are close to the object surface, thereby improving the key point detection accuracy; in order to establish a robust correspondence between the observed image points RGB, RGB-D and the normalized object coordinate space NOCS, an instance adaptive key point detection module IAKD-ECA that integrates an efficient attention mechanism is designed, which can adaptively detect the key points of instances with different shapes, thereby more effectively displaying objects and avoiding focusing on abnormal points; Initialize a set of class-shared learnable query detectors Where N kpt and C represent the number of key points and feature dimensions respectively, and each query represents a key point detector; obj Input to Q cat , this process will Q cat Convert to F obj Conditional instance adaptive detector First, the conversion process is to initialize Q cat Injected into the ordinary attention layer Att-Layers for local enhancement, and then injected into the ECA module. With the adaptability and efficient feature representation of the ECA module, the instance adaptive detector Q is generated. ins ; Then, calculate Q ins and F obj The cosine similarity between them generates a key point heat map Finally, the key points in the camera space are obtained by weighted summation And its corresponding characteristics P kpt =softmax(H)×P obj (1) F kpt =softmax(H)×F obj (2) Using a set of sparse key points to represent the geometric information of an object, the detected key points will be concentrated in a small area and concentrated on non-surface or abnormal points. To solve this problem, the key points can be evenly distributed in different parts of the object, and the diversity loss L is further used. div To force the detected key points to disperse from each other; at the same time, in order to make the key points located on the surface of the object and exclude outliers, a chamfer distance loss L is used ocd To constrain P kpt Distribution of Among them, P kpt i ,P kpt j Represent the i-th and j-th key points respectively, P obj * Represents the key points on the surface of the object and removes outliers; by constraining the key points to be close to P obj * , the IAKD-ECA module can automatically learn to filter out outliers during training. The pose estimation module is used to predict the position of the key points in the normalized coordinate space NOCS based on the detected key points and their features using a multi-layer perceptron MLP and SE attention mechanism for feature mapping; the key point features are then input into the pose regression network to regress the rotation matrix R, translation vector T and scale factor S respectively, where the rotation parameter R uses a 6D representation and the translation parameter T is calculated by the residual between the predicted reference value and the point cloud mean to ensure the stability and accuracy of the pose estimation; the GFGA module in AG-Pose is used to obtain the key point feature F with geometric information gfga Then, based on the pose estimation module in DPDN, a more accurate pose estimation method is designed; gfga Pair them and apply MLP to extract features, then calculate the features according to F gfga To predict NOCS coordinates we get Next, by connecting the key points P in the camera space kpt And its corresponding feature F kpt , applying the variants of MLP and SE models, relying on their adaptive fusion of local features and global information, to obtain the feature vector f pose : Finally, three parallel MLPs are applied to regress R, t, and s respectively: R,t,s=[MLP R (f pose ),MLP t (f pose ),MLP s (f pose )] (7)。 9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the computer program, the category-level 6D pose estimation method based on dual feature fusion and channel attention is implemented as described in any one of claims 1 to 7.
10. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the category-level 6D pose estimation method based on dual feature fusion and channel attention is implemented as described in any one of claims 1 to 7.
Citation Information
Cited By
Geometric perception key point-based category-level 6D attitude estimation method
CN121304789A
A class-level 6D pose estimation method based on geometric perception key points
CN121304789B