Category-Level Object 6D Pose Estimation Method and System Based on Dynamic Keypoint Detection
Through a method based on dynamic key point detection, RGB images and point cloud features are extracted and fused, object key points are adaptively extracted, and pose estimation is performed using multi-scale pose prediction network, which solves the problem of poor performance in noise and occlusion situations, and achieves a more efficient and robust pose estimation effect.
Patent Information
- Application Number
- CN202311546440.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-11-20
- Publication Date
- 2025-06-10
- Estimated Expiration
- 2043-11-20
AI Technical Summary
The existing class-level 6D object position estimation method performs poorly in noise points and occlusion situations, and the calculation amount and storage requirements are too large, making it difficult to actually implement it.
Using a method based on dynamic key point detection, by receiving RGB images and point cloud data, image and point cloud features are extracted and fusion is performed. The dynamic key point detection network is used to extract object key points and input them into a multi-scale pose prediction network for pose estimation.
This method can adaptively extract key points of the object under noise and occlusion, improve the accuracy and robustness of pose estimation, reduce the calculation amount and storage requirements, and achieve more efficient algorithm implementation.
Smart Images

Figure CN117456003B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer vision technology, and specifically to a method and system for class-level object 6D pose estimation based on dynamic key point detection. Background Art
[0002] Object 6D pose estimation technology is an important technology in the fields of computer vision and robotics. It is used to accurately determine the pose of a three-dimensional object in six degrees of freedom, namely three-dimensional translation and rotation. This technology has a wide range of applications in many fields, such as automated manufacturing, robot operation, augmented reality, virtual reality, driverless cars, and so on.
[0003] Although the current instance-level 6D object pose estimation method based on fixed key point detection already has good performance and robustness, the characteristic that it is only effective for a single instance makes this type of method still have a large gap from actual implementation. Therefore, a class-level 6D object pose estimation method based on a normalized class coordinate space is proposed. Such methods can detect the object poses of different instances under the same class of objects, have high versatility, and are closer to the needs of actual production.
[0004] The current class-level 6D object pose estimation methods can generally be divided into two categories. One is the method of directly regressing pose parameters based on the extracted features. However, due to the complexity and non-convexity of the three-dimensional space rotation matrix group, it is difficult to optimize this type of method, and it often cannot reach the accuracy and robustness required by the actual application scenario. The other is the method based on the prediction of the coordinates of the dense object coordinate system, aiming to detect the position of each point in the scene in the normalized object coordinate system, and through these corresponding relationships, use PnP or a pose prediction network for post-processing to output the object pose. These methods transform the regression of the object pose into the prediction of the position of each point in the scene in the object coordinate system, making the network easier to fit and optimize. Although the above methods avoid the problem of difficult optimization of the three-dimensional rotation group, there are often many noises in the three-dimensional point cloud. Directly predicting the coordinates of all points in the scene in the object coordinate system will affect the performance of the network due to the existence of noise points. Moreover, the scale of the scene point cloud is often huge. Adopting the method of predicting all points will result in too large a computational overhead and a large storage requirement for the device, which is not conducive to the actual implementation of the algorithm. Therefore, a class-level object 6D pose estimation method based on dynamic key point detection is proposed to solve the problems existing in the above-mentioned existing methods. Summary of the Invention
[0005] (1) Technical Problems to be Solved
[0006] In view of the deficiencies of the prior art, the present invention provides a method and system for 6D pose estimation of category-level objects based on dynamic key point detection, which can adaptively extract the key points of an object from an observation scene and achieve good results even when there are many noise points and serious occlusion in the scene.
[0007] (II) Technical Solution
[0008] To achieve the above objectives, the present invention is implemented through the following technical solutions:
[0009] In a first aspect, a method for 6D pose estimation of category-level objects based on dynamic key point detection is provided, including:
[0010] Receiving image data of an object, where the image data includes an RGB image and a point cloud, and the point cloud is formed by randomly sampling pixels in a depth map and projecting them into the scene in combination with the internal parameters of a camera;
[0011] Respectively extracting the image features of the RGB image and the point cloud features of the point cloud, and splicing and fusing the image features and the point cloud features to obtain the fused features;
[0012] Inputting the fused features into a preset dynamic key point detection network to extract the key points of the object;
[0013] Inputting the key points into a preset multi-scale pose prediction network, aggregating local structure information into the key points to obtain key points with multi-scale information, predicting the positions of the key points in the object space coordinate system through the key point features with multi-scale information, and splicing the positions of the output key points in the scene, the features of the key points in the scene, and the positions and features of the key points in the object space coordinate system to form multiple sets of corresponding relationships, and outputting the final 6D pose of the object through a multi-layer perceptron.
[0014] Preferably, the feature extractor of the RGB image adopts a Resnet18 convolutional neural network.
[0015] Preferably, the step of respectively extracting the image features of the RGB image and the point cloud features of the point cloud specifically includes:
[0016] Inputting the input RGB image into a Resnet18 convolutional neural network to extract the feature map f rgb ∈R h×w×c ;
[0017] Inputting the point cloud into a Pointnet++ point cloud feature extraction network to extract the structural features f point ∈R N ×C ; where N is the number of points in the point cloud.
[0018] Preferably, the process of splicing and fusing the image features and the point cloud features to obtain the fused features specifically includes
[0019] Projecting the structural features of the point cloud onto the feature map of the image through the internal parameters of the camera, and extracting the corresponding feature f of the structural features of the point cloud on the feature map of the image through bilinear interpolation point→rgb ∈R N×C ;
[0020] After splicing the feature map of the image and the structural features of the point cloud, passing through a multi-layer MLP to output the fused feature f fusion ∈R N×C .
[0021] Preferably, the process of inputting the fused features into a preset dynamic key point detection network to extract the key points of the object specifically includes:
[0022] Introducing an attention mechanism and a Transformer Layer for dynamically detecting the key points of the object Denote N s KPTqueries that are randomly initialized and continuously updated during the training process, used to represent N s key points in the scene;
[0023] Making the queries representing different key points interact with the fused feature f fusion ∈R N×C extracted from the scene through a cross-attention layer, and performing scene-adaptive update on the KPT query:
[0024] f′ kpt =MHCA(f fusion ; f kpt )
[0025] Using a similarity-based heatmap generation strategy, after calculating the similarity between each KPT query and the scene points, generating the 3D position and 3D features of the key points in a heatmap-weighted manner:
[0026] heatmap = Softmax(Similarity(f′ kpt , f fusion ))
[0027]
[0028] where represents the weight map of the similarity calculation of each key point detector in the scene, is the coordinate of the finally detected key point.
[0029] Preferably, inputting the key points into a preset multi-scale pose prediction network, and aggregating local structure information into the key points to obtain key points with multi-scale information, specifically including:
[0030] For each detected 3D key point By extracting the fusion features of the nearest neighbor scene points, and aggregating local structure information into the key points through cross attention:
[0031]
[0032] where knn represents k-nearest neighbor points in Euclidean space, and index represents the indexing operation.
[0033] Preferably, predicting the position of the key points in the object space coordinate system through the key point features with multi-scale information, and splicing the position of the output key points in the scene, the features of the key points in the scene, and the position and features of the key points in the object space coordinate system to form multiple groups of corresponding relationships, and outputting the final 6D pose of the object through a multi-layer perceptron, specifically including:
[0034] Predicting the position of the key points in the object coordinate space through the key point features:
[0035]
[0036] And splicing the position of the output key points in the scene, the features of the key points in the scene, and the position and features of the key points in the object space coordinate system to form N s groups of corresponding relationships, and outputting the final 6D pose of the object through a multi-layer perceptron:
[0037]
[0038] In a second aspect, a category-level object 6D pose estimation system based on dynamic key point detection is provided, characterized in that the system includes:
[0039] A receiving module, configured to receive image data of an object, where the image data includes an RGB image and a point cloud, and the point cloud is formed by randomly sampling pixels in the depth map and projecting them into the scene by combining the internal parameters of the camera;
[0040] A feature extraction and fusion module, configured to extract image features of the RGB image and point cloud features of the point cloud respectively, and splice and fuse the image features and the point cloud features to obtain fused features;
[0041] A key point extraction module, configured to input the fused features into a preset dynamic key point detection network to extract key points of the object;
[0042] A processing and output module, configured to input the key points into a preset multi-scale pose prediction network, aggregate local structure information into the key points to obtain key points with multi-scale information, predict the positions of the key points in the object space coordinate system through the key point features with multi-scale information, and splice the positions of the output key points in the scene, the features of the key points in the scene, and the positions and features of the key points in the object space coordinate system to form multiple groups of corresponding relationships, and output the final 6D pose of the object through a multi-layer perceptron.
[0043] In a third aspect, a computer-readable storage medium storing one or more programs is provided, where the one or more programs include instructions that, when executed by a computing device, cause the computing device to execute any of the methods described above.
[0044] In a fourth aspect, a computing device is provided, including:
[0045] One or more processors, a memory, and one or more programs, where the one or more programs are stored in the memory and configured to be executed by the one or more processors, and the one or more programs include instructions for executing any of the methods described above.
[0046] (III) Advantageous Effects
[0047] The method and system for category-level object 6D pose estimation based on dynamic key point detection according to the present invention can adaptively extract the key points of an object from an observation scene, and can achieve good results even when there are many noise points and serious occlusion situations in the scene. Secondly, this patent designs two modules to separately consider the local features of key points and a pose prediction network based on corresponding relationships. It can better extract the local spatial geometric features around the key points and use the corresponding relationships to regress the object pose. And it is trained in a digital twin simulation system, and finally can greatly improve the accuracy of category-level object 6D pose estimation on existing data sets. Description of the Drawings
[0048] Figure 1 It is a flowchart of the method for category-level object 6D pose estimation based on dynamic key point detection according to the present invention;
[0049] Figure 2 It is an analysis diagram of the method in an embodiment of the present invention. Detailed Embodiments
[0050] Next, in combination with the accompanying drawings of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts belong to the scope of protection of the present invention.
[0051] Embodiment
[0052] As Figure 1-2 shown, an embodiment of the present invention provides a category-level object 6D pose estimation method based on dynamic key point detection, including:
[0053] Receiving image data of an object, where the image data includes an RGB image and a point cloud, and the point cloud is formed by randomly sampling pixels in a depth map and projecting them into a scene in combination with internal camera parameters;
[0054] Respectively extracting the image features of the RGB image and the point cloud features of the point cloud, and splicing and fusing the image features and the point cloud features to obtain the fused features;
[0055] Inputting the fused features into a preset dynamic key point detection network to extract the key points of the object;
[0056] Inputting the key points into a preset multi-scale pose prediction network, aggregating local structure information into the key points to obtain key points with multi-scale information, predicting the positions of the key points in the object space coordinate system through the key point features with multi-scale information, and splicing the positions of the output key points in the scene, the features of the key points in the scene, and the positions and features of the key points in the object space coordinate system to form multiple groups of corresponding relationships, and outputting the final 6D pose of the object through a multi-layer perceptron.
[0057] Further, for the input RGBD image, Resnet18 is used as the feature extractor of the RGB image. For the depth map D, pixels in the depth map are randomly sampled and projected into the scene in combination with the internal camera parameters to form a point cloud. For input data of different modalities, the network first sends the input RGB image into the Resnet18 convolutional neural network to extract the feature map f rgb ∈R h×w×c . For the input point cloud features, the network inputs them into the Pointnet++ point cloud feature extraction network to extract the structural features f point ∈R N×C . Where N is the number of points in the point cloud. Then, the scene point cloud is projected onto the image feature map through the internal camera parameters, and the corresponding features of the scene point cloud on the feature map are extracted through bilinear interpolation f point→rgb ∈R N×CFinally, the features of the two modalities are concatenated and then passed through a multi-layer MLP to output the fused feature f. fusion ∈R N×C .
[0058] Furthermore, for each scene input, after extracting the multi-modal fusion features, we designed a dynamic keypoint detection network as shown in Figure 1 to adaptively extract object keypoints in the scene. To achieve adaptive dynamic detection of object keypoints in different scenes, this project plans to introduce an attention mechanism and Transformer Layer to achieve scene adaptability. denotes N s randomly initialized KPT queries that will be continuously updated during the training process, used to represent N s keypoints in the scene. These queries representing different keypoints are interacted with the fused feature f fusion ∈R N×C extracted from the scene through the cross attention layer, and the KPT queries are updated adaptively to the scene:
[0059] f′ kpt = MHCA(f fusion ; f kpt )
[0060] The updated KPT queries aggregate scene-adaptive features for the subsequent keypoint detection module. Next, using a similarity-based heatmap generation strategy, after calculating the similarity between each KPT query and the scene points, the 3D positions of the keypoints and the 3D features are generated by weighting the heatmap. Specifically:
[0061] heatmap = Softmax(Similarity(f′ kpt , f fusion ))
[0062]
[0063] where represents the weight map for calculating the similarity of each keypoint detector in the scene, is the coordinate of the finally detected keypoint. The dynamically detected keypoints can adapt to different scenes and changes. No matter how the position, angle, lighting conditions, etc. of the target object change, the keypoints can be accurately detected, and this set of keypoints can be generalized to different instance objects of the same category, making the model more generalizable and more suitable for the task of category-level object pose estimation. This enables better precise and robust pose estimation in the future.
[0064] Furthermore, this patent designs a multi-scale pose prediction network, which is mainly divided into two modules: a local feature aggregation module and a pose prediction network based on correspondence:
[0065] Local feature aggregation module. In order to enable each key point to better extract local information in the scene and generate multi-scale features, this project proposes a local feature aggregation module for the position of key points. Specifically, for each detected 3D key point By extracting the fused features of its nearest neighbor scene points, the local structural information is aggregated into the key points through cross attention:
[0066]
[0067] where knn represents the k-nearest neighbor points in the Euclidean space, and index represents the indexing operation. Aggregating local features through local attention enables the key points to have multi-scale information and can better predict the pose of scene objects.
[0068] Pose prediction network based on correspondence. In order to regress the object pose according to the correspondence output by the network, an advanced deep neural network is used to fit the traditional least squares algorithm, making the pose output by fitting the correspondence more robust. This method first predicts the position of the key points in the object coordinate space through the key point features:
[0069]
[0070] And the position of the output key points in the scene, the features of the key points in the scene, and the position and features of the key points in the object space coordinate system are concatenated to form N s groups of correspondences, and the final object pose is output through a multi-layer perceptron. Specifically:
[0071]
[0072] The prediction based on the correspondence relationship is more in line with the relationship of least squares fitting of object coordinate pairs in mathematics, enabling the network to more easily learn the pose mapping relationship. Moreover, in this method, the correspondences generated by adaptive key points can remove the influence of noise points in the scene, only selecting the most representative part of the key points, making the overall calculation accuracy higher, the robustness stronger, and the calculation more efficient.
[0073] This patent also designs a method for using digital twin simulation technology to assist in the training of models and verify the performance of models. To collect 6D pose data of an object, we deploy virtual RGBD sensors into the digital twin environment, and these sensors simulate perception devices such as cameras and lidar in the real world. These virtual sensors will record the position and pose information of the object in real time, generating a large amount of simulated data. We use this simulated data to train and verify the deep learning model. By conducting simulation experiments in the digital twin environment, the performance of the model, including its accuracy, robustness, and generalization ability, is verified.
[0074] In this patent, the Resnet convolutional neural network in the RGBD multi-modal feature extraction backbone network and the pointnet++ point cloud feature extraction network are existing technologies in the background art. Based on this technology, this patent newly adds an adaptive key point detection network, designs a new pose estimation network based on local feature aggregation, and conducts simulation experiments in the digital twin environment to verify the effectiveness of the technology.
[0075] Another embodiment of the present invention provides a category-level object 6D pose estimation system based on dynamic key point detection, characterized in that the system includes:
[0076] A receiving module for receiving image data of an object, the image data including an RGB image and a point cloud, and the point cloud is formed by randomly sampling pixels in the depth map and projecting them into the scene by combining the internal parameters of the camera.
[0077] A feature extraction and fusion module for respectively extracting the image features of the RGB image and the point cloud features of the point cloud, and splicing and fusing the image features and the point cloud features to obtain the fused features.
[0078] A key point extraction module for inputting the fused features into a preset dynamic key point detection network to extract the key points of the object.
[0079] A processing and output module for inputting the key points into a preset multi-scale pose prediction network, aggregating local structure information into the key points to obtain key points with multi-scale information, predicting the position of the key points in the object space coordinate system through the key point features with multi-scale information, and splicing the position of the output key points in the scene, the features of the key points in the scene, and the position and features of the key points in the object space coordinate system to form multiple sets of corresponding relationships, and outputting the final 6D pose of the object through a multi-layer perceptron.
[0080] Embodiments of the present application may be provided as a method or a computer program product. Therefore, the present application may take the form of an all-hardware embodiment, an all-software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code. The solutions in the embodiments of the present application can be implemented using various computer languages, for example, object-oriented programming languages such as Java and interpreted scripting languages such as JavaScript, etc.
[0081] The present application is described with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to embodiments of the present application. It should be understood that each flow and / or block in the flowchart and / or block diagram, and the combination of flows and / or
[0082] or blocks in the flowchart and / or block diagram can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, such that the instructions executed by the processor of the computer or other programmable data processing devices generate means for implementing the functions specified in Figure 1 one flow or multiple flows and / or Figure 1 one block or multiple blocks.
[0083] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, such that the instructions stored in the computer-readable memory generate a manufactured article including instruction means that implement the functions specified in Figure 1 one flow or multiple flows and / or Figure 1 one block or multiple blocks.
[0084] These computer program instructions can also be loaded onto a computer or other programmable data processing device, such that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process, so that the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in Figure 1 one flow or multiple flows and / or Figure 1 one block or multiple blocks.
[0085] It should be noted that in this article, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the terms "include", "comprise" or any other variant thereof are intended to cover non-exclusive inclusion, such that a process, method, article or device comprising a series of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "comprising an..." does not exclude the presence of additional identical elements in the process, method, article or device comprising said element.
Claims
1. A method for category - level object 6D pose estimation based on dynamic key - point detection, characterized in that, it includes: Receiving the image data of the object, where the image data includes RGB images and point clouds. Among them, the point cloud is formed by randomly sampling the pixels in the depth map and projecting them into the scene by combining the internal parameters of the camera; Respectively extracting the image features of the RGB image and the point cloud features of the point cloud, and splicing and fusing the image features and the point cloud features to obtain the fused features; Inputting the fused features into a preset dynamic key - point detection network to extract the key - points of the object, specifically including: Introduce the attention mechanism and Transformer Layer for dynamic detection of object key points, denote N s KPTqueries that are randomly initialized and continuously updated during the training process, used to represent N s key points in the scene; Match the queries representing different key points with the fused feature f extracted from the scene fusion ∈R N×C Perform interaction through the cross-attention layer and adaptively update the KPT queries for the scene: f′ kpt = MHCA(f fusion ; f kpt ) Using a similarity - based heat - map generation strategy, calculating the similarity between each KPT query and the scene points, and generating the 3D position and 3D features of the key - points in a heat - map weighted manner: heatmap = Softmax(Similarity(f′ kpt , f fusion )) Among them represents the weight map for calculating the similarity of each key point detector in the scene, which are the coordinates of the key points for the final detection; Inputting the key - points into a preset multi - scale pose prediction network, aggregating local structure information into the key - points to obtain key - points with multi - scale information, predicting the position of the key - points in the object space coordinate system through the key - point features with multi - scale information, and splicing the position of the output key - points in the scene, the features of the key - points in the scene, and the position and features of the key - points in the object space coordinate system to form multiple groups of corresponding relationships, and outputting the final 6D pose of the object through a multi - layer perceptron; The step of inputting the key - points into a preset multi - scale pose prediction network and aggregating local structure information into the key - points to obtain key - points with multi - scale information specifically includes: For each detected 3D key point By extracting the fused features of the nearest neighbor scene points, and aggregating the local structural information into the key points through cross attention: where knn represents the k - nearest neighbors in Euclidean space, and index represents the indexing operation; The step of predicting the position of the key - points in the object space coordinate system through the key - point features with multi - scale information, and splicing the position of the output key - points in the scene, the features of the key - points in the scene, and the position and features of the key - points in the object space coordinate system to form multiple groups of corresponding relationships, and outputting the final 6D pose of the object through a multi - layer perceptron specifically includes: Predicting the position of the key - points in the object coordinate space through the key - point features: And the positions of the output key points in the scene, the features of the key points in the scene, and the positions and features of the key points in the object space coordinate system are concatenated to form N s sets of corresponding relationships, and the final 6D pose of the object is output through a multi-layer perceptron:
2. The method for category - level object 6D pose estimation based on dynamic key - point detection according to claim 1, characterized in that: The feature extractor of the RGB image adopts the Resnet18 convolutional neural network.
3. The method for category - level object 6D pose estimation based on dynamic key - point detection according to claim 2, characterized in that: The step of respectively extracting the image features of the RGB image and the point cloud features of the point cloud specifically includes: The input RGB image is fed into the Resnet18 convolutional neural network to extract the feature map f of the image rgb ∈R h×w×c , where h, w, and c are the height, width, and number of channels of the feature map, respectively; Input the point cloud into the Pointnet++ point cloud feature extraction network to extract the structural feature f of the point cloud point ∈R N×C ; where N is the number of point clouds, and C is the dimension of each point cloud feature output by the Pointnet++ network.
4. The method for category - level object 6D pose estimation based on dynamic key - point detection according to claim 3, characterized in that: The step of splicing and fusing the image features and the point cloud features to obtain the fused features specifically includes Project the structural features of the point cloud onto the feature map of the image through the internal parameters of the camera, and extract the corresponding feature f of the structural features of the point cloud on the feature map of the image through bilinear interpolation point→rgb ∈R N×C ; The fused feature f is obtained by concatenating the feature map of the image and the structural features of the point cloud and then outputting through a multi-layer MLP fusion ∈R N×C .
5. A category - level object 6D pose estimation system based on dynamic key - point detection, which executes the method for category - level object 6D pose estimation based on dynamic key - point detection according to claim 1, characterized in that, the system includes: A receiving module, configured to receive image data of an object, where the image data includes an RGB image and a point cloud, and the point cloud is formed by randomly sampling pixels in a depth map and projecting them into a scene by combining internal parameters of a camera; A feature extraction and fusion module, configured to extract image features of the RGB image and point cloud features of the point cloud respectively, and splice and fuse the image features and the point cloud features to obtain fused features; A key point extraction module, configured to input the fused features into a preset dynamic key point detection network to extract key points of the object; A processing and output module, configured to input the key points into a preset multi-scale pose prediction network, aggregate local structure information into the key points to obtain key points with multi-scale information, predict the positions of the key points in the object space coordinate system through the key point features with multi-scale information, and splice the positions of the output key points in the scene, the features of the key points in the scene, and the positions and features of the key points in the object space coordinate system to form multiple groups of corresponding relationships, and output the final 6D pose of the object through a multi-layer perceptron.
6. A computer-readable storage medium storing one or more programs, wherein, the one or more programs include instructions that, when executed by a computing device, cause the computing device to execute any of the methods according to claims 1-4.
7. A computing device, wherein, comprising: one or more processors, a memory, and one or more programs, wherein the one or more programs are stored in the memory and configured to be executed by the one or more processors, and the one or more programs include instructions for executing any of the methods according to claims 1-4.