A 6dof object pose estimation method based on a single RGB image

Through hierarchical network architecture and feature extraction methods, three-dimensional point clouds are reconstructed and poses are estimated from a single RGB image, which solves the reconstruction problems in existing technologies and achieves efficient and accurate estimation of the three-dimensional shape and pose of objects.

CN116958262BActive Publication Date: 2025-09-16TONGJI UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310976771.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-08-04
Publication Date
2025-09-16
Estimated Expiration
2043-08-04

AI Technical Summary

Technical Problem

Existing technologies find it difficult to efficiently reconstruct an object's three-dimensional point cloud and estimate its pose from a single RGB image, especially in the absence of depth information and 3D models. Existing methods also require large amounts of data or complex input conditions.

Method used

A hierarchical network architecture is adopted to recover the object's contour map, depth map and surface normal map from the RGB image through the feature extraction network. The structure-aware network and heterogeneous network are combined to perform 3D point cloud reconstruction and pose estimation. The ResNet and PointNet models are used for feature extraction and fusion to gradually restore the object's 3D shape and pose.

Benefits of technology

It achieves high-quality 3D point cloud reconstruction and object pose estimation from a single RGB image, reduces dependence on background and texture information, improves the interpretability and accuracy of reconstruction, and is suitable for a variety of scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116958262B_ABST
    Figure CN116958262B_ABST
Patent Text Reader

Abstract

The present invention provides a 6dof object pose estimation method based on a single RGB image, which belongs to the field of computer vision and computer graphics technology, including feature extraction, three-dimensional point cloud reconstruction and pose estimation of RGB images. Feature extraction is achieved by building a feature extraction network architecture. Three-dimensional point cloud reconstruction is to first obtain the intermediate information of the object based on the various low-level (geometry, reflectivity) and high-level (connectivity, symmetry) characteristics of the object, and then further generate a 3D point cloud model of the object. Pose estimation uses a heterogeneous network to process RGB data and point cloud data respectively, integrates the features of the two data through a fusion network, thereby predicting the pose information of the object. The 6dof object pose detection method proposed in the present invention focuses on the problems such as small data volume, difficult acquisition of RGBD data format, and no object 3D model, and can ensure the accuracy and generalizability of target object pose detection, and can be effectively applied to real scenes.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of computer vision and computer graphics, and in particular relates to a 6Dof object pose estimation method based on a single RGB image. Background Art

[0002] With the development of artificial intelligence and machine vision technology, the research on object pose estimation has received widespread attention at home and abroad. It can be applied to many aspects such as robot grasping, autonomous driving, augmented reality, digital twins, etc., aiming to estimate the rotation and translation of the object relative to the specified standard frame.

[0003] Common methods for object pose estimation include template matching-based methods, point-based methods, and descriptor-based methods. For example, template matching-based methods render synthetic image patches from different viewpoints distributed on a sphere around the object's 3D model and store them as a template database. This template database is then used to sequentially search the input image using a sliding window. A representative template matching-based method is LineMOD, which has proposed an efficient and robust pose detection strategy for color, depth, and RGB-D images and provided the first dataset with labeled poses. This dataset is still used as a benchmark for object detection and pose estimation. An alternative to template matching methods is the random forest learning method. Most of these methods rely on RGB-D information or 3D models of the object. However, typical mobile phones and computer cameras cannot provide images with depth information, and textured 3D models of objects are even more difficult to obtain. Furthermore, methods that rely solely on RGB image input require excessively large amounts of data, require multiple viewing angles, and are difficult to obtain. Universality and simple input have always been the goals of object pose estimation research. Universality means that the pose estimator can be applied to any object without requiring specific training on the object or its category. When estimating object pose, if only a single RGB image is used as the input of the estimator without additional object masks, depth maps or 3D models, the requirement of simple input can be fully met.

[0004] However, reconstructing an object's 3D point cloud model and pose estimation from a single RGB image has been a long-standing and widely-regarded research problem in the fields of computer vision and computer graphics. Summary of the Invention

[0005] The object of the present invention is to provide a 6dof object pose estimation method based on a single RGB image, characterized in that it comprises the following steps:

[0006] S1: Build a feature extraction network architecture to extract features from RGB images and restore the RGB images to transition images without background color, texture, and lighting information. The transition images include the object's contour map, depth map, and surface normal map.

[0007] S2: The complete object is assumed to be composed of multiple geometric primitives. The structure-aware network module learns to predict the geometric shape of each component and the arrangement relationship between the components to obtain a predicted hierarchical structure graph. The hierarchical structure graph is combined with the object's contour map, depth map, and surface normal map in S1 to form a four-channel image. The 3D shape estimator is trained using the four-channel image to complete the reconstruction of the 3D point cloud.

[0008] S3: Use a heterogeneous network model to extract features from the RGB data and the point cloud data obtained in S2 respectively. Two types of features from the extracted features are used as the input of the fusion network model to output the 3D bounding box of the target object and complete the 6dof pose detection of the object.

[0009] Furthermore, in S1, the feature extraction network is specifically: the first encoder of the feature extractor based on the ResNet-18 network model; the feature extraction is specifically: inputting the RGB image into the first encoder, completing feature downsampling through convolution operation and residual block input, and outputting a feature map.

[0010] Furthermore, in S1, the transition image is realized by decoding the feature map through the first decoder, which is specifically performed as follows: the feature map is input into the first decoder, and the first decoder converts the feature map into a contour map, a depth map and a surface normal map through four groups of 5×5 transposed convolution operations and a ReLU layer.

[0011] Furthermore, in S2, obtaining the predicted hierarchical structure diagram specifically includes the following steps:

[0012] S21: The segmentation network recursively splits the shape representation into representations of its parts;

[0013] S22: Structure-aware networks focus on learning a hierarchical arrangement of primitives, i.e. assigning parts of an object to primitives at each depth level.

[0014] S23: Recover the geometric network of primitive parameters and obtain the predicted hierarchical structure graph.

[0015] Furthermore, in S2, the encoder of the 3D shape estimator is a network model implemented based on ResNet-34. The encoder of the 3D shape estimator achieves better feature extraction effect than the first encoder by deepening the number of network layers. The reconstruction of the 3D point cloud by training the 3D shape estimator specifically includes the following steps:

[0016] S24: Input the four-channel image of size [256, 256, 4] into the encoder of the 3D shape estimator, perform a convolution operation, and then output a feature map of size [128, 128, 64];

[0017] S25: Input the residual block, perform average pooling on the feature map and use the fully connected layer to map the feature map to the 512-dimensional feature space to obtain the feature vector Zs, thereby completing the global feature extraction;

[0018] S26: Use the feature vector Zp of the ground-truth point cloud corresponding to the image object extracted by the encoder of the point cloud generator as prior knowledge, and train the encoder of the 3D shape estimator by calculating the difference between ZsH and Zp. The training process is completed on the ShapeNetCore dataset;

[0019] S27: After the encoder training of the 3D shape estimator is completed, the feature vector Zs* of the target object is decoded into a three-dimensional cloud point with a resolution of 2048 points through the decoder of the point cloud generator, completing the reconstruction from a single image to a three-dimensional point cloud.

[0020] Furthermore, the loss functions of the encoder and decoder for training the 3D shape estimator are expressed as:

[0021]

[0022] In the loss function, X gt ∈N×3, is the ground truth of the point cloud; X pred ∈N×3, is the reconstructed point cloud.

[0023] Furthermore, in S3, the ResNet network model is used to extract visual features from the RGB image, and the PointNet network model is used to extract features from the point cloud data generated in S2. The two types of features are global features and single-point features.

[0024] Furthermore, by deleting all batch normalization layers of the PointNet network structure, the prediction accuracy of the bounding box is improved.

[0025] Furthermore, S3 is specifically as follows: the fusion network model is a dense fusion network model, which takes the input three-dimensional points as dense spatial positioning points and predicts the spatial offset from the three-dimensional point to the nearby bounding box corner position for each input three-dimensional point. The density fusion network model processes the joint input through multiple layers to predict a 3D bounding box and the score of each three-dimensional point. The prediction with the highest score is the final prediction.

[0026] Furthermore, the loss function of the dense fusion network model is expressed as:

[0027]

[0028] In the loss function of the dense fusion network model, N is the number of input point cloud points; is the offset between the true 3D box corner point and the i-th input point; Represents the predicted offset; L score is the score loss function; L stn is the introduced spatial transformation regularization loss.

[0029] Compared with the prior art, the beneficial effects of the present invention are mainly reflected in:

[0030] 1. This paper proposes a hierarchical network architecture for 3D point cloud reconstruction. It does not directly complete the 3D point cloud reconstruction task through a single RGB image. Instead, it relies on intermediate information extracted from the image, including the object's contour map, depth map, and surface normal map, to gradually restore the object's 3D shape. It removes background, color, and texture information that are not needed for point cloud reconstruction, reduces the burden of domain transfer, and improves the quality of point cloud generation.

[0031] 2. The present invention proposes a structure-aware representation method that takes into account high-level information of objects, including connectivity and symmetry. Based on the decomposition of components and the relationship between components, the object is modeled in the form of primitives. That is, geometrically complex objects are modeled with more primitives, while simple objects are modeled with fewer primitives, making the 3D reconstruction interpretable.

[0032] 3. A heterogeneous network is proposed to extract features from RGB data and point cloud data respectively, and then the two types of features are fused and further abstracted. Finally, the 3D point cloud is regarded as a spatial positioning point and dense prediction is performed to obtain the 3D contour box of the object, completing the 6dof pose detection of the object. BRIEF DESCRIPTION OF THE DRAWINGS

[0033] Figure 1 Schematic diagram of the input graph, intermediate result graph and final result graph of an embodiment of the present invention.

[0034] Figure 2 The figure is a schematic diagram of the workflow of the 6dof object pose estimation method based on a single RGB image of the present invention.

[0035] Figure 3 This is a diagram of the network model structure of the present invention. DETAILED DESCRIPTION

[0036] The following is a more detailed description of a 6dof object pose estimation method based on a single RGB image of the present invention in conjunction with a schematic diagram, which shows a preferred embodiment of the present invention. It should be understood that those skilled in the art can modify the present invention described herein while still achieving the beneficial effects of the present invention. Therefore, the following description should be understood as being widely known to those skilled in the art and not as a limitation of the present invention.

[0037] like Figure 1-3 As shown in FIG, a 6dof object pose estimation method based on a single RGB image includes the following steps:

[0038] S1: The encoder of the feature extractor is based on a ResNet-18 network model. It first resizes the RGB image and feeds it into encoder E1, where it performs a convolution operation and outputs a feature map of size [128, 128, 64]. A series of residual blocks then operate on the input feature map, gradually halving the input feature size and downsampling the features to reduce computational complexity. The number of channels is also gradually doubled, ultimately outputting a feature map of size [8, 8, 512]. Decoder D1 converts the output [8, 8, 512] feature map through four sets of 5×5 transposed convolutions and ReLU layers, converting it into 256×256 contour, depth, and surface normal maps. After generating the contour, depth, and surface normal maps, the contour map is used to mask the depth and surface normal maps to determine the precise location of the object to be reconstructed, resulting in a higher-quality 3D reconstructed point cloud.

[0039] S2: Next, a hierarchical structure prediction network is constructed. This network consists of three main components: (i) a segmentation network that recursively segments the shape representation into its component representations; (ii) a structure network that focuses on learning the hierarchical arrangement of primitives, assigning object parts to primitives at each depth level; and (iii) a geometry network that recovers the primitive parameters. Ultimately, a predicted hierarchical structure map is obtained. The processed surface normal map, depth map, and hierarchical structure map are combined into a four-channel image, and a 3D shape estimator is trained to complete the reconstruction of the 3D point cloud. The encoder E2 of the 3D shape estimator is a network model based on ResNet-34. Similar to the encoder E1 in step S1, the network is deeper to achieve better feature extraction. First, a convolution operation is performed on the input image of size [256, 256, 4], outputting a feature map of size [128, 128, 64]. Then, after a series of residual blocks similar to those in step S1, the feature map is average pooled and mapped to a 512-dimensional feature space using a fully connected layer to obtain a feature vector, completing global feature extraction.

[0040] After obtaining the feature vector Zs, the feature vector Zp of the true point cloud corresponding to the image object extracted by the encoder E3 in the point cloud generator is used as prior knowledge, and the difference between Zs and Zp is calculated to train the encoder E1 of the 3D shape estimator. This process is trained on the ShapeNetCore dataset. When the encoder E1 is trained, the decoder D3 in the point cloud generator is used to decode the feature vector Zs* of the target object into a three-dimensional point cloud with a resolution of 2048 points. At this point, the network model uses three sets of encoders and decoders, with the contour map, depth map, surface normal map, and hierarchical structure map as intermediaries, to complete the reconstruction from a single image to a three-dimensional point cloud by learning strong prior knowledge. Specifically, the loss function for training the point cloud encoder and decoder is:

[0041]

[0042] In the formula, X gt ∈N×3 is the ground truth of the point cloud, X pred ∈N×3 is the reconstructed point cloud.

[0043] S3: Visual feature extraction is performed on the RGB image using the ResNet network model, and feature extraction is performed on the point cloud data generated in step S2 using the PointNet network model, including global and single-point features. The PointNet network architecture was modified to remove all batch normalization layers to improve bounding box prediction accuracy. The fusion network uses a dense fusion network, taking as input the image features extracted by the CNN and the corresponding point cloud features generated by the PointNet subnetwork. Its task is to combine these features and output a 3D bounding box for the target object. The key idea of ​​this dense fusion network model is to treat the input 3D points as dense spatial anchor points. Instead of directly regressing the absolute positions of the 3D bounding box corners, it predicts the spatial offset of each input 3D point from the point to the positions of nearby 3D bounding box corners. The dense fusion network uses several layers to process the combined input, predicting a 3D bounding box and a score for each point. During testing, the prediction with the highest score is selected as the final prediction. Specifically, the loss function of the dense fusion network is:

[0044]

[0045] In the formula, N is the number of input point cloud points, is the offset between the true 3D box corner point and the i-th input point, Represents the predicted offset, L score is the score loss function, L stn is the introduced spatial transformation regularization loss.

[0046] The above description is merely a preferred embodiment of the present invention and does not limit the present invention in any way. Any person skilled in the art who, without departing from the scope of the present invention, makes any equivalent substitution, modification, or other changes to the technical solution and technical content disclosed in the present invention shall be deemed to be within the scope of the present invention and still fall within the scope of protection of the present invention.

Claims

1. A 6dof object pose estimation method based on a single RGB image, characterized in that: The steps include: S1: Building a feature extraction network architecture to extract features from the RGB image, and restoring the RGB image to a transition image without background color, texture, and lighting information. The transition image includes a contour map, a depth map, and a surface normal map of the object; S2: The complete object is assumed to be composed of multiple geometric primitives. The structure-aware network module learns to predict the geometric shape of each component and the arrangement relationship between the components to obtain a predicted hierarchical structure graph. The hierarchical structure graph is combined with the contour map, depth map, and surface normal map of the object in S1 into a four-channel image. The 3D shape estimator is trained using the four-channel image to complete the reconstruction of the 3D point cloud. S3: Use a heterogeneous network model to extract features from the RGB data and the point cloud data obtained in S2 respectively. Two types of features from the extracted features are used as the input of the fusion network model to output the 3D bounding box of the target object and complete the 6dof pose detection of the object.

2. The 6dof object pose estimation method based on a single RGB image according to claim 1, characterized in that In S1, the feature extraction network is specifically: a first encoder of a feature extractor based on a ResNet-18 network model; feature extraction is specifically: inputting an RGB image into the first encoder, completing feature downsampling through a convolution operation and a residual block input, and outputting a feature map.

3. The 6dof object pose estimation method based on a single RGB image according to claim 1, characterized in that In S1, the transition image is realized by decoding the feature map through the first decoder, which is specifically performed as follows: the feature map is input into the first decoder, and the first decoder converts the feature map into the contour map, depth map and surface normal map through four groups of 5×5 transposed convolution operations and ReLU layers.

4. The 6dof object pose estimation method based on a single RGB image according to claim 1, characterized in that In S2, obtaining the predicted hierarchical structure diagram specifically includes the following steps: S21: The segmentation network recursively splits the shape representation into representations of its parts; S22: Structure-aware networks focus on learning a hierarchical arrangement of primitives, i.e. assigning parts of an object to primitives at each depth level. S23: Recover the geometric network of primitive parameters and obtain the predicted hierarchical structure graph.

5. The 6dof object pose estimation method based on a single RGB image according to claim 1, characterized in that In S2, the encoder of the 3D shape estimator is a network model implemented based on ResNet-34. The encoder of the 3D shape estimator achieves better feature extraction effect than the first encoder by deepening the number of network layers. The reconstruction of the three-dimensional point cloud by training the 3D shape estimator specifically includes the following steps: S24: Input the four-channel image of size [256, 256, 4] into the encoder of the 3D shape estimator, perform a convolution operation, and then output a feature map of size [128, 128, 64]; S25: Input the residual block, perform average pooling on the feature map and use the fully connected layer to map the feature map to the 512-dimensional feature space to obtain the feature vector Zs, thereby completing the global feature extraction; S26: Use the feature vector Zp of the ground-truth point cloud corresponding to the image object extracted by the encoder of the point cloud generator as prior knowledge, and train the encoder of the 3D shape estimator by calculating the difference between ZsH and Zp. The training process is completed on the ShapeNetCore dataset; S27: After the encoder training of the 3D shape estimator is completed, the feature vector Zs* of the target object is decoded into a three-dimensional cloud point with a resolution of 2048 points through the decoder of the point cloud generator, completing the reconstruction from a single image to a three-dimensional point cloud.

6. The 6dof object pose estimation method based on a single RGB image according to claim 5, characterized in that The loss function for training the encoder and decoder of the 3D shape estimator is expressed as: In the loss function, X gt ∈N×3, is the ground truth of the point cloud; X pred ∈N×3, is the reconstructed point cloud.

7. The 6dof object pose estimation method based on a single RGB image according to claim 1, characterized in that In S3, visual features are extracted from the RGB image using the ResNet network model, and features are extracted from the point cloud data generated in S2 using the PointNet network model. The two types of features are global features and single-point features.

8. The 6dof object pose estimation method based on a single RGB image according to claim 7, characterized in that By deleting all batch normalization layers of the PointNet network structure, the prediction accuracy of the bounding box is improved.

9. The 6dof object pose estimation method based on a single RGB image according to claim 7, characterized in that Specifically, S3 is as follows: the fusion network model is a dense fusion network model, which uses the input three-dimensional points as dense spatial positioning points and predicts the spatial offset from the three-dimensional point to the corner position of the nearby bounding box for each input three-dimensional point. The dense fusion network model processes the joint input through multiple layers to predict a 3D bounding box and the score of each three-dimensional point. The prediction with the highest score is the final prediction.

10. The 6dof object pose estimation method based on a single RGB image according to claim 9, characterized in that The loss function of the dense fusion network model is expressed as: In the loss function of the dense fusion network model, N is the number of input point cloud points; is the offset between the true 3D box corner point and the i-th input point; Represents the predicted offset; L score is the score loss function; L stn is the introduced spatial transformation regularization loss.

Citation Information

Patent Citations

  • Target pose estimation method fusing RGB-D visual features

    CN112270249A

  • Method for tracking head mounted display device and head mounted display system

    US20230098910A1