A real-time monocular 6D pose estimation method and system applicable to symmetrical objects

CN117372521BActive Publication Date: 2026-09-18ZHEJIANG UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202311341800.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-10-17
Publication Date
2026-09-18
Estimated Expiration
2043-10-17

AI Technical Summary

Technical Problem

所以这些方法在学习针对对称物体的特征图时会出现同样的输入对应不同的特征图,导致网络收敛性出现问题,该问题在遮挡情况下更加严重,导致网络输出的特征图完全失效

Benefits of technology

[0041] 1. 6D pose estimation from a single RGB image is an important and challenging topic in computer vision. Recent work on 6D pose estimation of objects based on deep neural networks has significantly improved accuracy under conditions of object occlusion, lighting variations, and cluttered backgrounds. However, a good solution has not yet been found for the convergence problem of network feature map learning caused by symmetry. To address this issue, this invention improves upon the binary encoding of the ZebraPose algorithm by annotating the symmetry information of the object's CAD model, obtaining the binary encoding of the object's CAD model, and then generating a binary encoded feature map with symmetry information. The feature map of this invention inherits the advantage of the discreteness of the binary encoding of the ZebraPose algorithm. Compared with other feature maps, the feature space that the network needs to learn is smaller. Other methods use continuous encoding, requiring the network to fit using regression. The feature map of this invention is discrete encoding, allowing the network to fit using classification, which is beneficial for the learning of deep neural networks. At the same time, the binary encoding with symmetry information in this invention contains symmetry information, avoiding the convergence problem of the network.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117372521B_ABST
    Figure CN117372521B_ABST
Patent Text Reader

Abstract

This invention discloses a real-time monocular 6D pose estimation method and system applicable to symmetrical objects. The invention performs symmetry annotation on the CAD model of the target object, obtaining a binary code of the CAD model containing symmetry information. The binary code is then rendered under the ground truth pose of the target object, yielding a binary code feature map corresponding to the RGB image. A dilated spatial pooling pyramid coding network is used to process the target object image, outputting the complete mask, visible portion mask, and binary code feature map of the target object. Finally, a pose estimation network is used to process the complete mask, visible portion mask, and binary code feature map of the target object, outputting the 6D pose estimation result. This invention overcomes the difficulty of traditional pose estimation algorithms in handling 6D pose estimation of symmetrical objects. Since it does not require the PnP algorithm, the pose estimation result is directly output by the network, resulting in fast operation, real-time performance, and excellent overall performance.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of computer vision and relates to a real-time monocular 6D pose estimation method and system applicable to symmetrical objects. Background Technology

[0002] The statements in this section are merely background information related to the present invention and do not necessarily constitute prior art.

[0003] Monocular 6D pose estimation of objects refers to obtaining the position and pose of an object in the world coordinate system or camera coordinate system using a single RGB image as input. The object's position is represented by a 3D translation vector t, and its pose is represented by a 3D rotation matrix R. 6D pose estimation technology is a hot research topic in computer vision and has applications in areas such as robot grasping, augmented reality, and autonomous driving.

[0004] Traditional 6D pose estimation techniques for objects mainly rely on matching manually extracted image feature points with features on the object model, and then solving the 6D pose of the object using a multi-point perspective (PnP) algorithm. However, these methods are limited by the representational power of manually extracted image feature points and the matching methods between image features and model features, making them difficult to use in situations with cluttered backgrounds, weak textures, target occlusion, or symmetrical targets.

[0005] In recent years, with the development of deep learning technology, two types of deep learning algorithms have emerged, leading to a series of algorithms. The first type, called the direct method, utilizes the powerful expressive capabilities of deep learning networks to construct an end-to-end network. It does not use multi-point perspective algorithms, taking a single RGB image as input and directly outputting the 6D pose estimation result of the object. This type of method has a simple network construction and high 6D pose estimation accuracy. For symmetrical objects, the true pose needs to be normalized to a small range, which can lead to discontinuous network output. The other type of algorithm is called the feature map method. This type of method divides 6D pose estimation into two stages. The first stage allows the neural network to learn artificially constructed feature maps, which already contain the correspondence between the feature maps and the object model points. Then, the 6D pose of the object is solved using multi-point perspective algorithms, such as the GDR-Net algorithm, or by using a deep neural network to output the 6D pose of the object, such as the ZebraPose algorithm. This type of algorithm is interpretable because it has the feature maps output in the first stage, exhibiting stronger overall performance than the first type of method. Because many objects in reality possess symmetry, the feature maps commonly used in current feature map methods, such as UV maps, feature points, normalized object coordinates, and binary encoding, do not consider the symmetry of objects. Therefore, when these methods learn feature maps for symmetrical objects, the same input may correspond to different feature maps, leading to convergence problems in the network. This problem is exacerbated under occlusion conditions, causing the network's output feature maps to become completely invalid. Summary of the Invention

[0006] Existing feature map methods have demonstrated excellent overall performance, but there is still room for improvement in their performance for pose estimation of symmetrical objects under occlusion. To address this issue, this invention proposes a real-time monocular 6D pose estimation method and system suitable for symmetrical objects.

[0007] The objective of this invention is achieved through the following technical solution:

[0008] In a first aspect, the present invention provides a real-time monocular 6D pose estimation method applicable to symmetrical objects, comprising:

[0009] The original object image is acquired and preprocessed to obtain the target object image;

[0010] Symmetry annotation is performed on the CAD model of the target object. The symmetry form of each point on the CAD model is determined by spatial region selection and threshold selection. The points on the CAD model are recursively clustered by a clustering algorithm to generate a binary code of the CAD model containing symmetry information. The mapping relationship between the binary code of the CAD model and the coordinates of the points on the CAD model is constructed. The binary code of the CAD model is rendered under the ground truth pose of the target object to obtain the binary code feature map corresponding to the RGB image as the label.

[0011] The target object image is processed using a trained Atrous Spatial Pyramid Pooling (ASPP) network to output the complete mask, visible part mask, and binary encoded feature map of the target object.

[0012] The trained pose estimation network is used to process the complete mask, visible part mask, and binary encoded feature map of the target object, and outputs 6D pose estimation results.

[0013] Further, the step of acquiring the original object image and performing preprocessing to obtain the target object image specifically involves:

[0014] Obtain the original object image; use an object detection algorithm to detect objects in the original image; expand the detected rectangular region into a square region, and expand the boundary by 10% to 50%; crop the original object image according to the coordinates of the expanded region to obtain the target object image.

[0015] Furthermore, the symmetrical annotation of the target object's CAD model specifically involves:

[0016] The target object's CAD model may have a set of points with various symmetry forms, such as continuous symmetry, discrete symmetry, discrete symmetry around an axis, and asymmetry.

[0017] Based on the shape of the target object's CAD model, select all possible symmetry forms contained in the target object's CAD model;

[0018] The symmetry form corresponding to each point on the CAD model is determined by manually selecting spatial regions and manually selecting thresholds.

[0019] The manual spatial region selection specifically involves: for each possible symmetry form, selecting the region covered by that symmetry form; the selection process is achieved by taking the intersection and union of the spatial regions multiple times.

[0020] The manual threshold selection specifically involves: calculating all points p in this symmetric form for the points on the model. i , looking for p i The nearest point q i Calculate index l = ∑ i (p i -q i ) 2 Each point is divided into different symmetry forms based on a manual threshold and l, and then added to the corresponding symmetry form point set.

[0021] For points in discrete symmetry and about-axis discrete symmetry forms, the corresponding point p of each point in that symmetry form is... i Add it to the model to ensure that all points p in this symmetric form i , looking for p i The nearest point q i =p i This leads to the calculation index l = 0;

[0022] Traverse the points in each symmetric form's point set, and take all points with index l = 0 as a subset; each point in the asymmetric form's point set forms its own subset; for each subset of all symmetric forms, take one point and divide it into two classes using a clustering algorithm to obtain the first binary code {0,1}. For regions where the first binary code is 0, divide them into two classes again using a clustering algorithm to obtain the second binary code {0,1}. Similarly, for regions where the first binary code is 1, divide them into two classes using a clustering algorithm to obtain the second binary code {0,1}. Recursively traverse this process to obtain a total of sixteen binary codes.

[0023] Finally, all sub-point sets are traversed, and the same binary code is assigned to the points in all sub-point sets to complete the encoding of the CAD model points. Thus, a one-to-one mapping between the coordinates of the CAD model points and the binary code is formed, and the mutual mapping relationship between the binary code of the CAD model and the coordinates of the CAD model points is constructed, which can realize the conversion between three-dimensional coordinates and binary codes.

[0024] Furthermore, the rendering of the binary encoding of the CAD model in the true pose of the target object specifically involves:

[0025] The coordinates of points in the CAD model are projected onto the image using rasterized rendering with the true pose.

[0026] Find the coordinates of the CAD model point closest to the coordinates on the image. Based on the mapping relationship between the binary code of the CAD model and the coordinates of the CAD model point, convert the found coordinates of the CAD model point into the corresponding binary code to obtain the binary code feature map.

[0027] Furthermore, training and testing sets are constructed by combining binary encoded feature maps, and an optimizer is used to train the dilated spatial pooling pyramid coding network and the pose estimation network.

[0028] Furthermore, data augmentation is performed during the construction of the training and test sets. Specifically, the obtained target object image is randomly translated and scaled, brightness and contrast are augmented, the target object is used as an occluder, Gaussian noise and block occlusion regions are randomly added, the corresponding camera intrinsic parameter matrix is ​​calculated, and the complete mask, visible part mask and binary encoded feature map of the target object are cropped. The cropped result should match the target object image.

[0029] Furthermore, the process of training the dilated spatial pooling pyramid coding network and the pose estimation network using an optimizer specifically involves:

[0030] Construct loss functions for the complete mask, visible part mask, and binary encoded feature map of the target object. Combine the pose loss function and the position loss function, and sum all the above loss functions proportionally to obtain the final loss function.

[0031] Secondly, the present invention provides a real-time monocular 6D pose estimation system suitable for symmetrical objects, comprising:

[0032] The object detection module is configured to acquire a preprocessed target object image based on the original object image;

[0033] The object pose estimation module is configured to cascade a pre-trained dilated spatial pooling pyramid coding network and a pose estimation network to output a 6D pose estimation result; the 6D pose estimation process is as follows:

[0034] Symmetry annotation is performed on the CAD model of the target object. The symmetry form of each point on the CAD model is determined by spatial region selection and threshold selection. The points on the CAD model are recursively clustered by a clustering algorithm to generate a binary code of the CAD model containing symmetry information. The mapping relationship between the binary code of the CAD model and the coordinates of the points on the CAD model is constructed. The binary code of the CAD model is rendered under the ground truth pose of the target object to obtain the binary code feature map corresponding to the RGB image as the label.

[0035] The target object image is processed using a trained dilated spatial pooling pyramid coding network to output the complete mask, visible part mask, and binary encoded feature map of the target object.

[0036] The trained pose estimation network is used to process the complete mask, visible part mask, and binary encoded feature map of the target object, outputting a 6D pose estimation result. The output 6D pose estimation result is a scale-independent translation vector SITE and a 6D matrix R. 6d We then convert these two back to the representations of the translation vector t and the rotation matrix R, respectively.

[0037] Furthermore, the object detection module also has a depth image preprocessing function, which can preprocess the depth map corresponding to the target object image; and adds a point cloud iterative optimization module to further optimize the output translation vector t and rotation matrix R.

[0038] Thirdly, the present invention provides a real-time monocular 6D pose estimation device suitable for symmetrical objects, including a memory and one or more processors. The memory stores executable code, and when the processor executes the executable code, it implements the real-time monocular 6D pose estimation method for symmetrical objects as described in the first aspect.

[0039] Fourthly, the present invention provides a computer-readable storage medium having a program stored thereon, which, when executed by a processor, implements the real-time monocular 6D pose estimation method for symmetrical objects as described in the first aspect.

[0040] Compared with the prior art, the beneficial effects of the present invention are as follows:

[0041] 1. 6D pose estimation from a single RGB image is an important and challenging topic in computer vision. Recent work on 6D pose estimation of objects based on deep neural networks has significantly improved accuracy under conditions of object occlusion, lighting variations, and cluttered backgrounds. However, a good solution has not yet been found for the convergence problem of network feature map learning caused by symmetry. To address this issue, this invention improves upon the binary encoding of the ZebraPose algorithm by annotating the symmetry information of the object's CAD model, obtaining the binary encoding of the object's CAD model, and then generating a binary encoded feature map with symmetry information. The feature map of this invention inherits the advantage of the discreteness of the binary encoding of the ZebraPose algorithm. Compared with other feature maps, the feature space that the network needs to learn is smaller. Other methods use continuous encoding, requiring the network to fit using regression. The feature map of this invention is discrete encoding, allowing the network to fit using classification, which is beneficial for the learning of deep neural networks. At the same time, the binary encoding with symmetry information in this invention contains symmetry information, avoiding the convergence problem of the network.

[0042] 2. This invention differs from other feature map algorithms or direct network output methods, which utilize multi-point perspective algorithms, but these algorithms cannot handle the correspondence between 2D points and multiple 3D points. This invention leverages the fitting characteristics of deep neural networks to propose a region-correspondence pose estimation network to handle the correspondence between 2D points and multiple 3D points, achieving the same or even better accuracy as existing methods while significantly reducing inference time. Attached Figure Description

[0043] Figure 1A flowchart of a real-time monocular 6D pose estimation method for symmetrical objects provided as an exemplary embodiment;

[0044] Figure 2 A framework diagram of a symmetric object pose estimation network (SymNet) provided for an exemplary embodiment;

[0045] Figure 3 The processing procedure of the input image during training and testing, provided as an exemplary embodiment;

[0046] Figure 4 A schematic diagram of symmetry annotation and binary encoding of a CAD model provided as an exemplary embodiment;

[0047] Figure 5 A structural diagram of a real-time monocular 6D pose estimation device for symmetrical objects, provided as an exemplary embodiment. Detailed Implementation

[0048] The present invention will be further described below with reference to the accompanying drawings and embodiments.

[0049] It should be noted that the following detailed descriptions are exemplary and intended to provide further illustration of the invention. Unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention pertains.

[0050] It should be noted that the terminology used herein is for the purpose of describing particular embodiments only and is not intended to limit the exemplary embodiments of the invention. As used herein, the singular form is intended to include the plural form as well, unless the context clearly indicates otherwise. Furthermore, it should be understood that when the terms “comprising” and / or “including” are used in this specification, they are known to include features, steps, operations, devices, components and / or combinations thereof.

[0051] Where there is no conflict, the embodiments and features in the embodiments of the present invention can be combined with each other.

[0052] This invention validates the method of this embodiment on all objects in the TLESS dataset. The TLESS dataset is a publicly available dataset containing 30 industrial parts with different combinations of symmetry forms, including continuous symmetry, discrete symmetry, discrete symmetry around an axis, and asymmetry. For example, object 1 contains continuous symmetry, discrete symmetry around an axis, and asymmetry. Experimental results show that the method of this embodiment performs comparably to or even better than state-of-the-art methods on the TLESS dataset.

[0053] like Figure 1 As shown, this embodiment provides a real-time monocular 6D pose estimation method suitable for symmetrical objects, which includes the following steps:

[0054] Step S1: Acquire the original object image and perform preprocessing to obtain the target object image; specifically:

[0055] Step S1.1: Use object detection algorithms, such as YOLO series algorithms, Mask R-CNN series algorithms, etc., to detect objects in the original image and obtain the rectangular region surrounding the object;

[0056] Step S1.2: Expand the short side of the rectangle so that the length of the short side is equal to that of the long side, and expand the rectangular area surrounding the object into a square area. If the square area surrounding the object exceeds the boundary of the image, fill it with black.

[0057] Step S1.3: Expand the square region according to the filling and expansion strategy used in training, such as using a random expansion ratio of 10% to 50%;

[0058] Step S1.4: Extract a square area and scale it to a predetermined size, such as 256×256 pixels. This image is the preprocessed target object image.

[0059] Step S2: Process the target object image using a dilated spatial pooling pyramid coding network to output the complete mask, visible part mask, and binary encoded feature map of the target object.

[0060] Step S3: Using a region-corresponding pose estimation network, input the complete mask of the target object, the visible part mask, and the binary encoded feature map, and output the 6D pose estimation result.

[0061] Specifically, in step S2, the process of obtaining the binary encoded feature maps required for training the dilated spatial pooling pyramid coding network is as follows:

[0062] Step A1: Perform symmetry annotation on the CAD model of the target object, including:

[0063] Step A1.1: Based on the shape of the target object's CAD model, select all possible symmetry forms contained in the target object's CAD model;

[0064] The target object's CAD model may have a set of points with various symmetry forms, such as continuous symmetry, discrete symmetry, discrete symmetry around an axis, and asymmetry.

[0065] A continuous symmetric point set contains a series of subsets of points. Points in each subset can be generated by rotating any point in the subset around the axis of symmetry at any angle, such as the cylindrical region on the top of a cola bottle can.

[0066] Discrete symmetric point sets contain a series of subsets of points. Points in each subset can be obtained by transforming any point using a symmetric form described by a series of predefined transformation matrices. For example, a cuboid contains only discrete symmetric points.

[0067] The discrete symmetric point set about an axis contains a series of subsets of points, and points in each subset can be rotated about the axis of symmetry in the object coordinate system through any point in the subset. The degree is generated, for example, the hexagonal support area at the bottom of a cola bottle can, where n is the number of discrete symmetries;

[0068] An asymmetric point set is a set of points that are not included in the three types of symmetry mentioned above, such as an irregular doll;

[0069] Step A1.2: Determine the symmetry form corresponding to each point on the CAD model by manually selecting the spatial region and manually selecting the threshold;

[0070] Manual spatial region selection specifically involves: for each possible form of symmetry, selecting the region covered by that form of symmetry. The selection process is achieved by taking the intersection and union of the spatial regions multiple times.

[0071] Manual threshold selection specifically involves calculating the threshold values ​​for all points p in this symmetry form on the model. i , looking for p i The nearest point q i Calculate index l = ∑ i (p i -q i ) 2 Each point is divided into different symmetry forms based on a manual threshold and l, and then added to the corresponding symmetry form point set.

[0072] Step A1.3: Complete the points for discrete symmetry and about-axis discrete symmetry, specifically as follows:

[0073] For points in discrete symmetry and about-axis discrete symmetry forms, the corresponding point p of each point in that symmetry form is... i Add it to the model to ensure that all points p in this symmetric form i , looking for p i The nearest point q i =p i This leads to the calculation index l = 0;

[0074] Step A1.4: Assign the same code to symmetrical points in each symmetry form. Recursively cluster the points on the model using a clustering algorithm to generate a binary code for the CAD model containing symmetry information. Construct a mapping relationship between the binary code and the coordinates of the points in the CAD model. Specifically:

[0075] Traverse the points in each symmetric form's point set, and take all points with index l = 0 as a subset; each point in the asymmetric form's point set forms its own subset; for each subset of all symmetric forms, take one point and divide it into two classes using a clustering algorithm to obtain the first binary code {0,1}. For regions where the first binary code is 0, divide them into two classes again using a clustering algorithm to obtain the second binary code {0,1}. Similarly, for regions where the first binary code is 1, divide them into two classes using a clustering algorithm to obtain the second binary code {0,1}. Recursively traverse this process to obtain a total of sixteen binary codes.

[0076] Finally, all subsets of points are traversed, and the same binary code is assigned to all points in each subset, thus completing the encoding of the CAD model points. This establishes a one-to-one mapping between the coordinates of the CAD model points and the binary code, and constructs the mutual mapping relationship between the binary code of the CAD model and the coordinates of the CAD model points, enabling the conversion between 3D coordinates and binary codes.

[0077] Step A2: Render the binary code of the CAD model in the true pose of the target object, including:

[0078] Step A2.1: Project the coordinates of the CAD model points onto the image using rasterized rendering with the true pose;

[0079] Step A2.2: Find the coordinates of the CAD model point closest to the coordinates on the image. Based on the mapping relationship between the binary code of the CAD model and the coordinates of the CAD model point, convert the found coordinates of the CAD model point into the corresponding binary code to obtain the binary code feature map corresponding to the RGB image.

[0080] Specifically, in step S3, the output 6D pose estimation result is a scale-independent translation vector SITE and a 6D matrix R. 6d We then convert these two back to the representations of the translation vector t and the rotation matrix R, respectively.

[0081] Specifically, training and testing sets are constructed by combining binary encoded feature maps, and an optimizer is used to train the dilated spatial pooling pyramid coding network and the pose estimation network.

[0082] Data augmentation can be performed during the construction of training and testing sets, including: random translation and scaling of the obtained target object image, augmentation of brightness and contrast, occlusion using the target object as an occluder, random addition of Gaussian noise and block occlusion regions; calculation of the corresponding camera intrinsic parameter matrix, and cropping of the complete mask, visible part mask, and binary encoded feature map of the target object. The cropped result should match the target object image.

[0083] like Figure 2 As shown, this embodiment presents the framework of the Symmetric Object Pose Estimation Network (SymNet).

[0084] like Figure 3 As shown, the input image in this embodiment is a preprocessed image of the target object. Specifically, during training, the ground truth detection box is randomly cropped, translated, and enlarged to obtain the target object image; during testing, the result of the target detector is cropped and enlarged.

[0085] The network in this embodiment is a backbone-dilated spatial pooling pyramid coding network-region-correspondence pose estimation network structure. The backbone branch uses a modified ResNet-34 to match the feature dimensions of the network output to the input of the dilated spatial pooling pyramid coding network. The backbone branch fuses multi-layer feature information before outputting. Given a target object image, the backbone branch outputs multi-layer feature information of the target object image. The coding network in this embodiment is built on the dilated spatial pooling pyramid coding network, which simultaneously outputs the complete mask, the visible part mask, and the binary encoded feature map of the target object. Through multi-scale information fusion, the accuracy of the network output is improved. The pose estimation network in this embodiment is built based on convolutional layers and fully connected layers.

[0086] like Figure 4 As shown, the symmetry annotation types of the CAD model in this embodiment include continuous symmetry, discrete symmetry around an axis, and asymmetry. Then, a clustering algorithm is used to recursively cluster the points on the CAD model, ultimately generating a binary code of the CAD model containing symmetry information. When rendering the encoded image, the model point coordinates are projected onto the image using rasterized rendering based on the ground truth pose of the target object. The coordinates of the CAD model point closest to the coordinates on the image are found. Based on the mapping relationship between the binary code of the CAD model and the coordinates of the CAD model point, the found CAD model point coordinates are converted into the corresponding binary code, resulting in a binary code feature map.

[0087] The training in this embodiment is implemented using PyTorch, employing the Ranger optimizer. The batch size is 32, and the basic learning rate is 1e-4. The network is trained for 500 epochs.

[0088] Table 1 shows the results of this embodiment compared with other methods on the TLESS dataset. All methods listed in the table are trained using simulated images (PBR only), not real images. The evaluation metrics in Table 1 adopt widely used BOP metrics, including Visible Surface Discrepancy (VSD), Maximum Symmetry-Aware Surface Distance (MSSD), and Maximum Symmetry-Aware Projection Distance (MSPD). AR in the table below represents the average value of the above metrics. For details, see Sundermeyer M, HodaňT, Labbe Y, et al. Bop challenge 2022 on detection, segmentation and pose estimation of specific rigid objects [C] / / Proceedings of the IEEE / CVF Conference on Computer Vision and Pattern Recognition.2023:2784-2793.

[0089] Table 1: Comparison of BOP metrics with state-of-the-art methods on the TLESS dataset.

[0090] CDPNv2 0.407 0.303 0.338 0.579 1.849 EPOS 0.467 0.380 0.403 0.619 1.992 ZebraPose 0.677 0.597 0.636 0.466 0.25 Surfemb 0.735 0.661 0.686 0.857 9.043 SymNet (Ours) 0.723 0.620 0.679 0.868 0.058

[0091] The method of this invention achieved a test speed of 0.058 seconds / frame on the GTX3090, while the ZebraPose test speed on the GTX3090 was only 0.25 seconds / frame.

[0092] On the other hand, the present invention also provides a real-time monocular 6D pose estimation system suitable for symmetrical objects, comprising:

[0093] The object detection module is configured to acquire a preprocessed target object image based on the original object image;

[0094] The object pose estimation module is configured to cascade a pre-trained dilated spatial pooling pyramid coding network and a pose estimation network to output a 6D pose estimation result; the 6D pose estimation process is as follows:

[0095] Symmetry annotation is performed on the CAD model of the target object. The symmetry form of each point on the CAD model is determined by spatial region selection and threshold selection. The points on the CAD model are recursively clustered by a clustering algorithm to generate a binary code of the CAD model containing symmetry information. The mapping relationship between the binary code of the CAD model and the coordinates of the points on the CAD model is constructed. The binary code of the CAD model is rendered under the ground truth pose of the target object to obtain the binary code feature map corresponding to the RGB image as the label.

[0096] The target object image is processed using a trained dilated spatial pooling pyramid coding network to output the complete mask, visible part mask, and binary encoded feature map of the target object.

[0097] The trained pose estimation network is used to process the complete mask, visible part mask, and binary encoded feature map of the target object, outputting a 6D pose estimation result. The output 6D pose estimation result is a scale-independent translation vector SITE and a 6D matrix R. 6d We then convert these two back to the representations of the translation vector t and the rotation matrix R, respectively.

[0098] In addition, the present invention also provides a 6D pose estimation system for symmetrical objects using an RGBD camera. Compared with the above system, the object detection module also has a depth image preprocessing function, which can preprocess the depth map corresponding to the target object image; and adds a point cloud iterative optimization module to further optimize the output translation vector t and rotation matrix R.

[0099] Corresponding to the aforementioned embodiments of the real-time monocular 6D pose estimation method applicable to symmetrical objects, the present invention also provides an embodiment of a real-time monocular 6D pose estimation device applicable to symmetrical objects.

[0100] See Figure 5 The real-time monocular 6D pose estimation device for symmetrical objects provided in this embodiment of the invention includes a memory and one or more processors. The memory stores executable code, and when the processor executes the executable code, it is used to implement a real-time monocular 6D pose estimation method for symmetrical objects in the above embodiment.

[0101] The embodiment of the real-time monocular 6D pose estimation device for symmetrical objects provided by this invention can be applied to any device with data processing capabilities, such as a computer or other similar device. The device embodiment can be implemented through software, hardware, or a combination of both. Taking software implementation as an example, as a logical device, it is formed by the processor of any data processing device loading the corresponding computer program instructions from non-volatile memory into memory for execution. From a hardware perspective, such as... Figure 5 The diagram shown is a hardware structure diagram of any device with data processing capabilities, including a real-time monocular 6D pose estimation device for symmetrical objects provided by the present invention. (Except for...) Figure 5 In addition to the processor, memory, network interface, and non-volatile memory shown, any data processing device in the embodiment may also include other hardware depending on the actual function of the data processing device, which will not be described in detail here.

[0102] The specific implementation process of the functions and roles of each unit in the above-mentioned equipment can be found in the implementation process of the corresponding steps in the above-mentioned method, and will not be repeated here.

[0103] For the device embodiments, since they basically correspond to the method embodiments, the relevant parts can be referred to in the description of the method embodiments. The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of the present invention according to actual needs. Those skilled in the art can understand and implement this without creative effort.

[0104] This invention also provides a computer-readable storage medium storing a program thereon, which, when executed by a processor, implements a real-time monocular 6D pose estimation method for symmetrical objects as described in the above embodiments.

[0105] The computer-readable storage medium can be an internal storage unit of any data processing device as described in any of the foregoing embodiments, such as a hard disk or memory. The computer-readable storage medium can also be an external storage device of any data processing device, such as a plug-in hard disk, smart media card (SMC), SD card, flash card, etc., equipped on the device. Furthermore, the computer-readable storage medium can include both internal storage units and external storage devices of any data processing device. The computer-readable storage medium is used to store the computer program and other programs and data required by the data processing device, and can also be used to temporarily store data that has been output or will be output.

[0106] The above embodiments are used to explain and illustrate the present invention, but not to limit the present invention. Any modifications and changes made to the present invention within the spirit and scope of the claims shall fall within the protection scope of the present invention.

Claims

1. A method for real-time monocular 6D pose estimation of symmetric objects, characterized in that, include: The original object image is acquired and preprocessed to obtain the target object image; Symmetry annotation is performed on the CAD model of the target object. The symmetry form of each point on the CAD model is determined by spatial region selection and threshold selection. The points on the CAD model are recursively clustered by a clustering algorithm to generate a binary code of the CAD model containing symmetry information. The mapping relationship between the binary code of the CAD model and the coordinates of the points on the CAD model is constructed. The binary code of the CAD model is rendered under the ground truth pose of the target object to obtain the binary code feature map corresponding to the RGB image as the label. The target object image is processed using a trained dilated spatial pooling pyramid coding network to output the complete mask, visible part mask, and binary encoded feature map of the target object. The trained pose estimation network is used to process the complete mask, visible part mask, and binary encoded feature map of the target object, and outputs 6D pose estimation results. The specific steps for performing symmetrical annotation on the CAD model of the target object are as follows: The target object's CAD model has a set of points with four types of symmetry: continuous symmetry, discrete symmetry, discrete symmetry around an axis, and asymmetry. Based on the shape of the target object's CAD model, select all symmetry forms contained in the target object's CAD model; The symmetry form corresponding to each point on the CAD model is determined by manually selecting spatial regions and manually selecting thresholds. The manual spatial region selection specifically involves: for each type of symmetry, selecting the area covered by that type of symmetry; The manual threshold selection specifically involves calculating all points on the model under this symmetry form. , looking for nearest point Calculation indicators According to manual threshold and Divide each point into different symmetry forms and add them to the corresponding symmetry form point set; For points in discrete symmetry and about-axis discrete symmetry forms, the corresponding points of each point in that symmetry form are... Add it to the model to ensure that all points in this symmetric form... , looking for nearest point This leads to the calculation of indicators ; Traverse the points in the point set for each symmetry form and calculate the index. All points are treated as a subset of points; Each point in the asymmetric point set forms a subset; for each subset of all symmetric points, one point is selected and divided into two classes using a clustering algorithm, yielding the first binary code. For regions where the first binary code is 0, a clustering algorithm is used to divide them into two categories to obtain the second binary code. Similarly, for regions where the first binary code is 1, a clustering algorithm is used to divide them into two categories to obtain the second binary code. By recursively traversing the above process, a total of sixteen binary bits are obtained. Traverse all subsets of points and assign the same binary code to all points in each subset to complete the encoding of the CAD model points. This establishes a one-to-one mapping between the coordinates of the CAD model points and the binary code, and constructs a mutual mapping relationship between the binary code of the CAD model and the coordinates of the CAD model points, thus realizing the conversion between 3D coordinates and binary code.

2. The real-time monocular 6D pose estimation method for symmetrical objects according to claim 1, characterized in that, The process of acquiring the original object image and performing preprocessing to obtain the target object image specifically involves: Obtain the original object image; use an object detection algorithm to detect objects in the original image; expand the detected rectangular region into a square region, and expand the boundary by 10% to 50%; crop the original object image according to the coordinates of the expanded region to obtain the target object image.

3. The real-time monocular 6D pose estimation method for symmetrical objects according to claim 1, characterized in that, The rendering of the binary encoding of the CAD model in the true pose of the target object specifically involves: The coordinates of points in the CAD model are projected onto the image using rasterized rendering with the true pose. Find the coordinates of the CAD model point closest to the coordinates on the image. Based on the mapping relationship between the binary code of the CAD model and the coordinates of the CAD model point, convert the found coordinates of the CAD model point into the corresponding binary code to obtain the binary code feature map.

4. The real-time monocular 6D pose estimation method for symmetrical objects according to claim 1, characterized in that, Training and testing sets are constructed by combining binary encoded feature maps, and an optimizer is used to train the dilated spatial pooling pyramid coding network and the pose estimation network.

5. The real-time monocular 6D pose estimation method for symmetrical objects according to claim 4, characterized in that, Data augmentation is performed during the construction of training and test sets. Specifically, the obtained target object images are randomly translated and scaled, brightness and contrast are augmented, the target objects are used as occluders, and Gaussian noise and block occlusion regions are randomly added. Calculate the corresponding camera intrinsic parameter matrix, and extract the complete mask, visible part mask, and binary encoded feature map of the target object. The extracted result should match the image of the target object.

6. The real-time monocular 6D pose estimation method for symmetrical objects according to claim 4, characterized in that, The process of training the dilated spatial pooling pyramid coding network and the pose estimation network using an optimizer specifically involves: Construct loss functions for the complete mask, visible part mask, and binary encoded feature map of the target object. Combine these with the pose loss function and the position loss function, and sum all the above loss functions proportionally to obtain the final loss function.

7. A real-time monocular 6D pose estimation system suitable for symmetrical objects, characterized in that, include: The object detection module is configured to acquire a preprocessed target object image based on the original object image; The object pose estimation module is configured to cascade a pre-trained dilated spatial pooling pyramid coding network and a pose estimation network to output a 6D pose estimation result; the 6D pose estimation process is as follows: Symmetry annotation is performed on the CAD model of the target object. The symmetry form of each point on the CAD model is determined by spatial region selection and threshold selection. The points on the CAD model are recursively clustered by a clustering algorithm to generate a binary code of the CAD model containing symmetry information. The mapping relationship between the binary code of the CAD model and the coordinates of the points on the CAD model is constructed. The binary code of the CAD model is rendered under the ground truth pose of the target object to obtain the binary code feature map corresponding to the RGB image as the label. The target object image is processed using a trained dilated spatial pooling pyramid coding network to output the complete mask, visible part mask, and binary encoded feature map of the target object. The trained pose estimation network is used to process the complete mask, visible part mask, and binary encoded feature map of the target object, and outputs 6D pose estimation results. The specific steps for performing symmetrical annotation on the CAD model of the target object are as follows: The target object's CAD model has a set of points with four types of symmetry: continuous symmetry, discrete symmetry, discrete symmetry around an axis, and asymmetry. Based on the shape of the target object's CAD model, select all symmetry forms contained in the target object's CAD model; The symmetry form corresponding to each point on the CAD model is determined by manually selecting spatial regions and manually selecting thresholds. The manual spatial region selection specifically involves: for each type of symmetry, selecting the area covered by that type of symmetry; The manual threshold selection specifically involves calculating all points on the model under this symmetry form. , looking for nearest point Calculation indicators According to manual threshold and Divide each point into different symmetry forms and add them to the corresponding symmetry form point set; For points in discrete symmetry and about-axis discrete symmetry forms, the corresponding points of each point in that symmetry form are... Add it to the model to ensure that all points in this symmetric form... , looking for nearest point This leads to the calculation of indicators ; Traverse the points in the point set for each symmetry form and calculate the index. All points in the asymmetric point set form a subset; all points in the asymmetric point set each form a subset; for all subsets of all symmetric forms, one point is taken and divided into two classes using a clustering algorithm, yielding the first binary code. For regions where the first binary code is 0, a clustering algorithm is used to divide them into two categories to obtain the second binary code. Similarly, for regions where the first binary code is 1, a clustering algorithm is used to divide them into two categories to obtain the second binary code. By recursively traversing the above process, a total of sixteen binary bits are obtained. Traverse all subsets of points and assign the same binary code to all points in each subset to complete the encoding of the CAD model points. This establishes a one-to-one mapping between the coordinates of the CAD model points and the binary code, and constructs a mutual mapping relationship between the binary code of the CAD model and the coordinates of the CAD model points, thus realizing the conversion between 3D coordinates and binary code.

8. A real-time monocular 6D pose estimation device for symmetrical objects, comprising a memory and one or more processors, wherein the memory stores executable code, characterized in that, When the processor executes the executable code, it implements the real-time monocular 6D pose estimation method for symmetrical objects as described in any one of claims 1-6.

9. A computer-readable storage medium having a program stored thereon, characterized in that, When the program is executed by the processor, it implements the real-time monocular 6D pose estimation method for symmetrical objects as described in any one of claims 1-6.