Class-level object pose estimation method and system based on robust corresponding relation

Through the robust correspondence method, the example segmentation network and geometric cross-attention layer are used to process class-level object position estimation, which solves the shape sensitivity and noise problems and achieves more accurate and stable position estimation.

CN120339569APending Publication Date: 2025-07-18DEEP SPACE EXPLORATION LABORATORY
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510444834.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-10
Publication Date
2025-07-18

AI Technical Summary

Technical Problem

Existing class-level object pose estimation methods are challenging in extracting shape sensitivity and pose invariance features, and noise-induced anomalies affect estimation accuracy.

Method used

Using a robust correspondence method, dense features are extracted through instance segmentation networks, combined with point cloud and image feature stitching, local and global geometric cross-attention layers are used to perform sparse feature interaction, and noise effects are eliminated through the abnormal point removal loss function, and finally robust pose and size estimation is performed.

Benefits of technology

It improves the accuracy and stability of object position estimation, ensures stronger generalization ability in complex scenarios, and eliminates interference from abnormal correspondence.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120339569A_ABST
    Figure CN120339569A_ABST
Patent Text Reader

Abstract

The invention discloses a category-level object pose estimation method and system based on a robust corresponding relation, and relates to the technical field of computer vision, and the method comprises the steps: obtaining an image, inputting the image into a pre-trained instance segmentation network model Mask-RCNN, outputting an extracted depth image, carrying out the extraction of dense features of the depth image, and obtaining a category-level object pose estimation result; dense point-by-point features are obtained; sparse feature interaction is carried out on the dense point-by-point features to obtain final key point features, robust pose and size estimation is carried out based on the final key point features to finally obtain pose and size estimation results of the object, the adverse effect of an abnormal corresponding relation on pose fitting is effectively eliminated, and the accuracy of pose fitting is improved. And the accuracy and the stability of object pose estimation are ensured.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of computer vision, and specifically, to a method and system for category-level object pose estimation based on robust correspondence relations. Background Art

[0002] Category-level object pose estimation is one of the core tasks in the field of computer vision, aiming to predict the pose and size of objects in a given category. This task has broad prospects in multiple applications such as robotic operation, augmented reality, and autonomous driving, and thus has received high attention from both academia and industry. Traditional object pose estimation methods usually rely on instance-level modeling, which requires providing a CAD model for each object instance, and this is not flexible enough and time-consuming in practical applications. Different from this, category-level object pose estimation methods do not rely on the CAD model of a single object, but model the object category to achieve general estimation of all objects in the same category. This category-level method has stronger generalization ability and can be applied in various complex practical scenarios.

[0003] Most of the existing category-level object pose estimation methods adopt a two-stage correspondence-based framework. In this framework, first, the correspondence relations between the observed object in the camera coordinate space and the object coordinate space are established, and then a pose fitting algorithm is used to accurately estimate the pose of the object. Although this method has made good progress, there are still some challenges. Among them, the first problem to be solved is how to extract features with shape sensitivity and pose invariance in the correspondence prediction stage. Since objects in the same category may have significant differences in shape, shape-sensitive features are crucial for accurately learning coordinate transformation. At the same time, pose-invariant features help maintain consistent mapping when the observation perspective of the object changes. Secondly, in the pose fitting stage, due to the possible presence of noise in the observed object point cloud, some abnormal correspondence relations are caused, and these abnormal correspondence relations will significantly affect the accurate fitting of the object pose. Therefore, removing abnormal correspondence relations and avoiding their interference become an important issue. Summary of the Invention

[0004] To solve the deficiencies mentioned in the above background art, the purpose of the present invention is to provide a method and system for category-level object pose estimation based on robust correspondence relations, ensuring the accuracy and stability of object pose estimation.

[0005] In a first aspect, the purpose of the present invention can be achieved by the following technical solutions: A method for category-level object pose estimation based on robust correspondence relations, the method comprising the following steps:

[0006] Obtain an image, input the image into a pre-trained instance segmentation network model Mask-RCNN, and output the extracted depth image. Extract dense features from the depth image to obtain dense point-by-point features;

[0007] Perform sparse feature interaction on the dense point-by-point features to obtain the final key point features, and perform robust pose and size estimation based on the final key point features to finally obtain the pose and size estimation results of the object.

[0008] Combined with the first aspect, in some implementation manners of the first aspect, the method further includes: when inputting the image into the pre-trained instance segmentation network model Mask-RCNN, obtain the cropped RGB image and the segmented depth image by extracting the segmentation mask of the object, and project the depth image back into the three-dimensional space using the camera internal parameters to obtain the observed point cloud.

[0009] Combined with the first aspect, in some implementation manners of the first aspect, the method further includes: the process of extracting dense features from the depth image:

[0010] For the observed point cloud Use the designed point cloud feature extraction network PoseInv-PointNet++ to extract geometric features containing pose invariance, and predict the affine transformation matrix based on T-Net Align the input point cloud P obj Remove the absolute coordinate injection part in the original point cloud feature extraction network PointNet++. For the image I obj , use the pre-trained image feature extraction network DINOv2 to extract semantic features that are robust to pose transformation from it, and obtain dense point-by-point features through point-by-point feature stitching of the point cloud and image features and compression of the feature dimensions

[0011] Combined with the first aspect, in some implementation manners of the first aspect, the method further includes: when performing sparse feature interaction on the dense point-by-point features, extract a sparse key point set from the original point cloud P obj

[0012] using the farthest point sampling method FPS and retrieve the corresponding key point features from the dense feature F obj Perform shape-sensitive interaction between the key point feature F and the point cloud feature F kpt and the point cloud feature F obj The calculation involves pairwise distance encoding and K-angle encoding. For the key point P n and the point set interacting with it Pairwise distance encoding By performing on is obtained by applying sine encoding; while the K - angle encoding needs to select P att from P n 's K - nearest neighbor point set By applying sine encoding to obtain, the final geometric descriptor is calculated as follows:

[0013]

[0014] Combined with the first aspect, in some implementations of the first aspect, the method further includes: the feature interaction process of the sparse feature interaction includes L geometric interaction modules, each geometric interaction module consists of a local geometric cross - attention layer and a global geometric self - attention layer. In the local geometric cross - attention layer, select key point P obj from P n 's K local nearest neighbors and the corresponding features and calculate the corresponding geometric descriptor Use the attention mechanism for feature aggregation to obtain the locally enhanced key - point features The process is represented by the following formula:

[0015]

[0016] where the initially input key - point feature GCA represents the geometric cross - attention operation, which is defined as follows:

[0017] GCA(q,C,E)=Attention(q,C + E,C + E),

[0018]

[0019] Combined with the first aspect, in some implementations of the first aspect, the method further includes: in the global geometric self - attention layer, key point P n performs global interaction with all key points, and takes the average of the global geometric descriptors of each key point in the m dimension to obtain the geometric structure encoding Through the geometric self - attention operation GSA, aggregate the global information into the key - point features, and the formula is:

[0020]

[0021] GSA(C,E)=Attention(C + E,C + E,C + E).

[0022] After passing through L geometric interaction modules, the final key-point features are obtained.

[0023] Combined with the first aspect, in some implementations of the first aspect, the method further includes: the process of performing robust pose and size estimation based on the final key-point features includes:

[0024] By predicting the NOCS coordinates of each key point And combining with the anomaly score The prediction of NOCS coordinates is implemented through an MLP-based network, and the NOCS loss function is calculated by L 2 distance, expressed as:

[0025]

[0026] The prediction process of the anomaly score is expressed by the formula as:

[0027] O = Sigmoid(MLP([F, MLP(P)]))

[0028] By combining with the NOCS loss, a robust correspondence loss function is formed:

[0029]

[0030] Where I n = 1 - O n , representing the score of each key point belonging to a normal point. By applying a threshold to the anomaly score O, abnormal correspondences are filtered out, and for the remaining correspondences, the rotation R and translation t are solved through the Umeyama algorithm to obtain the pose of the object. The size estimation of the object is predicted through an MLP regression network and supervised using L 2 loss.

[0031] In a second aspect, to achieve the above object, the present invention discloses a category-level object pose estimation system based on robust correspondence, including:

[0032] A feature extraction module, configured to obtain an image, input the image into a pre-trained instance segmentation network model Mask-RCNN, output the obtained depth image, and perform extraction of dense features on the depth image to obtain dense point-by-point features;

[0033] A pose estimation module, configured to perform sparse feature interaction on the dense point-by-point features to obtain the final key-point features, perform robust pose and size estimation based on the final key-point features, and finally obtain the pose and size estimation results of the object.

[0034] In another aspect of the present invention, to achieve the above object, a terminal device is disclosed, which includes a memory, a processor, and a computer program stored in the memory and capable of running on the processor. The memory stores a computer program capable of running on the processor. When the processor loads and executes the computer program, a method for class-level object pose estimation based on robust correspondence as described above is adopted.

[0035] In yet another aspect of the present invention, to achieve the above object, a computer-readable storage medium is disclosed. The computer-readable storage medium stores a computer program. When the computer program is loaded and executed by a processor, a method for class-level object pose estimation based on robust correspondence as described above is adopted.

[0036] Advantages of the present invention:

[0037] By improving the general point cloud feature extraction network and introducing a pose-invariant geometric descriptor in the present invention, features that are robust to object pose changes are extracted. Further, by introducing local geometric cross-attention and global geometric self-attention operations, features for object shape perception can be effectively captured, promoting the prediction of correspondences. Aiming at the problem of abnormal points in the observed point cloud caused by depth cameras and object segmentation masks, through the design of a correspondence loss function based on outlier removal, the adverse effects of abnormal correspondences on pose fitting are effectively eliminated, ensuring the accuracy and stability of object pose estimation. Brief Description of the Drawings

[0038] To more clearly illustrate the technical solutions in the embodiments of the present invention or in the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, for those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings;

[0039] Figure 1 It is a schematic flowchart of the method of the present invention;

[0040] Figure 2 It is a schematic framework flowchart of the present invention;

[0041] Figure 3 It is a schematic system structure diagram of the present invention. Detailed Embodiments

[0042] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0043] Embodiment 1:

[0044] As Figure 1 shown, a category-level object pose and size estimation method based on robust correspondence relationships, the method includes the following steps:

[0045] S101: Obtain an image, input the image into a pre-trained instance segmentation network model Mask-RCNN, output the extracted depth image, and extract dense features from the depth image to obtain dense point-by-point features;

[0046] Specifically, dense feature extraction. Given the input RGB-D image, first use the pre-trained instance segmentation network Mask-RCNN to extract the segmentation mask of the object, further obtain the cropped RGB image and the segmented depth image, and project the depth image back into the three-dimensional space using the camera intrinsics to obtain the observed point cloud of the object. Then, extract the pose-invariant point-by-point dense features of the observed image and the point cloud respectively, and integrate the information of the two modalities through dense fusion. Specifically, for the observed point cloud P obj use the designed point cloud feature extraction network PoseInv-PointNet++ to extract the geometric features with pose invariance from it, and introduce T-Net to predict the affine transformation matrix to achieve the alignment of the input point cloud P obj and thus eliminate the influence of pose changes on feature extraction. In addition, the absolute coordinate injection part in the original point cloud feature extraction network PointNet++ is also removed to ensure that the extracted features are independent of pose. For the image I obj , use the pre-trained image feature extraction network DINOv2 to extract the semantic features that are robust to pose transformation from it. Finally, through the point-by-point feature stitching of the point cloud and image features and the compression of the feature dimension, dense point-by-point features are obtained as the input of the subsequent module.

[0047] S102: Perform sparse feature interaction on the dense point-by-point features to obtain the final key point features, and perform robust pose and size estimation based on the final key point features to finally obtain the object pose and size estimation results.

[0048] Specifically, sparse feature interaction. To reduce the computational overhead, a set of sparse key points is used to represent the shape of an object, and a sparse key point set is extracted from the original point cloud P obj by the farthest point sampling method FPS and the corresponding key point features are retrieved from the dense feature F obj Shape-sensitive interaction is performed between the key point feature F and the point cloud feature F kpt by introducing pose-invariant geometric descriptors based on distance and angle, while ensuring pose invariance at the same time. Its calculation involves a pairwise distance encoding and a K-angle encoding. For the key point P obj and the set of points it interacts with n The pairwise distance encoding is obtained by applying sine encoding to ; while the K-angle encoding first requires selecting the K-nearest neighbor point set of P from P att Then, the K-angle encoding n is obtained by applying sine encoding to The final geometric descriptor is calculated as follows:

[0049]

[0050] The feature interaction process includes L geometric interaction modules, each of which consists of a local geometric cross-attention layer and a global geometric self-attention layer. In the local geometric cross-attention layer, first, K obj nearest neighbors of the key point P n are selected from P local along with the corresponding features and the corresponding geometric descriptor is calculated Then, the attention mechanism is used for feature aggregation to obtain the locally enhanced key point feature This process can be represented by the following formula:

[0051]

[0052] where the initial input key point feature

[0053] GCA represents the geometric cross-attention operation, which is defined as follows:

[0054] GCA(q,C,E)=Attention(q,C+E,C+E),

[0055] ​​In the global geometric self-attention layer, key point P n performs global interaction with all key points. First, the global geometric descriptor of each key point takes the average in the m dimension to obtain its geometric structure encoding Then, through the geometric self-attention operation GSA, the global information is aggregated into the key point features, thereby realizing globally enhanced shape-sensitive feature learning. The specific formula is:

[0056]

[0057] GSA(C,E) = Attention(C + E, C + E, C + E).

[0058] After L geometric interaction modules, the final key point features are obtained. At this time, the key point features have fully fused local and global shape information and maintained pose invariance.

[0059] Specifically, for the problem of outliers in the observed point cloud caused by the noise of the depth camera and the object segmentation mask, a correspondence prediction method based on outlier removal is proposed. By predicting the NOCS coordinates of each key point and combining the outlier score effectively removes outliers and outlier correspondences, improving the robustness of pose estimation. The prediction of NOCS coordinates is achieved through a network based on MLP, and the NOCS loss function is calculated through the L 2 distance, and the specific expression is:

[0060]

[0061] The prediction process of the outlier score is expressed by the formula:

[0062] O = Sigmoid(MLP([F, MLP(P)])).

[0063] By combining with the NOCS loss, a robust correspondence loss function is formed:

[0064]

[0065] where I n = 1 - O n , representing the score of each key point belonging to normal points. Finally, in the inference stage, by applying a threshold to the outlier score O to filter out outlier correspondences, and then solving for the rotation R and translation t of the remaining correspondences through the Umeyama algorithm, the pose of the object is obtained. The size estimation of the object is predicted through an MLP regression network and uses L 2Supervise the losses.

[0066] Embodiment 2: Second aspect, as Figure 3 shown, to achieve the above object, the present invention discloses a category-level object pose estimation system based on robust correspondence relationships, including:

[0067] A feature extraction module 11, configured to obtain an image, input the image into a pre-trained instance segmentation network model Mask-RCNN, output a depth image obtained by extraction, and extract dense point-by-point features from the depth image to obtain dense point-by-point features;

[0068] A pose estimation module 12, configured to perform sparse feature interaction on the dense point-by-point features to obtain final key point features, perform robust pose and size estimation based on the final key point features, and finally obtain the object pose and size estimation results.

[0069] Based on the same inventive concept, the present invention further provides a computer device, which includes: one or more processors, and a memory for storing one or more computer programs; the program includes program instructions, and the processor is configured to execute the program instructions stored in the memory. The processor may be a central processing unit (CPU), or may also be other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. It is the computing core and control core of the terminal, and is used to implement one or more instructions. Specifically, it is used to load and execute one or more instructions in the computer storage medium to implement the above method.

[0070] It should be further noted that, based on the same inventive concept, the present invention also provides a computer storage medium, on which a computer program is stored. When the computer program is run by a processor, it executes the above method. The storage medium may be any combination of one or more computer-readable media. The computer-readable media may be a computer-readable signal medium or a computer-readable storage medium. The computer-readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electro-magnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples (non-exhaustive list) of the computer-readable storage medium include: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present invention, the computer-readable storage medium may be any tangible medium that contains or stores a program, and this program can be used by or in combination with an instruction execution system, apparatus, or device.

[0071] In the description of this specification, the descriptions referring to terms such as "one embodiment", "example", "specific example", etc. mean that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present disclosure. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described may be combined in any one or more embodiments or examples in a suitable manner.

[0072] The above has shown and described the basic principles, main features, and advantages of the present disclosure. Those skilled in the art of this industry should understand that the present disclosure is not limited by the above embodiments. The above embodiments and the descriptions in the specification only illustrate the principles of the present disclosure. Without departing from the spirit and scope of the present disclosure, the present disclosure will have various changes and improvements, and these changes and improvements all fall within the scope of the present disclosure claimed.

Claims

1. A method for category-level object pose estimation based on robust correspondences, characterized in that, The method includes the following steps: Obtain an image, input the image into a pre-trained instance segmentation network model Mask-RCNN, output the extracted depth image, and extract dense features from the depth image to obtain dense point-by-point features; Perform sparse feature interaction on the dense point-by-point features to obtain the final key-point features, and perform robust pose and size estimation based on the final key-point features to finally obtain the pose and size estimation results of the object.

2. The class-level object pose estimation method based on robust correspondence relationships according to claim 1, characterized in that Inputting the image into the pre-trained instance segmentation network model Mask-RCNN obtains a cropped RGB image and a segmented depth image by extracting the segmentation mask of the object, and projects the depth image back into the three-dimensional space using the camera intrinsics to obtain the observed point cloud.

3. A method for category-level object pose estimation based on robust correspondence relations according to claim 1, characterized in that The process of extracting dense features from the depth image: For the observed point cloud Use the designed point cloud feature extraction network PoseInv-PointNet++ to extract geometric features with pose invariance, and predict the affine transformation matrix based on T-Net Align the input point cloud P obj Remove the absolute coordinate injection part in the original point cloud feature extraction network PointNet++. For the image I obj Use the pre-trained image feature extraction network DINOv2 to extract semantic features robust to pose transformation from it. Through point-by-point feature splicing of point cloud and image features and compression of feature dimensions, dense point-by-point features are obtained 4. A method for category-level object pose estimation based on robust correspondence relations according to claim 1, wherein When performing sparse feature interaction on dense point-by-point features, a sparse key point set is extracted from the original point cloud P by the farthest point sampling method FPS obj and the corresponding key point features are retrieved from the dense feature F Shape-sensitive interaction is performed between the key point feature F obj and the point cloud feature F The calculation involves pairwise distance encoding and K-angle encoding. For the key point P kpt and the set of points it interacts with obj Pairwise distance encoding n is obtained by applying sine encoding to ; while K-angle encoding requires selecting the K-nearest neighbor point set of P from P att and is obtained by applying sine encoding to n The final geometric descriptor is calculated as follows:

5. A method for category-level object pose estimation based on robust correspondence relations according to claim 4, wherein The feature interaction process of the sparse feature interaction includes L geometric interaction modules. Each geometric interaction module consists of a local geometric cross-attention layer and a global geometric self-attention layer. In the local geometric cross-attention layer, key points P obj are selected from P n with K local nearest neighbor points and the corresponding features and calculate the corresponding geometric descriptors Feature aggregation is performed using the attention mechanism to obtain the locally enhanced key point features The process is represented by the following formula: Among them, the key point features of the initial input GCA represents a geometric cross-attention operation, which is defined as follows: GCA(q, C, E) = Attention(q, C + E, C + E), 6. A method for category-level object pose estimation based on robust correspondence relations according to claim 5, characterized in that The key point P in the global geometric self-attention layer n performs global interaction with all key points, and for the global geometric descriptor of each key point takes the average in the m dimension to obtain the geometric structure encoding Through the geometric self-attention operation GSA, aggregate the global information into the key point features, and the formula is: GSA(C, E) = Attention(C + E, C + E, C + E). After passing through L geometric interaction modules, the final key-point features are obtained 7. A method for category-level object pose estimation based on robust correspondence relations according to claim 1, characterized in that, The process of performing robust pose and size estimation based on the final key-point features includes: By predicting the NOCS coordinates of each key point and combining the anomaly scores The prediction of NOCS coordinates is achieved through an MLP-based network, and the NOCS loss function is calculated by the L 2 distance, expressed as: The prediction process of the anomaly score is expressed by the formula: O = Sigmoid(MLP([F, MLP(P)])) By combining with the NOCS loss, a robust correspondence loss function is formed: where I n = 1 - O n , representing the score that each key point belongs to a normal point. By applying a threshold to the abnormal score O, the abnormal correspondences are filtered out. For the remaining correspondences, the rotation R and translation t are solved by the Umeyama algorithm to obtain the pose of the object. The size estimation of the object is predicted by an MLP regression network and supervised using the L 2 loss.

8. A category-level object pose estimation system based on robust correspondence relationships, characterized in that, Including: A feature extraction module for obtaining an image, inputting the image into a pre-trained instance segmentation network model Mask-RCNN, outputting the extracted depth image, and extracting dense features from the depth image to obtain dense point-by-point features; A pose estimation module for performing sparse feature interaction on the dense point-by-point features to obtain the final key-point features, and performing robust pose and size estimation based on the final key-point features to finally obtain the pose and size estimation results of the object.

9. A terminal device, comprising a memory, a processor, and a computer program stored in the memory and capable of running on the processor, characterized in that The memory stores a computer program that can run on a processor. When the processor loads and executes the computer program, it adopts a method for class-level object pose estimation based on robust correspondence as described in any one of claims 1 to 7.

10. A computer-readable storage medium storing a computer program therein, characterized in that, When the computer program is loaded and executed by the processor, it adopts a method for class-level object pose estimation based on robust correspondence as described in any one of claims 1 to 7.