Class-level object pose estimation method and system based on spherical shape proxy
Through the method based on spherical shape proxy, image and point cloud features are combined with Transformer encoder and spherical projection, the semantic inconsistency caused by object shape differences in class-level object position estimation is solved, and the object position estimation without the CAD model is realized, which improves the generalization ability of the method.
Patent Information
- Application Number
- CN202510589413.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-08
- Publication Date
- 2025-08-15
AI Technical Summary
The existing class-level object position estimation methods are not consistent in semantics when the shape of objects are large, resulting in difficult to effectively learn and generalize the corresponding relationships, and it is impossible to effectively estimate the rotation and translation of the object.
Using a method based on spherical shape agent, the input image and point cloud are received, and the translation and size of the object are predicted using the point cloud processing network model, and the image and point cloud features are extracted after normalization. Combined with the Transformer encoder and spherical projection, the spherical NOCS coordinates are obtained, and the Umeyama algorithm is used to solve the rotation results of the object.
It realizes that the rotation and translation of objects can be effectively estimated without relying on the object CAD model under different object shapes, which improves the generalization ability and practicality of the method.
Smart Images

Figure CN120495413A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer vision technology, and in particular to a category-level object pose estimation method and system based on a spherical shape agent. Background Art
[0002] Object pose estimation, which aims to predict the rotation and translation of observed objects, has become a key task in computer vision and has broad applications in fields such as robotics, augmented reality, and autonomous driving. While many traditional instance-level object pose estimation methods have achieved promising performance, their reliance on CAD models of objects limits their generalization capabilities. To address this issue, category-level object pose estimation has emerged. It aims to infer the rotation, translation, and size of any object in a given category without requiring a CAD model, offering greater generalization and practical potential.
[0003] Most existing category-level object pose estimation methods adopt a correspondence-based framework. They first establish a correspondence between the observed object in the camera coordinate space and the normalized object coordinate space (NOCS), and then use a pose fitting algorithm to solve the object's pose. However, most existing methods use a point-based representation, using discrete points obtained by back-projection from the observed depth image as the representation of the observation data. The NOCS coordinates of these points are then obtained through geometric normalization and alignment of the object shape, a process that is shape-dependent. As a result, observation points on different object shapes, even if they have similar semantics, will be mapped to different NOCS coordinates, resulting in inconsistent semantics. This makes it difficult to effectively learn and generalize correspondences when the object shapes vary greatly. Summary of the Invention
[0004] In order to address the deficiencies mentioned in the above background technology, the purpose of the present invention is to provide a category-level object pose estimation method and system based on spherical shape agents.
[0005] In a first aspect, the purpose of the present invention can be achieved by the following technical solution: a category-level object pose estimation method based on a spherical shape agent, the method comprising the following steps:
[0006] Receive input images and point clouds, input the input images and point clouds into a pre-established point cloud processing network model, obtain the translation and size of the object, normalize the point cloud based on the translation and size of the object, and obtain a processed point cloud;
[0007] Perform feature extraction on the input image and point cloud to obtain image features and point cloud features. Extract high-level information of the image features and point cloud features respectively. Perform feature concatenation based on the high-level information of the image features and point cloud features to obtain the features of each point. Perform spherical projection on the processed point cloud and the features of each point to obtain a new feature representation on the sphere.
[0008] The new feature representation on the sphere is input into the pre-established Transformer encoder. Combined with the learnable position encoding of each spherical anchor point, the feature representation of the spherical anchor point is obtained, and the spherical NOCS coordinates are obtained. Based on the Umeyama algorithm, the spherical anchor point coordinates and the spherical NOCS coordinates are solved to obtain the rotation result of the object.
[0009] In conjunction with the first aspect, in certain implementations of the first aspect, the method further includes: inputting the input image and point cloud into a pre-established point cloud processing network model to obtain the translation and size of the object:
[0010] For the input image I and point cloud P O , predicting the translation of objects through the point cloud processing network model PointNet++ and size
[0011] In combination with the first aspect, in some implementations of the first aspect, the method further includes: normalizing the point cloud based on the translation and size of the object to obtain a processed point cloud
[0012] In conjunction with the first aspect, in certain implementations of the first aspect, the method further includes: the process of extracting high-level information of image features and point cloud features respectively includes:
[0013] For image features, the underlying RGB texture information is used and the pre-trained image feature extraction network DINOv2 is used to extract high-level semantic information; for point cloud features, the underlying radial distance is used and the improved point cloud feature extraction network ColorPointNet++ is used to extract high-level geometric information.
[0014] In conjunction with the first aspect, in certain implementations of the first aspect, the method further includes: a process of acquiring the new feature representation on the spherical surface:
[0015] The unit sphere is evenly divided using the HEALPix grid, so that the area of each grid area is equal, and the center of each grid is used as a spherical anchor point. m∈{1,…,M}, and project the processed point cloud P onto the spherical mesh, the feature F of the corresponding point PThen it is assigned to the corresponding spherical anchor point to obtain the new feature representation on the sphere
[0016] In conjunction with the first aspect, in certain implementations of the first aspect, the method further includes: inputting the new feature representation on the spherical surface into a pre-established Transformer encoder, and combining the learnable position encoding of each spherical anchor point to obtain a feature representation of the spherical anchor point:
[0017] Through the self-attention mechanism of the Transformer encoder, combined with the learnable position encoding of each spherical anchor point The features between anchor points are interacted and propagated globally to obtain the enhanced feature representation of each spherical anchor point.
[0018]
[0019] where F = F L , F 0 =F A The Attention operation is defined as:
[0020]
[0021] Q l =F l +E pos ,K l =F l +E pos ,V l =F l ,.
[0022] In combination with the first aspect, in certain implementations of the first aspect, the method further includes: a process of acquiring the spherical NOCS coordinates:
[0023] Use the MLP predictor to predict the spherical NOCS coordinates corresponding to each spherical anchor point Since the anchor point and NOCS coordinates on the spherical surface are both located on the unit sphere, through L 2 The norm normalizes the output of the MLP to obtain the final spherical NOCS coordinates:
[0024]
[0025] Use hyperbolic correspondence loss function for supervision:
[0026]
[0027] In a second aspect, in order to achieve the above-mentioned object, the present invention discloses a category-level object pose estimation system based on a spherical shape agent, comprising:
[0028] A data processing module is used to receive input images and point clouds, input the input images and point clouds into a pre-established point cloud processing network model, obtain the translation and size of the object, and normalize the point cloud based on the translation and size of the object to obtain a processed point cloud;
[0029] The feature processing module is used to extract features from the input image and point cloud to obtain image features and point cloud features, extract high-level information of the image features and point cloud features respectively, perform feature splicing based on the high-level information of the image features and point cloud features to obtain the features of each point, and perform spherical projection based on the processed point cloud and the features of each point to obtain a new feature representation on the spherical surface;
[0030] The rotation estimation module is used to input the new feature representation on the spherical surface into the pre-established Transformer encoder, combine the learnable position encoding of each spherical anchor point to obtain the feature representation of the spherical anchor point, obtain the spherical NOCS coordinates, and solve the spherical anchor point coordinates and spherical NOCS coordinates based on the Umeyama algorithm to obtain the rotation result of the object.
[0031] In another aspect of the present invention, in order to achieve the above-mentioned purpose, a terminal device is disclosed, including a memory, a processor, and a computer program stored in the memory and capable of running on the processor. The memory stores a computer program capable of running on the processor, and when the processor loads and executes the computer program, it adopts the above-mentioned category-level object pose estimation method based on a spherical shape agent.
[0032] In another aspect of the present invention, in order to achieve the above-mentioned purpose, a computer-readable storage medium is disclosed, in which a computer program is stored. When the computer program is loaded and executed by a processor, the category-level object pose estimation method based on a spherical shape agent as described above is adopted.
[0033] Beneficial effects of the present invention:
[0034] The present invention uses the unit sphere as a proxy shape shared between different objects, and learns coordinate transformations that are independent of object shape through spherical representation, thereby avoiding the negative impact of shape changes of objects within a class on object pose estimation. BRIEF DESCRIPTION OF THE DRAWINGS
[0035] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, those skilled in the art can derive other drawings based on these drawings without inventive effort.
[0036] Figure 1 It is a schematic flow chart of the method of the present invention;
[0037] Figure 2 Schematic diagram of the framework of the category-level object pose estimation method based on spherical shape agent of the present invention;
[0038] Figure 3 It is a schematic diagram of the system structure of the present invention. DETAILED DESCRIPTION
[0039] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making any creative efforts shall fall within the scope of protection of the present invention.
[0040] Example 1:
[0041] like Figure 1 As shown, a category-level object pose estimation method based on a spherical shape agent comprises the following steps:
[0042] S101: receiving an input image and a point cloud, inputting the input image and the point cloud into a pre-established point cloud processing network model, obtaining the translation and size of the object, and normalizing the point cloud based on the translation and size of the object to obtain a processed point cloud;
[0043] The process of inputting the input image and point cloud into the pre-established point cloud processing network model to obtain the translation and size of the object:
[0044] For the input image I and point cloud P O , predicting the translation of objects through the point cloud processing network model PointNet++ and size
[0045] The point cloud is normalized based on the translation and size of the object to obtain the processed point cloud
[0046] S102: performing feature extraction on the input image and point cloud to obtain image features and point cloud features, extracting high-level information of the image features and point cloud features respectively, performing feature concatenation based on the high-level information of the image features and point cloud features to obtain features of each point, and performing spherical projection on the processed point cloud and the features of each point to obtain a new feature representation on the spherical surface;
[0047] The process of extracting high-level information of image features and point cloud features respectively includes:
[0048] For image features, we use low-level RGB texture information and high-level semantic information extracted by the pre-trained image feature extraction network DINOv2. For point cloud features, we use low-level radial distance and high-level geometric information extracted by the improved point cloud feature extraction network ColorPointNet++. ColorPointNet++ is our improved version of the original point cloud feature extraction network PointNet++. By replacing the input XYZ coordinates with RGB color values, it can effectively extract point cloud features with pose invariance.
[0049] The process of obtaining the new feature representation on the sphere:
[0050] The four features of each observation point are spliced together in the feature dimension and dimensionally compressed to obtain the features of each point Next, the HEALPix grid is used to evenly divide the unit sphere, ensuring that the area of each grid area is equal, and the center of each grid is used as a spherical anchor point. ,m∈{1,…,M}, and project the point cloud P onto these spherical grids, the features of the corresponding points F P Then it is assigned to the corresponding spherical anchor point to obtain the new feature representation on the sphere For a certain mesh, if multiple points are projected onto it, we use the feature of the point with the largest radial distance; if no point is projected onto it, we set the feature of the mesh to zero.
[0051] S103: Input the new feature representation on the spherical surface into the pre-established Transformer encoder, combine the learnable position encoding of each spherical anchor point to obtain the feature representation of the spherical anchor point, obtain the spherical NOCS coordinates, and solve the feature representation coordinates of the spherical anchor point and the spherical NOCS coordinates based on the Umeyama algorithm to obtain the rotation result of the object.
[0052] The process of inputting the new feature representation on the spherical surface into the pre-established Transformer encoder and combining it with the learnable position encoding of each spherical anchor point to obtain the feature representation of the spherical anchor point:
[0053] Through the self-attention mechanism of the Transformer encoder, combined with the learnable position encoding of each spherical anchor point The features between anchor points are interacted and propagated globally to obtain the enhanced feature representation of each spherical anchor point.
[0054]
[0055] where F = F L , F 0 =F A , the Attention operation is defined as:
[0056]
[0057] Q l =F l +E pos ,K l =F l +E pos ,V l =F l ,.
[0058] Use the MLP predictor to predict the spherical NOCS coordinates corresponding to each spherical anchor point Since the anchor point and NOCS coordinates on the spherical surface are both located on the unit sphere, the L 2 The norm normalizes the output of the MLP to obtain the final spherical NOCS coordinates:
[0059]
[0060] Use hyperbolic correspondence loss function for supervision:
[0061]
[0062] After obtaining the spherical NOCS coordinates, the Umeyama algorithm is used to solve the object's rotation R∈SO(3) based on the correspondence between the spherical anchor points and the spherical NOCS coordinates.
[0063] This invention can be applied to applications related to 3D understanding and interaction, estimating the six-degree-of-freedom pose and three-dimensional dimensions of objects. In practice, it can be installed as software on a personal phone or computer, providing real-time object pose estimation results. It can also be installed on a company's backend server, providing large-scale 3D understanding results, thereby providing fundamental support for subsequent applications such as robotics, augmented reality, and autonomous driving.
[0064] Example 2: The second aspect, as Figure 3As shown, in order to achieve the above-mentioned purpose, the present invention discloses a category-level object pose estimation system based on a spherical shape agent, comprising:
[0065] The data processing module 11 is used to receive an input image and a point cloud, input the input image and the point cloud into a pre-established point cloud processing network model, obtain the translation and size of the object, and normalize the point cloud based on the translation and size of the object to obtain a processed point cloud;
[0066] A feature processing module 12 is configured to perform feature extraction on the input image and point cloud to obtain image features and point cloud features, extract high-level information of the image features and point cloud features respectively, perform feature concatenation based on the high-level information of the image features and point cloud features to obtain features of each point, and perform spherical projection based on the processed point cloud and the features of each point to obtain a new feature representation on the spherical surface;
[0067] The rotation estimation module 13 is used to input the new feature representation on the spherical surface into the pre-established Transformer encoder, combine the learnable position encoding of each spherical anchor point, obtain the feature representation of the spherical anchor point, obtain the spherical NOCS coordinates, and solve the spherical anchor point coordinates and spherical NOCS coordinates based on the Umeyama algorithm to obtain the rotation result of the object.
[0068] Based on the same inventive concept, the present invention also provides a computer device, which includes: one or more processors and a memory for storing one or more computer programs; the program includes program instructions, and the processor is used to execute the program instructions stored in the memory. The processor may be a central processing unit (CPU), or other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field-programmable gate arrays (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. It is the computing core and control core of the terminal, which is used to implement one or more instructions, specifically for loading and executing one or more instructions in a computer storage medium to implement the above method.
[0069] It should be further explained that, based on the same inventive concept, the present invention also provides a computer storage medium having a computer program stored thereon, which executes the above method when executed by a processor. The storage medium can be any combination of one or more computer-readable media. The computer-readable medium can be a computer-readable signal medium or a computer-readable storage medium. The computer-readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electrical, magnetic, infrared, or semiconductor system, device or component, or any combination thereof. More specific examples (a non-exhaustive list) of computer-readable storage media include: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination thereof. In the present invention, a computer-readable storage medium can be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, device or component.
[0070] Throughout this specification, references to terms such as "one embodiment," "example," or "specific example" indicate that a specific feature, structure, material, or characteristic described in conjunction with that embodiment or example is included in at least one embodiment or example of the present disclosure. In this specification, schematic representations of these terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in any one or more embodiments or examples.
[0071] The above shows and describes the basic principles, main features and advantages of the present disclosure. Those skilled in the art should understand that the present disclosure is not limited to the above embodiments. The above embodiments and descriptions are merely illustrative of the principles of the present disclosure. Various changes and improvements may be made to the present disclosure without departing from the spirit and scope of the present disclosure, and such changes and improvements shall fall within the scope of the present disclosure.
Claims
1. A category-level object pose estimation method based on spherical shape proxy, characterized in that: The method comprises the following steps: Receive input images and point clouds, input the input images and point clouds into a pre-established point cloud processing network model, obtain the translation and size of the object, normalize the point cloud based on the translation and size of the object, and obtain a processed point cloud; Perform feature extraction on the input image and point cloud to obtain image features and point cloud features. Extract high-level information of the image features and point cloud features respectively. Perform feature concatenation based on the high-level information of the image features and point cloud features to obtain the features of each point. Perform spherical projection on the processed point cloud and the features of each point to obtain a new feature representation on the sphere. The new feature representation on the sphere is input into the pre-established Transformer encoder. Combined with the learnable position encoding of each spherical anchor point, the feature representation of the spherical anchor point is obtained, and the spherical NOCS coordinates are obtained. Based on the Umeyama algorithm, the spherical anchor point coordinates and the spherical NOCS coordinates are solved to obtain the rotation result of the object.
2. The method for class-level object pose estimation based on spherical shape proxy according to claim 1, characterized in that: The process of inputting the input image and point cloud into the pre-established point cloud processing network model to obtain the translation and size of the object: For the input image I and point cloud P O , predicting the translation of objects through the point cloud processing network model PointNet++ and size 3. The method for class-level object pose estimation based on spherical shape proxy according to claim 1, characterized in that: The point cloud is normalized based on the translation and size of the object to obtain the processed point cloud 4. The method for class-level object pose estimation based on spherical shape proxy according to claim 1, characterized in that: The process of extracting high-level information of image features and point cloud features respectively includes: For image features, the underlying RGB texture information is used and the pre-trained image feature extraction network DINOv2 is used to extract high-level semantic information; for point cloud features, the underlying radial distance is used and the improved point cloud feature extraction network ColorPointNet++ is used to extract high-level geometric information.
5. The method for class-level object pose estimation based on spherical shape proxy according to claim 1, characterized in that: The process of obtaining the new feature representation on the sphere: The unit sphere is evenly divided using the HEALPix grid, so that the area of each grid area is equal, and the center of each grid is used as a spherical anchor point. And project the processed point cloud P onto the spherical grid, the feature F of the corresponding point P Then it is assigned to the corresponding spherical anchor point to obtain the new feature representation on the sphere 6. The method for class-level object pose estimation based on spherical shape proxy according to claim 1, characterized in that: The process of inputting the new feature representation on the spherical surface into the pre-established Transformer encoder and combining it with the learnable position encoding of each spherical anchor point to obtain the feature representation of the spherical anchor point: Through the self-attention mechanism of the Transformer encoder, combined with the learnable position encoding of each spherical anchor point The features between anchor points are interacted and propagated globally to obtain the enhanced feature representation of each spherical anchor point. where F = F L , F 0 =F A , the Attention operation is defined as: Q l =F l +E pos ,K l =F l +E pos ,V l =F l ,。 7. The method for class-level object pose estimation based on spherical shape proxy according to claim 1, characterized in that: The process of obtaining the spherical NOCS coordinates: Use the MLP predictor to predict the spherical NOCS coordinates corresponding to each spherical anchor point Since the anchor point and NOCS coordinates on the spherical surface are both located on the unit sphere, through L 2 The norm normalizes the output of the MLP to obtain the final spherical NOCS coordinates: Use hyperbolic correspondence loss function for supervision:
8. A category-level object pose estimation system based on spherical shape proxy, characterized in that: include: A data processing module is used to receive input images and point clouds, input the input images and point clouds into a pre-established point cloud processing network model, obtain the translation and size of the object, and normalize the point cloud based on the translation and size of the object to obtain a processed point cloud; The feature processing module is used to extract features from the input image and point cloud to obtain image features and point cloud features, extract high-level information of the image features and point cloud features respectively, perform feature splicing based on the high-level information of the image features and point cloud features to obtain the features of each point, and perform spherical projection based on the processed point cloud and the features of each point to obtain a new feature representation on the spherical surface; The rotation estimation module is used to input the new feature representation on the spherical surface into the pre-established Transformer encoder, combine the learnable position encoding of each spherical anchor point to obtain the feature representation of the spherical anchor point, obtain the spherical NOCS coordinates, and solve the spherical anchor point coordinates and spherical NOCS coordinates based on the Umeyama algorithm to obtain the rotation result of the object.
9. A terminal device comprising a memory, a processor, and a computer program stored in the memory and capable of running on the processor, characterized in that: The memory stores a computer program that can be run on a processor. When the processor loads and executes the computer program, it adopts a category-level object pose estimation method based on a spherical shape agent according to any one of claims 1 to 7.
10. A computer-readable storage medium storing a computer program, wherein: When the computer program is loaded and executed by the processor, the method for class-level object pose estimation based on a spherical shape agent according to any one of claims 1 to 7 is adopted.