Zero sample category level object attitude estimation method and device

Through a two-stage iterative optimization strategy based on general semantic features, the problem of no category and pose ambiguity in object pose prediction is solved, and the object pose estimation with high precision and strong generalization ability is achieved, which is suitable for diverse object scenes.

CN120339383APending Publication Date: 2025-07-18INST OF SOFTWARE - CHINESE ACAD OF SCI
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510242655.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2024-12-31
Filing Date
2025-03-03
Publication Date
2025-07-18

AI Technical Summary

Technical Problem

Existing object pose prediction methods cannot effectively deal with unseen object categories, and are limited by the diversity of training data and drastic changes in object poses, resulting in feature degradation and pose ambiguity, affecting prediction accuracy and generalization ability.

Method used

A two-stage iterative optimization strategy based on general semantic features is adopted. First, rough pose estimation is performed through 2D general semantic features, and the object pose transformation relationship is calculated using Umeyama algorithm and RANSAC algorithm. Then, the pose is optimized by 3D general semantic features, combining the total loss function and regular term optimization parameters to achieve accurate pose prediction.

Benefits of technology

It improves the accuracy and generalization ability of object posture prediction, and can handle diverse object categories, especially when the shape and texture are different, with high accuracy and automation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120339383A_ABST
    Figure CN120339383A_ABST
Patent Text Reader

Abstract

The invention discloses a zero sample category level object attitude estimation method and device, and belongs to the field of computer vision and image processing. In order to solve the problem that a traditional non-general deep learning model cannot accurately process unseen category objects in object attitude prediction, the method mainly adopts a general semantic model to be combined with 2D and 3D general semantic features to carry out two-stage iterative optimization, and carries out rough attitude estimation and attitude optimization. The accuracy of object attitude prediction can be effectively improved, and the problems of object shape difference and attitude ambiguity are solved. The method has high automation degree, precision and generalization ability, and can be widely applied to the requirements of professional and popular object attitude estimation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of computer vision and image processing, and particularly relates to a zero-shot category-level object pose estimation method and device guided by general semantic features. Background Art

[0002] Objects have the characteristics of a large variety of types, complex shapes and textures, and are the main interaction objects in daily life. Therefore, helping the computer understand the position of objects in space and realizing accurate spatial perception of the interaction scene is crucial for augmented reality and virtual reality (AR / VR) applications. The intelligent helmet can obtain the accurate motion trajectory of objects in real time, improve the realism and efficiency of the user's interaction with objects, and realize natural interaction with virtual objects. At the same time, object pose estimation plays a key role in embodied intelligence applications, which can help robots autonomously understand diverse object shapes and poses, improve the efficiency and execution accuracy of the robot's operation task planning for objects, support adaptive perception of complex scenes and solve mutation problems, and greatly improve the stability of the embodied intelligence system.

[0003] However, the existing object pose prediction works mainly focus on known instance objects and known category objects. They need to collect a large amount of three-dimensional annotation data of objects to optimize or retrain the network. This method is time-consuming and laborious, and the model is limited by the types of objects and cannot be extended to new object categories. The general semantic model extracts general semantic features through pre-training on large-scale data, and these features can perceive the semantic association relationships between different objects to establish a pose mapping matrix. However, due to the feature degradation caused by the diversity of object poses and the pose ambiguity caused by the diversity of object shapes, the traditional object pose prediction method based on general semantic features has poor accuracy. To achieve high-precision object pose prediction, the following key problems need to be solved:

[0004] 1) Traditional object pose prediction methods (such as PoseCNN and NOCS) are limited by the diversity of training data and cannot predict the poses of unseen object categories; moreover, the acquisition and annotation of three-dimensional data are time-consuming and laborious, while the construction of two-dimensional data is more convenient and fast. Therefore, there is an urgent need to build a technical framework that uses two-dimensional models to solve three-dimensional tasks, reduce the cost of data acquisition, and improve the generality of the technology;

[0005] 2) General semantic models (such as DINO and Stable Diffusion) are pre-trained on a large amount of easy-to-collect data. The general semantic features extracted by them focus more on the structural information of objects, thereby supporting the establishment of semantic associations between unseen object categories and realizing posture mapping between objects. However, existing methods based on general semantic models do not consider the characteristic problem of applying general semantic features to object posture prediction tasks. In view of the problem of diverse object types and complex textures, it is necessary to consider the object features that different general semantic models focus on when establishing object semantic similarity, and construct general features with strong generalization ability to establish mapping relationships between objects;

[0006] 3) At present, general semantic features are greatly affected by the posture of objects. Drastic changes in the posture of objects will lead to feature degradation and make it impossible to establish an accurate mapping relationship between objects, thus affecting the accuracy of posture prediction. At the same time, due to the large differences in the geometric shapes of objects, direct posture mapping based on the association relationship between objects will cause posture ambiguity problems, and fall into the local optimal solution during the posture optimization process, affecting the accuracy of object posture prediction.

[0007] In summary, object pose prediction is a common problem in many mainstream applications, and a zero-shot category-level object pose model with high precision and strong generalization ability is its core issue. Summary of the invention

[0008] The purpose of the present invention is to provide a zero-sample category-level object pose estimation solution based on universal semantic features. Taking a single-frame RGBD image as input, a universal semantic feature combination method suitable for object pose prediction is constructed, and a two-stage iterative optimization strategy based on multimodal universal semantic features is designed. This can solve the problem of low pose prediction accuracy caused by large changes in object pose and obvious shape differences.

[0009] The technical solution adopted by the present invention to achieve the above-mentioned purpose is as follows:

[0010] A zero-sample category-level object pose estimation method comprises the following steps:

[0011] 1) Data preprocessing: preprocess the original color image and depth image of the target object to obtain the target color image and target depth image; perform perspective rendering according to the mesh model of the reference object to obtain the reference color image and reference depth image;

[0012] 2) Coarse pose estimation: Extract corresponding point pairs between the target color image and the reference color image, and extract the depth information of the corresponding point pairs from the target depth image and the reference depth image; Map the corresponding point pairs to the camera coordinate system according to the depth information and the camera intrinsics to obtain the key point cloud of the target object and the key point cloud of the reference object, and calculate the object pose transformation relationship from the key point cloud of the reference object to the key point cloud of the target object; Through iterative calculation, obtain the roughly estimated object pose transformation relationship, which consists of the rotation matrix, translation matrix and scaling size of the object;

[0013] 3) Pose optimization: Optimize the roughly estimated object pose transformation relationship by calculating the total loss function, and obtain the accurate pose prediction result of the final target object according to the optimized object pose transformation relationship.

[0014] Further, the steps of preprocessing the original color image of the target object in step 1) include: Using the pre-trained model Mask R-CNN to perform class detection and pixel segmentation on the original color image of the target object, and extracting the target mask; Calculate the bounding box of the object based on the target mask image, and perform foreground pixel extraction, cropping, resizing and padding on the original color image to obtain the target color image;

[0015] The steps of preprocessing the original depth image of the target object include: Using the target mask image to perform foreground pixel extraction, cropping, resizing and padding on the original depth image of the target object to obtain the target depth image.

[0016] Further, the steps of extracting the corresponding point pairs between the target color image and the reference color image in step 2) include:

[0017] Extract the 2D general semantic features of the target color image and the reference color image to obtain the target feature map and the reference feature map;

[0018] Calculate the cosine similarity between the pixels of the target feature map and the reference feature map to obtain the similarity score matrix;

[0019] Based on the similarity score matrix, select the pixel position q with the highest score of the target feature map pixel p on the reference feature map, then select the pixel p' with the highest score of q on the target feature map, and calculate the L2 distance between p and p' in the two-dimensional plane to obtain the cyclic distance matrix;

[0020] Sort the cyclic distance matrix D in ascending order, and select the first M values as the corresponding point pairs from the target color image to the reference color image.

[0021] Further, in step 2), both the target color image and the reference color image are used to extract 2D general semantic features through the general semantic models DINOv2 and Stable Diffusion.

[0022] Further, the step of calculating the object pose transformation relationship from the key point cloud of the reference object to the key point cloud of the target object in step 2) includes: calculating the object pose transformation relationship from the key point cloud of the reference object to the key point cloud of the target object by using the Umeyama algorithm, and excluding the influence of outliers by using the RANSAC algorithm. The object pose transformation relationship consists of the rotation matrix, translation matrix and scaling size of the object.

[0023] Further, the step of obtaining a roughly estimated object pose through iterative calculation in step 2) includes: rotating the reference object according to the calculated object pose transformation relationship by using the rotation matrix, and then rendering to obtain a new reference color image; extracting the corresponding point pairs between the target color image and the new reference color image, and extracting the depth information of the corresponding point pairs from the target depth image and the depth image corresponding to the new reference color image; mapping the corresponding point pairs to the camera coordinate system according to the depth information and the camera internal parameters to obtain the key point cloud of the target object and the key point cloud of the reference object, and calculating the object pose transformation relationship from the key point cloud of the reference object to the key point cloud of the target object to obtain a roughly estimated object pose transformation relationship.

[0024] Further, in step 1), four reference color images and corresponding four reference depth images are rendered from four perspectives of left, right, top-down and bottom-up according to the mesh model of the reference object; in step 2), four groups of corresponding point pairs between the target color image and the four reference color images are extracted and the pose of the reference object is optimized by an iterative method; when calculating the roughly object pose transformation relationship, calculate the average value of the cosine similarity of each group of corresponding point pairs from the target color image to the new reference color image, select the group of corresponding point pairs with the highest similarity, and calculate the roughly estimated object pose transformation relationship according to this group of corresponding point pairs.

[0025] Further, the total loss function in step 3) includes a pose optimization loss function and a regularization term loss function. The pose optimization loss function consists of a mask loss, a Chamfer loss and a general semantic feature alignment loss. The regularization term loss function consists of a pose regularization loss, a center point regularization loss and a deformation regularization loss.

[0026] Further, the general semantic feature alignment loss in step 3) is calculated based on the cosine similarity of the 3D general semantic features of the target object point cloud and the reference object point cloud. The 3D general semantic features are obtained by normalizing the target object point cloud and the reference object point cloud and then inputting them into the DGCNN network for extraction.

[0027] A zero-shot category-level object pose estimation device for implementing the above method, which includes:

[0028] A data preprocessing module, which is used to preprocess the original color image and depth image of the target object to obtain the target color image and target depth image; perform perspective rendering according to the mesh model of the reference object to obtain the reference color image and reference depth image;

[0029] A rough pose estimation module, which is used to extract the corresponding point pairs between the target color image and the reference color image, and extract the depth information of the corresponding point pairs from the target depth image and the reference depth image; map the corresponding point pairs to the camera coordinate system according to the depth information and the camera internal parameters to obtain the key point cloud of the target object and the key point cloud of the reference object, and calculate the object pose transformation relationship from the key point cloud of the reference object to the key point cloud of the target object; through iterative calculation, obtain the roughly estimated object pose transformation relationship, which consists of the rotation matrix, translation matrix and scaling size of the object;

[0030] A pose optimization module, which is used to optimize the roughly estimated object pose transformation relationship by calculating the total loss function, and obtain the accurate pose prediction result of the final target object according to the optimized object pose transformation relationship.

[0031] The beneficial effects achieved by the present invention are as follows:

[0032] 1. The present invention solves the problem of low accuracy of object pose prediction through a zero-shot category-level object pose estimation scheme guided by general semantic features, especially in the case of large differences in object shape and texture.

[0033] 2. The general semantic model adopted by the present invention effectively overcomes the problem of poor generalization ability of traditional non-general semantic models when facing unseen categories, thereby improving the prediction accuracy of object poses.

[0034] 3. The two-stage iterative optimization strategy of the present invention realizes the accurate estimation of object poses by establishing key point mapping between objects using 2D general semantic features in the rough pose estimation stage and further optimizing object shape and pose using 3D general semantic features in the pose optimization stage.

[0035] 4. The present invention can effectively solve the pose ambiguity problem caused by large differences in object shapes, and prevent the degradation of general semantic features through iterative optimization, further improving the accuracy of pose prediction.

[0036] 5. The present invention has a high degree of automation, can be widely applied to various professional and popular object pose estimation scenarios, has strong generalization ability, and can handle diverse object categories.

[0037] 6. The technical solution of the present invention has been verified in practice, proving that it has high precision, high automation and excellent generalization ability, and can meet different application requirements. Description of the Drawings

[0038] Figure 1 It is a schematic diagram of the application of the present invention.

[0039] Figure 2 It is the overall flowchart of a zero-shot category-level object pose estimation method in the embodiment.

[0040] Figure 3 It is a schematic diagram of pose optimization.

[0041] Figure 4 It is the pose prediction result of the present invention during actual testing.

[0042] Figure 5 It is the deformation result of the reference object in the pose optimization stage of the present invention. Specific Embodiments

[0043] To make the objectives, technical solutions, and advantages of the present invention clearer and more understandable, the following further elaborates on the present invention in detail with reference to specific embodiments and the accompanying drawings.

[0044] The embodiment of the present invention provides a zero-shot category-level object pose estimation method, and its principle is as Figure 1 shown. First, collect the model of common object categories in life from the public mesh model dataset Objverse, and select one mesh model as the reference object model for each category model; then, using the observed target object image and the reference object model as inputs, use the general semantic model to predict the three-dimensional pose of the target object. Among them, the general semantic model uses a pre-trained model and can achieve pose prediction for new object categories without retraining. The overall process of this method is as Figure 2 shown, and specifically includes the following three parts:

[0045] I. Data Preprocessing

[0046] 1. Target Image Preprocessing

[0047] Process the original color image and depth image data of the target object into the required format to obtain the target color image and the target depth image, that is, the target RGBD image. Among them, for the original color image, use the pre-trained model Mask R-CNN to perform category detection and pixel segmentation on the objects in the scene, take the foreground pixels as 1 and the background pixels as 0 to obtain the target mask image (that is, the mask image that only retains the foreground target object); then, based on the target mask image, calculate the bounding box of the object in the image, and crop, resize, and fill the image based on the bounding box to obtain a target color image with a resolution of 256×256. Perform the same cropping, resizing, and filling processing on the original depth image to obtain the target color image.

[0048] 2. Preprocessing of Reference Images

[0049] For common reference object categories, including bottles, bowls, cameras, cans, laptops, and mugs, a mesh model is selected from the Objaverse database as a reference object model respectively to form a local mesh model library. Then, based on Pytorch3D, each reference object model is rendered from four perspectives: left, right, top-down, and bottom-up to obtain a reference color image and a reference depth image with a resolution of 256×256, that is, a reference RGBD image, where the reference mask image can be obtained from the pixels with a depth greater than 0 in the reference depth image.

[0050] II. Coarse Pose Estimation

[0051] In the coarse pose estimation stage, the preprocessed target RGBD image and reference RGBD image are used as inputs. Based on the features extracted by the general semantic model, corresponding points between the reference color image and the target color image are constructed and the key point cloud is obtained by back-projection through the camera intrinsics and the depth image. Then, the Umeyama algorithm and the RANSAC algorithm are used to calculate the transformation matrix from the reference object to the target object, so as to obtain the coarse pose estimation result. The steps of the coarse pose estimation are as follows:

[0052] 1. Calculation of Corresponding Point Pairs

[0053] 1.1 Input the target color image I t and the reference color image I r into the general semantic models DINOv2 and StableDiffusion to extract 2D general semantic features, obtaining the target feature map F(I t ) and the reference feature map F(I r ), where F(·) represents the extraction of 2D general semantic features, and the specific composition includes:

[0054] F=(λ D2 ‖F DD2 ‖2,λ SD ‖F SD ‖2)

[0055] Among them, F D2 represents the features extracted by DINOv2, F SD represents the features extracted by Stable Diffusion, ‖·‖2 represents the normalization of the features, and λ D2 and λ SD are hyperparameters, which are set to 0.7 and 0.3 respectively.

[0056] 1.2 Calculate the cosine similarity between the pixels of the two feature images to obtain the similarity score matrix S, and the calculation formula is as follows:

[0057] S(p, q) = d cos (F(I t ) p , F(I r ) q , p ∈ [1, N t , q ∈ [1, N r

[0058] Among them, N t and N r respectively represent the number of pixels of the target feature map and the reference feature map, d cos (·, ·) represents the cosine similarity calculation function. The DINOv2 model selects the "vits14" version, and the Stable Diffusion model selects the "v1-5" version.

[0059] 1.3 Select the pixel position q with the highest score on the reference feature map for the pixel p of the target feature map based on the similarity score matrix S, then select the pixel p' with the highest score of q on the target feature map, and calculate the L2 distance between p and p' in the two-dimensional plane to obtain the cyclic distance matrix D. Its calculation formula is as follows:

[0060]

[0061] Among them, d(·, ·) represents the L2 distance calculation function.

[0062] 1.4 Sort the cyclic distance matrix D in ascending order, and then select the first M values as the corresponding point pairs from the target color image to the reference color image. Each target color image establishes corresponding point pairs with four reference color images, for a total of four groups of corresponding point pairs.

[0063] 2. Object pose transformation relationship calculation

[0064] Extract the depth values of the corresponding point pairs on the target depth image and the reference depth image, and then use the camera intrinsics to map the corresponding point pairs to the camera coordinate system to obtain the target object key point cloud P t and the reference object key point cloud P r . Based on the three-dimensional corresponding point pairs composed of the target object key point cloud P t and the reference object key point cloud P r , use the Umeyama algorithm to calculate the pose transformation relationship from the reference object to the target object, and use the RANSAC algorithm to exclude the influence of outliers. The calculation formula is as follows:

[0065]

[0066] Among them, represents the object pose transformation relationship, ​respectively represent the rotation matrix, translation matrix, and scaling size of the object. The scaling size is unified into the rotation matrix, and finally the rotation matrix of the rough pose estimation can be abbreviated as

[0067] 3. Iteratively calculate the object pose

[0068] Through the above operations, the first iteration of the rough pose estimation is completed. Then, after rotating the reference object using the rotation matrix after the first iteration, a new reference color image is obtained by rendering it Then, repeat step 1 to obtain four groups of corresponding point pairs, calculate the average cosine similarity of each group of corresponding point pairs from the target color image to the new reference color image, select the group of corresponding point pairs with the highest similarity, and input it into step 2 to calculate the object pose in the rough pose estimation stage

[0069] III. Pose optimization

[0070] In the pose optimization stage, according to the mask image and depth image of the target object, the mesh model, mask image, and depth image of the reference object, and the object pose transformation relationship obtained in the rough pose estimation stage, the pose and shape of the reference object are jointly optimized using 3D general semantic features and the point cloud information of the object, and the object pose is optimized. The optimized object pose is used as the final pose prediction result, as Figure 3 shown. The pose optimization stage can solve the problem of the decrease in pose accuracy caused by the shape difference between the target object and the reference object. The steps of pose optimization are as follows

[0071] 1. Based on the object pose transformation relationship obtained in the rough pose estimation stage Perform a pose transformation on the reference object, and then normalize the point cloud of the transformed reference object and input it into the DGCNN network to extract 3D general semantic features; similarly, normalize the point cloud of the target object and input it into the DGCNN network to extract 3D general semantic features

[0072] 2. Define the optimization parameters in the pose optimization stage, including the rotation matrix ΔR, translation matrix ΔT, scaling size Δs, and displacement amount ΔV of each vertex of the reference object point cloud. Then the final pose of the object can be expressed as

[0073]

[0074] where represents the rotation of the object represents the translation of the object, and the shape of the reference object in the model coordinate system can be calculated by the following formula

[0075]

[0076] Position V of the final reference object f can be expressed as:

[0077]

[0078] 3. Define the total loss function to optimize the parameters in the second step. During the optimization process, use the pose optimization loss function L p and the regularization term loss function L r . The total loss function is defined as:

[0079] L = L p + L r

[0080] 3.1 Pose optimization loss function

[0081] The pose optimization loss function L p aims to optimize the shape of the reference object to be as similar as possible to the target object. The loss functions used include: mask loss L m , Chamfer loss L c and general semantic feature alignment loss L g . The pose optimization loss function L p can be expressed as:

[0082] L p = λ m L m + λ c L c + λ g L g

[0083] where λ m , λ c , λ g are hyperparameters, which are set to 1, 0.1, 1 respectively.

[0084] (1) Mask loss L m is used to evaluate the difference between the reference mask image M r and the target mask image M t . Its calculation formula is:

[0085]

[0086] (2) Chamfer loss L c is used to evaluate the difference between the reference object point cloud and the target object point cloud. Its calculation formula is:

[0087]

[0088] where, represents the position of the reference object point cloud in the camera coordinate system, and N represents the number of points in the reference object point cloud. represents the distance of the target object point cloud to the position of the closest point.

[0089] (3) General semantic feature alignment loss L g is used to constrain that the positions of the reference object point cloud and the target object point cloud for the point clouds with high 3D general semantic feature similarity also need to be as close as possible. Its calculation formula is:

[0090]

[0091] where represent the reference object point cloud and the target object point cloud respectively, represents the cosine similarity calculation function based on 3D general semantic features between two points, and N g represents satisfying the number of point pairs.

[0092] 3.2 Regularization term loss function

[0093] Regularization term loss function L r aims to constrain that the change amount of the optimization parameters in the optimization process is as small as possible. The loss functions used include: pose regularization loss L ps , center point regularization loss L ce and deformation regularization loss L d . The regularization term loss function L r can be expressed as:

[0094] L r = λ ps L ps + λ ce L ce + λ d L d

[0095] where λ ps , λ ce , λ d are hyperparameters, which are set to 20, 1, 1 respectively.

[0096] (1) Pose regularization loss L ps is used to constrain that the optimized pose is as close as possible to the pose obtained in the rough pose estimation stage. Its calculation formula is as follows:

[0097]

[0098] where ‖·‖ 2 represents the MSE loss, and N vIndicates the number of vertices of the reference object model.

[0099] (2) Central point regularization loss L ce It is used to constrain the central point of the optimized reference object point cloud to be as close as possible to the central point before optimization. Its calculation formula is as follows:

[0100]

[0101] Where mean(·) represents the point cloud mean calculation function.

[0102] (3) Deformation regularization loss L d It is used to optimize the deformation of the reference object vertices to be as small as possible. Its calculation formula is as follows:

[0103]

[0104] At the same time, during the optimization process, referring to the optimization scheme of Pytorch3D, normal constraint, edge constraint and Laplacian constraint are added.

[0105] In the pose optimization stage, the parameters related to the object pose are optimized based on the total loss function L, and a total of 80 iterations are performed to obtain the final pose of the target object.

[0106] Figure 4 It is the pose prediction result of the present invention during actual testing. When tested on the REAL275 and Wild6D datasets, the first row is the target object to be detected, and the second row is the result of the present invention. It can be seen from this figure that the prediction result of this patent has a very small gap with the target pose and a very high accuracy.

[0107] Figure 5 Shows the deformation result of the reference object in the pose optimization stage. The reference object in the first column will optimize its own shape and pose simultaneously during the pose optimization stage, making its shape approach the target object in the second column, and solving the pose deviation caused by the shape influence. The third column shows the deformation process of the reference object. Among them, the shape of the "cup" in the first row will expand as a whole, and the fuselage of the "camera" in the second row will narrow as a whole.

[0108] The embodiment of the present invention also provides a zero-shot category-level object pose estimation device, which includes:

[0109] A data preprocessing module, which is responsible for preprocessing the data of the target object and the reference object, including obtaining foreground pixels of the image, rendering color images and depth images.

[0110] The rough pose estimation module, according to the input information of the target object and the reference object, establishes corresponding points between objects based on 2D general semantic features, and then uses Umeyama and RANSAC to calculate the pose mapping between objects; combined with an iterative optimization strategy, it solves the feature degradation caused by pose differences.

[0111] The pose optimization module uses the object pose obtained by the rough pose estimation module as the initial state, and combines the geometric information provided by the point cloud of the 3D general semantic features to optimize the shape and pose of the reference object.

[0112] In another embodiment, the general semantic features DINOv2, Stable Diffusion, and DGCNN of the present invention can be modified or equivalently replaced.

[0113] In another embodiment, the present invention further provides an electronic device (such as a computer, a server, etc.), which includes a memory and a processor. The memory stores a computer program, and the computer program is configured to be executed by the processor. The computer program includes instructions for executing the steps in the above-mentioned method.

[0114] In another embodiment, the present invention further provides a computer-readable storage medium (such as ROM / RAM, a disk, an optical disc), and the computer-readable storage medium stores a computer program. When the computer program is executed by a computer, the steps of the above-mentioned method are implemented.

[0115] The above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them. Those of ordinary skill in the art can modify or equivalently replace the technical solutions of the present invention without departing from the principles and scope of the present invention. The protection scope of the present invention shall be subject to what is described in the claims.

Claims

1. A zero-shot class-level object pose estimation method, characterized in that, It includes the following steps: 1) Preprocess the original color image and depth image of the target object to obtain the target color image and target depth image; Perform perspective rendering based on the mesh model of the reference object to obtain the reference color image and reference depth image; 2) Extract the corresponding point pairs of the target color image and the reference color image, and extract the depth information of the corresponding point pairs from the target depth image and the reference depth image; Map the corresponding point pairs to the camera coordinate system according to the depth information and the camera internal parameters to obtain the key point cloud of the target object and the key point cloud of the reference object, and calculate the object pose transformation relationship from the key point cloud of the reference object to the key point cloud of the target object; Through iterative calculation, obtain a roughly estimated object pose transformation relationship, which consists of the rotation matrix, translation matrix and scaling size of the object; 3) Optimize the roughly estimated object pose transformation relationship by calculating the total loss function, and obtain the accurate pose prediction result of the final target object according to the optimized object pose transformation relationship.

2. The method according to claim 1, characterized in that The steps for preprocessing the original color image of the target object in step 1) include: using the pre-trained model Mask R-CNN to perform class detection and pixel segmentation on the original color image of the target object, and extracting the target mask; calculating the bounding box of the object based on the target mask image, and performing foreground pixel extraction, cropping, resizing and padding on the original color image to obtain the target color image; The steps for preprocessing the original depth image of the target object include: using the target mask image to perform foreground pixel extraction, cropping, resizing and padding on the original depth image of the target object to obtain the target depth image.

3. The method according to claim 1, characterized in that, The steps for extracting the corresponding point pairs of the target color image and the reference color image in step 2) include: Extract the 2D general semantic features of the target color image and the reference color image to obtain the target feature map and the reference feature map; Calculate the cosine similarity between the pixels of the target feature map and the reference feature map to obtain the similarity score matrix; Based on the similarity score matrix, select the pixel position q with the highest score of the target feature image pixel p on the reference feature map, then select the pixel p' with the highest score of q on the target feature map, and calculate the L2 distance between p and p' in the two-dimensional plane to obtain the cyclic distance matrix; Arrange the cyclic distance matrix D in ascending order, and select the first M values as the corresponding point pairs from the target color image to the reference color image.

4. The method according to claim 3, wherein In step 2), both the target color image and the reference color image are used to extract 2D general semantic features through the general semantic model DINOv2 and Stable Diffusion.

5. The method according to claim 1, wherein The steps for calculating the object pose transformation relationship from the key point cloud of the reference object to the key point cloud of the target object in step 2) include: using the Umeyama algorithm to calculate the object pose transformation relationship from the key point cloud of the reference object to the key point cloud of the target object, and using the RANSAC algorithm to exclude the influence of outliers. The object pose transformation relationship consists of the rotation matrix, translation matrix and scaling size of the object.

6. The method according to claim 1, wherein The steps of obtaining a roughly estimated object pose through iterative calculation in step 2) include: according to the calculated object pose transformation relationship, rotating the reference object using the rotation matrix, and then obtaining a new reference color image after rendering; extracting the corresponding point pairs between the target color image and the new reference color image, and extracting the depth information of the corresponding point pairs from the target depth image and the depth image corresponding to the new reference color image; mapping the corresponding point pairs to the camera coordinate system according to the depth information and the camera internal parameters to obtain the target object key point cloud and the reference object key point cloud, and calculating the object pose transformation relationship from the reference object key point cloud to the target object key point cloud to obtain a roughly estimated object pose transformation relationship.

7. The method according to claim 1, wherein In step 1), four reference color images and corresponding four reference depth images are obtained by rendering the mesh model of the reference object from four perspectives: left, right, top-down, and top-up; in step 2), four groups of corresponding point pairs between the target color image and the four reference color images are extracted and the pose of the reference object is optimized by an iterative method. When calculating the roughly estimated object pose transformation relationship, calculate the average cosine similarity of each group of corresponding point pairs from the target color image to the new reference color image, select the group of corresponding point pairs with the highest similarity, and calculate the roughly estimated object pose transformation relationship according to this group of corresponding point pairs.

8. The method according to claim 1, wherein In step 3), the total loss function includes a pose optimization loss function and a regularization term loss function. The pose optimization loss function is composed of a mask loss, a Chamfer loss, and a general semantic feature alignment loss, and the regularization term loss function is composed of a pose regularization loss, a center point regularization loss, and a deformation regularization loss.

9. The method according to claim 8, wherein In step 3), the general semantic feature alignment loss is calculated based on the cosine similarity of the 3D general semantic features of the target object point cloud and the reference object point cloud. The 3D general semantic features are obtained by normalizing the target object point cloud and the reference object point cloud and then inputting them into the DGCNN network for extraction.

10. A zero-shot category-level object pose estimation device for implementing the method according to any one of claims 1-9, characterized in that, Including: A data preprocessing module for preprocessing the original color image and depth image of the target object to obtain a target color image and a target depth image; Rendering the perspective according to the mesh model of the reference object to obtain a reference color image and a reference depth image; A rough pose estimation module for extracting the corresponding point pairs between the target color image and the reference color image, and extracting the depth information of the corresponding point pairs from the target depth image and the reference depth image; Mapping the corresponding point pairs to the camera coordinate system according to the depth information and the camera internal parameters to obtain the target object key point cloud and the reference object key point cloud, and calculating the object pose transformation relationship from the reference object key point cloud to the target object key point cloud; Through iterative calculation, obtain a roughly estimated object pose transformation relationship, which consists of the rotation matrix, translation matrix, and scaling size of the object; A pose optimization module for optimizing the roughly estimated object pose transformation relationship by calculating the total loss function, and obtaining the accurate pose prediction result of the final target object according to the optimized object pose transformation relationship.