Class-level object pose estimation method and system
By adopting a combination technology of causal learning theory and knowledge distillation method in class-level object position estimation, the problem of extracting unified representations from complex intra-class features is solved, the estimation accuracy and generalization ability are improved, and the robustness is enhanced.
Patent Information
- Application Number
- CN202510171600.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-17
- Publication Date
- 2025-06-06
AI Technical Summary
In the prior art, it is difficult to extract unified representations from complex and changeable in-class features in class-level object position estimation, resulting in limited generalization capabilities and low accuracy.
The front door adjustment backbone network is constructed using causal learning theory, and the characteristic interaction between the target object and other similar objects is achieved through the self-attention layer and the cross-attention layer, and the class semantic information of the large-scale three-dimensional basic model is injected through the knowledge distillation method.
It improves the accuracy and generalization ability of class-level object position estimation, and enhances the robustness of the network to uncertain factors such as environmental noise, object deformation or occlusion.
Smart Images

Figure CN120107356A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the fields of deep learning and three-dimensional computer vision, and in particular to a method and system for estimating the pose of a category-level object. Background Art
[0002] With the rapid development of robotics and autonomous driving technologies, object pose estimation algorithms have gradually become a key research area. These algorithms are widely used in fields such as dexterous robot manipulation, autonomous driving, and augmented reality to perceive the position and pose of target objects in the environment. Accurate six-degree-of-freedom object pose estimation is crucial for achieving precise robotic manipulation and task execution in complex environments.
[0003] Traditional pose estimation algorithms usually rely on high-precision 3D CAD models of objects. However, the production and acquisition costs of 3D models are high, and the accuracy of pose estimation is largely restricted by the quality of the CAD model. In addition, such methods can usually only process trained object instances, resulting in limited generalization capabilities.
[0004] To solve these problems, category-level object pose estimation has become a new research trend. Category-level object pose estimation is a computer vision task that aims to estimate the 6-DOF pose of different instances in the same category, including 3D rotation, 3D translation, and 3D size. Unlike instance-level pose estimation, category-level methods do not require a predefined object model and only rely on a single RGB image as input. Through training, the model can predict the 6-DOF pose and size of unseen instances of the same object. This method significantly improves the category generalization ability of the pose estimation model by mining similar features of objects of the same category, thereby broadening its application scenarios. However, the main challenge facing category-level object pose estimation is how to extract a unified representation from complex and changeable intra-class features. Existing technologies only mine features from limited training instances, and these datasets often have inherent biases (such as uneven category distribution, scene similarity, etc.), which may cause the model to learn incorrect prediction logic, thereby reducing the accuracy of pose estimation. In addition, the generalization of the model is also limited by the size of the training set. Summary of the invention
[0005] The purpose of the present invention is to provide a method and system for class-level object pose estimation in order to overcome the defects of the above-mentioned prior art, thereby improving the accuracy of object pose estimation.
[0006] The purpose of the present invention can be achieved by the following technical solutions:
[0007] A method for class-level object pose estimation comprises the following steps:
[0008] Acquire a monocular image and a corresponding depth image of the target object, and obtain a three-dimensional point cloud of the target object according to the monocular image and the corresponding depth image;
[0009] Obtaining RGB features of the target object through a two-dimensional feature extractor according to the monocular image, and obtaining point cloud features of the target object according to the three-dimensional point cloud through a three-dimensional feature extractor;
[0010] The RGB features and the point cloud features are concatenated by channel as input to the object characterization point extraction network to obtain the characterization point features;
[0011] The representation point features and the sampling reference features are used as inputs of the front-gate adjustment backbone network to obtain cross-attention layer features and self-attention layer features respectively, the sampling reference features are obtained by random sampling in the feature queue, and the front-gate adjustment backbone network is constructed by the front-gate adjustment strategy of causal learning theory;
[0012] The self-attention layer features and the cross-attention layer features are weightedly fused by an adaptive weighted fusion algorithm to obtain a pose estimation output feature;
[0013] The pose estimation output features are input into the MLP network for decoding to obtain the six-degree-of-freedom pose and size of the target object.
[0014] Furthermore, the two-dimensional feature extractor adopts a Vision Transformer network and uses a basic model encoder with frozen parameters; the three-dimensional feature extractor adopts a PointNet++ neural network; and the object representation point extraction network adopts a Vision Transformer network to locate key points on the object surface through feature similarity.
[0015] Furthermore, the front-gate adjustment backbone network is trained through knowledge distillation, the point cloud features are average-pooled to obtain average features, the average features are input into the MLP network for processing and the knowledge distillation output features are obtained through residual connection, and the knowledge distillation output features are aligned with the corresponding three-dimensional features extracted from the point cloud features through a large-scale three-dimensional basic model, thereby supervising the training of the front-gate adjustment backbone network.
[0016] Furthermore, the average feature is:
[0017]
[0018] In the formula, is the average feature, linear is the linear operation, GELU is the activation function, is the intermediate feature, μ is the balance parameter, AvgPool is the average pooling operation, F PPoint cloud features.
[0019] Furthermore, the feature alignment improves feature similarity according to an L2 loss function, and the L2 loss function is:
[0020]
[0021] Where, L KD is the knowledge distillation loss, B is the training batch size, The corresponding 3D features extracted from the large-scale 3D basic model, is the average characteristic.
[0022] Furthermore, the cross attention layer feature is obtained by performing a cross attention operation on the representation point feature after passing through a linear layer and the sampling reference feature, and its calculation formula is:
[0023] F x =CrossAttention(linear(F kpt ),F samp )
[0024] In the formula, F samp is the sampling reference feature, F x is the cross attention layer feature, F kpt To represent point features, CrossAttention is the cross attention operation, and linear is the linear operation.
[0025] Furthermore, the self-attention layer feature is obtained by performing a self-attention operation on the representation point feature after passing through a linear layer, and its calculation formula is:
[0026] F m =SelfAttention(linear(F kpt ))
[0027] In the formula, F m is the self-attention layer feature, F kpt To represent point features, SelfAttention is the self-attention operation, and linear is the linear operation.
[0028] Furthermore, the output characteristics are:
[0029] F f =ω a ⊙(layernorm(F x +F m ))+(1-ω a )⊙F kpt
[0030] In the formula, ω ais the weighted fusion weight coefficient of the cross-attention layer feature and the self-attention layer feature, layernorm is the normalization operation, ⊙ is the matrix element-by-element multiplication, F x is the cross attention layer feature, F m is the self-attention layer feature, F kpt To represent point features.
[0031] Furthermore, the weighted fusion weight coefficient of the cross-attention layer feature and the self-attention layer feature is:
[0032] ω a =σ(layernorm(F x +F m )W 1 +F kpt W 2 )
[0033] In the formula, ω a is the weighted fusion weight coefficient of the cross-attention layer features and the self-attention layer features, σ is the S-shaped growth curve function, W 1 and W 2 are learnable parameters.
[0034] According to another aspect of the present invention, there is provided a category-level object pose estimation system, comprising:
[0035] A three-dimensional point cloud acquisition module is used to acquire a monocular image and a corresponding depth image of a target object, and obtain a three-dimensional point cloud of the target object according to the monocular image and the corresponding depth image;
[0036] An RGB feature and point cloud feature acquisition module, used to obtain the RGB features of the target object through a two-dimensional feature extractor according to the monocular image, and to obtain the point cloud features of the target object according to the three-dimensional point cloud through a three-dimensional feature extractor;
[0037] A characterization point feature acquisition module is used to splice the RGB features and point cloud features by channel as input to the object characterization point extraction network to obtain characterization point features;
[0038] A cross-attention layer feature and self-attention layer feature acquisition module, used to use the representation point feature and the sampling reference feature as inputs of the front-gate adjustment backbone network to obtain the cross-attention layer feature and the self-attention layer feature respectively, wherein the sampling reference feature is obtained by random sampling in the feature queue, and the front-gate adjustment backbone network is constructed by the front-gate adjustment strategy of causal learning theory;
[0039] An output feature acquisition module is used to weightedly fuse the self-attention layer features and the cross-attention layer features through an adaptive weighted fusion algorithm to obtain a pose estimation output feature;
[0040] The posture acquisition module is used to input the posture estimation output features into the MLP network for decoding to obtain the six-degree-of-freedom posture and size of the object.
[0041] Compared with the prior art, the present invention has the following beneficial effects:
[0042] 1. Based on the causal learning theory, the present invention introduces the front-gate adjustment strategy, constructs the front-gate adjustment backbone network, realizes the feature interaction between the target object and other objects of the same type through the self-attention layer and the cross-attention layer, and realizes the extraction of a unified representation from the complex and changeable intra-class features, thereby ensuring the logical guidance in the data fitting process and effectively improving the accuracy of category-level object pose estimation.
[0043] 2. The feature-based knowledge distillation method of the present invention utilizes a large-scale three-dimensional model to supervise the training of the front-gate adjusted backbone network, injects the rich category semantic information in the large-scale three-dimensional base model into the object feature embedding, and further enhances the category-level generalization ability of the network.
[0044] 3. The present invention performs weighted fusion of self-attention layer features and cross-attention layer features through an adaptive weight fusion algorithm, which can be adaptively adjusted in a variety of scenarios and conditions, thereby improving the robustness of the network to uncertain factors such as environmental noise, object deformation or occlusion. BRIEF DESCRIPTION OF THE DRAWINGS
[0045] Figure 1 A schematic diagram of the flow chart of the method for estimating the pose of a category-level object proposed in the present invention;
[0046] Figure 2 This is a schematic diagram of the knowledge distillation process;
[0047] Figure 3 Flowchart of adjusting backbone network for front door;
[0048] Figure 4 This is a schematic diagram of the structure of the category-level object pose estimation system proposed by the present invention;
[0049] Legend: 1. 3D point cloud acquisition module; 2. RGB feature and point cloud feature acquisition module; 3. Representation point feature acquisition module; 4. Cross-attention layer feature and self-attention layer feature acquisition module; 5. Output feature acquisition module; 6. Posture acquisition module. DETAILED DESCRIPTION
[0050] The present invention is described in detail below in conjunction with the accompanying drawings and specific embodiments. This embodiment is implemented based on the technical solution of the present invention, and provides a detailed implementation method and specific operation process, but the protection scope of the present invention is not limited to the following embodiments.
[0051] Abbreviations involved:
[0052] Multilayer Perceptron: Multilayer Perceptron, MLP
[0053] Example 1
[0054] This embodiment provides a method for estimating the pose of a class-level object. Figure 1 As shown, the following steps are included:
[0055] S1. Obtain a monocular image and a corresponding depth image of a target object, and obtain a three-dimensional point cloud of the target object according to the monocular image and the corresponding depth image.
[0056] S2. Obtain RGB features of the target object through a two-dimensional feature extractor according to the monocular image, and obtain point cloud features of the target object according to the three-dimensional point cloud through a three-dimensional feature extractor.
[0057] The 2D feature extractor uses the Vision Transformer network and uses a base model encoder with frozen parameters; the 3D feature extractor uses the PointNet++ neural network; the object representation point extraction network uses the Vision Transformer network to locate the key points on the object surface through feature similarity. The above 2D features and 3D features are aligned on the dimension channel.
[0058] S3. Concatenate the RGB features and point cloud features by channel as the input of the object representation point extraction network to obtain the representation point features.
[0059] S4. Use the representation point features and sampling reference features as the input of the front-gate adjusted backbone network to obtain the cross-attention layer features and self-attention layer features respectively.
[0060] Through knowledge distillation training, the front gate is adjusted to adjust the backbone network, such as Figure 2 As shown in the figure, the point cloud features are averaged and pooled to obtain the average features, which are then input into the MLP network for processing and the knowledge distillation output features are obtained through residual connection. The knowledge distillation output features are aligned with the corresponding 3D features extracted from the point cloud features through the large-scale 3D base model, thereby supervising the training of the front-door adjustment backbone network. The 3D base model uses the point cloud Transformer network architecture and freezes the backbone network parameters.
[0061] After average pooling, the point cloud features are passed through an MLP network consisting of two linear layers and one GELU activation layer to obtain the average feature, which is:
[0062]
[0063] In the formula, is the average feature, linear is the linear operation, GELU is the activation function, is the intermediate feature, μ is the balance parameter, AvgPool is the average pooling operation, F P Point cloud features.
[0064] Feature alignment improves feature similarity according to the L2 loss function, which is:
[0065]
[0066] Where, L KD is the knowledge distillation loss, B is the training batch size, The corresponding 3D features extracted from the large-scale 3D basic model, is the average characteristic.
[0067] A large-scale three-dimensional basic model is used to extract features of several similar samples from the training set to form a feature queue, and the features in the queue are randomly sampled to obtain sampling reference features.
[0068] like Figure 3 As shown in the figure, the cross attention layer feature is obtained by performing a cross attention operation on the representation point feature after passing through a linear layer and the sampled reference feature. The calculation formula is:
[0069] F x =CrossAttention(linear(F kpt ),F samp )
[0070] In the formula, F samp is the sampling reference feature, F x is the cross attention layer feature, F kpt To represent point features, CrossAttention is the cross attention operation, and linear is the linear operation.
[0071] The self-attention layer feature is obtained by performing the self-attention operation after the representation point feature passes through a linear layer. The calculation formula is:
[0072] F m =SelfAttention(linear(F kpt ))
[0073] In the formula, F m is the self-attention layer feature, F kpt To represent point features, SelfAttention is the self-attention operation, and linear is the linear operation.
[0074] The randomly sampled reference feature shape is Ns ×C, where N s is the number of features after sampling, and C is the feature dimension
[0075] S5. The self-attention layer features and the cross-attention layer features are weightedly fused through an adaptive weighted fusion algorithm to obtain the pose estimation output features.
[0076] The output features of pose estimation are:
[0077] F f =ω a ⊙(layernorm(F x +F m ))+(1-ω a )⊙F kpt
[0078] In the formula, ω a is the weighted fusion weight coefficient of the cross-attention layer feature and the self-attention layer feature, layernorm is the normalization operation, ⊙ is the matrix element-by-element multiplication, F x is the cross attention layer feature, F m is the self-attention layer feature, F kpt To represent point features.
[0079] The weighted fusion weight coefficient of the cross-attention layer features and the self-attention layer features is:
[0080] ω a =σ(layernorm(F x +F m )W 1 +F kpt W 2 )
[0081] In the formula, ω a is the weighted fusion weight coefficient of the cross-attention layer features and the self-attention layer features, σ is the S-shaped growth curve function, W 1 and W 2 are learnable parameters.
[0082] S6. Input the output features into the MLP network for decoding to obtain the six-degree-of-freedom pose and size of the object.
[0083] Send the output features into the feature queue and update the features in the feature queue according to the first-in-first-out rule.
[0084] Example 2
[0085] This embodiment provides a category-level object pose estimation system, such as Figure 4 As shown, including:
[0086] The three-dimensional point cloud acquisition module 1 is used to acquire a monocular image and a corresponding depth image of the target object, and obtain a three-dimensional point cloud of the target object according to the monocular image and the corresponding depth image;
[0087] RGB feature and point cloud feature acquisition module 2, used to obtain the RGB features of the target object through a two-dimensional feature extractor according to the monocular image, and obtain the point cloud features of the target object according to the three-dimensional point cloud through a three-dimensional feature extractor;
[0088] The characterization point feature acquisition module 3 is used to splice the RGB features and the point cloud features by channel as the input of the object characterization point extraction network to obtain the characterization point features;
[0089] The cross-attention layer feature and self-attention layer feature acquisition module 4 is used to use the representation point feature and the sampling reference feature as the input of the front-door adjustment backbone network to obtain the cross-attention layer feature and the self-attention layer feature respectively. The sampling reference feature is obtained by random sampling in the feature queue. The front-door adjustment backbone network is constructed by the front-door adjustment strategy of causal learning theory.
[0090] The output feature acquisition module 5 is used to perform weighted fusion of the self-attention layer features and the cross-attention layer features through an adaptive weighted fusion algorithm to obtain the pose estimation output features;
[0091] The posture acquisition module 6 is used to input the posture estimation output features into the MLP network for decoding to obtain the six-degree-of-freedom posture and size of the object.
[0092] The rest is the same as in Example 1.
[0093] The preferred specific embodiments of the present invention are described in detail above. It should be understood that a person skilled in the art can make many modifications and changes based on the concept of the present invention without creative work. Therefore, any technical solution that can be obtained by a person skilled in the art through logical analysis, reasoning or limited experiments based on the concept of the present invention on the basis of the prior art should be within the scope of protection determined by the claims.
Claims
1. A method for class-level object pose estimation, characterized in that: The following steps are involved: Acquire a monocular image and a corresponding depth image of the target object, and obtain a three-dimensional point cloud of the target object according to the monocular image and the corresponding depth image; Obtaining RGB features of the target object through a two-dimensional feature extractor according to the monocular image, and obtaining point cloud features of the target object according to the three-dimensional point cloud through a three-dimensional feature extractor; The RGB features and the point cloud features are concatenated by channel as input to the object characterization point extraction network to obtain the characterization point features; The representation point features and the sampling reference features are used as inputs of the front-gate adjustment backbone network to obtain cross-attention layer features and self-attention layer features respectively, the sampling reference features are obtained by random sampling in the feature queue, and the front-gate adjustment backbone network is constructed by the front-gate adjustment strategy of causal learning theory; The self-attention layer features and the cross-attention layer features are weightedly fused by an adaptive weighted fusion algorithm to obtain a pose estimation output feature; The pose estimation output features are input into the MLP network for decoding to obtain the six-degree-of-freedom pose and size of the target object.
2. The method for class-level object pose estimation according to claim 1, characterized in that: The two-dimensional feature extractor adopts the Vision Transformer network and uses a basic model encoder with frozen parameters; the three-dimensional feature extractor adopts the PointNet++ neural network; the object representation point extraction network adopts the Vision Transformer network to locate the key points on the object surface through feature similarity.
3. The method for class-level object pose estimation according to claim 1, characterized in that: The front-door adjustment backbone network is trained through knowledge distillation, the point cloud features are average-pooled to obtain average features, the average features are input into the MLP network for processing and the knowledge distillation output features are obtained through residual connection, and the knowledge distillation output features are aligned with the corresponding three-dimensional features extracted from the point cloud features through a large-scale three-dimensional basic model, thereby supervising the training of the front-door adjustment backbone network.
4. The method for class-level object pose estimation according to claim 3, characterized in that: The average characteristics are: In the formula, is the average feature, linear is the linear operation, GELU is the activation function, is the intermediate feature, μ is the balance parameter, AvgPool is the average pooling operation, F P Point cloud features.
5. The method for class-level object pose estimation according to claim 3, characterized in that: The feature alignment improves feature similarity according to the L2 loss function, and the L2 loss function is: Where, L KD is the knowledge distillation loss, B is the training batch size, The corresponding 3D features extracted from the large-scale 3D basic model, is the average characteristic.
6. The method for class-level object pose estimation according to claim 1, characterized in that: The cross attention layer feature is obtained by performing a cross attention operation on the representation point feature after passing through a linear layer and the sampling reference feature. The calculation formula is: F x =CrossAttention(linear(F kpt ),F samp ) In the formula, F samp is the sampling reference feature, F x is the cross attention layer feature, F kpt To represent point features, CrossAttention is the cross attention operation, and linear is the linear operation.
7. The method for class-level object pose estimation according to claim 1, characterized in that: The self-attention layer feature is obtained by performing a self-attention operation on the representation point feature after passing through a linear layer, and its calculation formula is: F m =SelfAttention(linear(F kpt )) In the formula, F m is the self-attention layer feature, F kpt To represent point features, SelfAttention is the self-attention operation, and linear is the linear operation.
8. The method for class-level object pose estimation according to claim 1, characterized in that: The output features of the pose estimation are: F f =ω a ⊙(layernorm(F x +F m ))+(1-ω a )⊙F kpt In the formula, ω a is the weighted fusion weight coefficient of the cross-attention layer feature and the self-attention layer feature, layernorm is the normalization operation, ⊙ is the matrix element-by-element multiplication, F x is the cross attention layer feature, F m is the self-attention layer feature, F kpt To represent point features.
9. The method for class-level object pose estimation according to claim 8, characterized in that: The weighted fusion weight coefficient of the cross attention layer feature and the self-attention layer feature is: oh a =σ(layernorm(F x +F m )W1+F kpt W2) In the formula, ω a is the weighted fusion weight coefficient of the cross-attention layer features and the self-attention layer features, σ is the S-shaped growth curve function, and W1 and W2 are learnable parameters.
10. A category-level object pose estimation system, characterized in that: include: A three-dimensional point cloud acquisition module (1) is used to acquire a monocular image and a corresponding depth image of a target object, and obtain a three-dimensional point cloud of the target object according to the monocular image and the corresponding depth image; An RGB feature and point cloud feature acquisition module (2) is used to obtain the RGB features of the target object through a two-dimensional feature extractor according to the monocular image, and to obtain the point cloud features of the target object according to the three-dimensional point cloud through a three-dimensional feature extractor; A characterization point feature acquisition module (3), used for splicing the RGB features and the point cloud features by channel as input of an object characterization point extraction network to obtain characterization point features; A cross-attention layer feature and self-attention layer feature acquisition module (4) is used to use the representation point feature and the sampling reference feature as inputs of the front-gate adjustment backbone network to obtain the cross-attention layer feature and the self-attention layer feature respectively, wherein the sampling reference feature is obtained by random sampling in the feature queue, and the front-gate adjustment backbone network is constructed by the front-gate adjustment strategy of causal learning theory; An output feature acquisition module (5) is used to perform weighted fusion on the self-attention layer features and the cross-attention layer features through an adaptive weighted fusion algorithm to obtain a pose estimation output feature; The posture acquisition module (6) is used to input the posture estimation output features into the MLP network for decoding to obtain the six-degree-of-freedom posture and size of the object.
Citation Information
Patent Citations
Class-level 6D pose and size estimation method and device
CN113012122A
Vehicle door and vehicle body side accurate matching adjustment method based on scanning measurement
CN114492016A
Image title automatic generation method based on causal reasoning
CN115239944A
Class-level object 6D pose estimation method and system based on dynamic key point detection
CN117456003A
Cited By
Class-level object pose estimation method and system based on memory mechanism and medium
CN122223108A