Global high-order pooling 6D object attitude estimation method based on deep learning
By introducing global high-order enhancement modules and attention mechanisms into deep learning networks, the problems of accurate CAD model dependence and environmental sensitivity in class-level 6D object pose estimation are solved, and a more accurate and robust pose estimation effect is achieved.
Patent Information
- Application Number
- CN202411533740.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-10-31
- Publication Date
- 2025-05-27
AI Technical Summary
The prior art has problems with accurate CAD model dependence, ambient lighting and complex background sensitivity, and high computational load in class-level 6D object pose estimation.
A global high-order pooled 6D object pose estimation method based on deep learning is proposed. By introducing a global high-order enhancement module, it is integrated in multiple stages of the network, using high-order statistical information to enhance feature representation, and pose estimation is realized through attention mechanism and multi-task learning.
This method can more accurately and reliably estimate the position of an object, improve the robustness and generalization ability of the model, accurately capture the pose changes of the target object in complex environments, and supports an end-to-end training process.
Smart Images

Figure FDA0005111039950000011 
Figure FDA0005111039950000012 
Figure FDA0005111039950000021
Abstract
Description
Technical Field
[0001] The present invention relates to the pose measurement technology at the intersection of the computer and mechanical fields. Specifically, it provides a novel category-level 6D object pose estimation method. By introducing a global high-order enhancement module, this method effectively captures and fuses the high-order geometric features of the object to achieve accurate prediction of the object's rotation, translation, and size. This technology has broad application prospects in the fields of mechanical automation and robot vision, can significantly improve the robot's ability to recognize and operate objects in complex environments, and thus promote the development of intelligent manufacturing and automation technologies. Through in-depth research and experimental verification, the present invention has demonstrated its superior performance in improving the accuracy and robustness of pose measurement, providing strong support for the technological progress in the intersection field of machinery and computer vision. Background Art
[0002] In the fields of modern industrial manufacturing and automation, precise pose measurement technology is the key to realizing efficient and accurate robot operations and automated production lines. Pose measurement, that is, determining the three-dimensional position and orientation of an object in space, is crucial for multiple applications such as robot guidance, automated assembly, quality control, and augmented reality. With the advancement of Industry 4.0, higher requirements are put forward for the flexibility and intelligence level of automated systems, which makes the research and development of pose measurement technology more urgent.
[0003] Traditional pose measurement methods, such as laser scanning or stereo vision-based systems, although can provide accurate measurement results in some cases, they are usually limited by the measurement environment, cost, and processing speed. For example, laser scanning may be affected by changes in reflectivity, while stereo vision systems may be sensitive to lighting conditions. In addition, the performance of these methods often degrades when dealing with complex backgrounds or dynamic scenes.
[0004] To overcome these limitations, researchers have begun to explore computer vision-based pose estimation methods. Early methods relied on detecting specific markers or feature points in images and then using these feature points to infer the pose of the object. However, these methods usually require the object to have obvious texture features or known geometric structures, which limits their application when dealing with textureless or irregularly shaped objects.
[0005] In recent years, the breakthroughs in deep learning technology have brought new possibilities to pose estimation. In particular, methods based on convolutional neural networks (CNNs) have achieved great success in fields such as image classification, object detection, and segmentation. These methods can handle more complex and diverse visual scenes by learning complex feature representations from large amounts of data. However, most of the existing deep learning-based pose estimation methods focus on the recognition and localization of single objects, and class-level pose estimation, that is, the generalization recognition and pose prediction of objects in different instances of the same class, remains an open problem.
[0006] Class-level pose estimation not only has to handle the deformations between objects and the variations within a class, but also has to be carried out without an accurate 3D model. This poses higher requirements for the generalization ability of the algorithm and the processing ability of high-dimensional data. In addition, when the existing deep learning-based methods extract geometric features, they often rely on global average pooling operations, which limits the network's utilization of the high-order statistical information of features, thus affecting the accuracy and robustness of pose estimation.
[0007] To address these challenges, the present invention proposes an innovative pose estimation network - HoPENet, which enhances the network's ability to capture high-order geometric features by introducing global high-order pooling technology. The design of HoPENet takes into account the high requirements for pose measurement accuracy and robustness in the mechanical field, and also makes full use of the advantages of computer vision in processing complex visual information. By integrating global high-order enhancement modules at multiple stages of the network, HoPENet can gradually learn and integrate high-order geometric information, thus achieving a more accurate and reliable estimation of the object's pose.
[0008] The technical background of the present invention is based on an in-depth analysis of existing pose measurement technologies and a comprehensive understanding of the needs of industrial automation. By identifying the limitations of existing technologies and integrating emerging technologies, the present invention aims to provide a new solution for pose measurement technologies at the intersection of the mechanical and computer fields to meet the requirements of future intelligent manufacturing for high-precision, high-efficiency, and high-robustness pose measurement. Summary of the Invention
[0009] To solve a series of problems existing in class-level 6D object pose estimation, such as the dependence on accurate CAD models, the sensitivity to environmental light and complex backgrounds, and the computational load when processing large-scale data, we propose a deep learning-based class-level 6D object pose estimation algorithm.
[0010] The technical solutions adopted by the present invention to solve the technical problems are as follows:
[0011] A global high-order pooling 6D object pose estimation method based on deep learning, which includes the following steps:
[0012] Step 1: Input data processing. The input data is the point cloud P, which is an N×3-dimensional matrix, where N represents the number of points, and each point has three coordinate values (x, y, z).
[0013] Step 2: Feature extraction. Use PointNet++ as the backbone network to process the input point cloud data P and extract geometric features F. These features are N×d-dimensional, where d represents the dimension of the features.
[0014] Step 3: Construct implicit shape priors. Use a set of queries Q as implicit shape priors. This set of queries is predefined for each object category, and the dimension of Q\ is the same as Q, i.e., N×d.
[0015] Step 4: Integration of global high-order enhancement modules. Integrate global high-order enhancement modules (GHoE modules) at each key stage of the network. These modules enhance the feature representation through high-order statistical information to capture complex relationships between features.
[0016] Step 5: Feature correlation and attention mechanism. Calculate the attention map A between the features F and the queries Q to select the features most relevant to the queries, and update the queries Q to Q′ accordingly.
[0017] Step 6: Feature similarity calculation and update. Calculate the similarity matrix M using the updated queries Q′ and the original features F, and further enhance the feature representation through sampling and residual connections.
[0018] Step 7: Feature fusion and further high-order enhancement. Fuse the sampled features M′ with the original features F and perform high-order enhancement through another GHoE module to obtain a richer feature representation.
[0019] Step 8: Pose estimation for multi-task learning. At the final stage of the network, use three independent multi-layer perceptrons (MLPs) to predict the rotation, translation, and size of the object respectively to achieve multi-task learning.
[0020] Step 9: Global feature integration. Integrate global features using global average pooling and attention mechanism, and perform final feature transformation through MLP to output accurate pose estimation results.
[0021] Step 10: Model training and optimization. Configure the model training process, including selecting the AdamW optimizer, setting the learning rate and weight decay, and determining the number of training epochs and batch size.
[0022] Step 11: Evaluation Metrics and Experimental Settings. According to the NOCS evaluation scheme, the mean average precision (mAP) and n°m cm at different IoU thresholds are used as evaluation criteria to evaluate the performance of the model.
[0023] The beneficial effects of the present invention are as follows:
[0024] 1). Deep Feature Learning and Multidimensional Data Fusion. HoPENet realizes deep feature learning of point cloud data through high-order pooling technology, and can extract richer and more abstract feature representations from the original data. The multidimensional data fusion strategy in the network not only integrates features from different point cloud regions, but also considers the mutual relationship between features, enhancing the model's understanding of complex scenes.
[0025] 2). Fine-grained Attention Mechanism and Feature Selection. By using a fine-grained attention allocation strategy, HoPENet can identify and strengthen the local features that are most critical for object pose estimation, while ignoring irrelevant background information. This mechanism greatly improves the sensitivity of the model to object features, enabling it to accurately capture the pose changes of target objects even in complex environments.
[0026] 3). Adaptive Learning Ability and Dynamic Adjustment. HoPENet has a strong adaptive learning ability and can dynamically adjust network parameters according to the characteristics of the input data to respond to different scenarios and objects in an optimal way. This adaptability enables HoPENet to maintain high-accuracy pose estimation even in the face of the diversity and uncertainty of data distributions.
[0027] 4). Multi-scale and Multi-view Feature Integration. The network can integrate feature information from different scales and views, providing a comprehensive geometric description of objects, so as to enable effective recognition and estimation under various observation conditions. This multi-scale feature processing ability makes HoPENet more flexible and accurate when dealing with objects of different sizes and orientations.
[0028] 5). Significant Improvement in Robustness and Generalization Ability. The design of HoPENet particularly emphasizes the robustness and generalization ability of the model, enabling it to work stably under various environmental conditions, even in the presence of occlusion, lighting changes, and object deformations. This robustness ensures the reliability of HoPENet in practical applications, especially in occasions that require high precision and high stability.
[0029] 6). Efficient process for end-to-end training and optimization. HoPENet supports an end-to-end training process. From data preprocessing to feature extraction, fusion, and finally pose estimation, the entire process can be efficiently completed within a unified framework. This end-to-end training method simplifies the model development and optimization process, improves R & D efficiency, and also enables the model parameters to be co-optimized to achieve the best performance.
[0030] 7). Model interpretability and visual analysis. HoPENet provides rich model interpretability. Through visualization tools, the decision-making process of the network can be intuitively displayed, including the attention distribution, feature selection, and the impact of high-order pooling. This visual analysis not only helps to understand the behavior of the model but also provides strong support for further model debugging and optimization. Description of the Drawings
[0031] As Figure 1 shown, it is the structure diagram of the high-order pose estimation network. This network includes components for feature extraction, high-order query, and pose estimation.
[0032] As Figure 2 shown, it is the structure of the global high-order enhancement module. Detailed Implementation Manner
[0033] The present invention will be further described in detail below with reference to the drawings.
[0034] The specific implementation steps are as follows:
[0035] Step 1: Input data processing. The input data is the point cloud P, which is an N×3-dimensional matrix, where N represents the number of points, and each point has three coordinate values (x, y, z).
[0036] Step 2: Feature extraction. Use PointNet++ as the backbone network to process the input point cloud data P and extract geometric features F. These features are N×d-dimensional, where d represents the dimension of the features.
[0037] Step 3: Construct an implicit shape prior. Use a set of queries Q as the implicit shape prior. This set of queries is predefined for each object category, and the dimension of Q is the same as Q, that is, N×d.
[0038] Step 4: Integration of the global high-order enhancement module. Integrate the global high-order enhancement module (GHoE module) at each key stage of the network. These modules enhance the feature representation through high-order statistical information to capture the complex relationships between features.
[0039] Step 5: Feature Association and Attention Mechanism. By calculating the attention map A between the feature F and the query Q, the feature most relevant to the query is selected, and the query Q is updated to Q′ accordingly. First, according to Equation 1, calculate the attention map
[0040] A i = Attn(F, Q) (1)
[0041] where A i , represents the i-th attention map, and N Q represents the number of queries. N P represents the number of points. By calculating the attention using Equation 2, the part with similar semantics to the query can be obtained. Therefore, the result of cross-attention is:
[0042] A = (A i F)W (2)
[0043] where
[0044] Step 6: Feature Similarity Calculation and Update. The updated Q′ is obtained according to Equation 3. Using the updated query Q′ and the original feature F, according to Equation 4, calculate the similarity matrix M, and further enhance the feature representation through sampling and residual connection. Finally, M′ is obtained according to Equation 5
[0045] Q′ = Q + A (3)
[0046] M = Norm(MLP)(FQ' T ) (4)
[0047] M′ = MQ′ + F (5)
[0048] where
[0049] Step 7: Feature Fusion and Further Higher-Order Enhancement. The sampled feature M′ is fused with the original feature F, and further higher-order enhancement is performed through another GHoE module to obtain a richer feature representation.
[0050] Step 8: Pose Estimation for Multi-Task Learning. At the final stage of the network, three independent multi-layer perceptrons (MLPs) are used to predict the rotation, translation, and size of the object respectively to achieve multi-task learning.
[0051] Step 9: Global Feature Integration. Use global average pooling and attention mechanism to integrate global features, and perform final feature transformation through MLP to output accurate pose estimation results.
[0052] Step Ten: Model Training and Optimization. Configure the model training process, including selecting the AdamW optimizer, setting the learning rate and weight decay, and determining the number of training epochs and batch size.
[0053] Step Eleven: Evaluation Metrics and Experimental Settings. According to the NOCS evaluation scheme, use the mean average precision (mAP) and n°m cm at different IoU thresholds as evaluation criteria to evaluate the performance of the model.
Claims
1. A global high-order pooling 6D object pose estimation method based on deep learning, which includes the following steps: Step 1: Input data processing. The input data is a point cloud P, which is an N×3-dimensional matrix, where N represents the number of points and each point has three coordinate values (x, y, z). Step 2: Feature extraction: Use PointNet++ as the backbone network to process the input point cloud data P and extract geometric features F, which are N×d-dimensional, where d represents the dimension of the feature. Step 3: Construct implicit shape prior. A set of queries Q is used as implicit shape prior. This set of queries is pre-defined for each object category, and the dimension of Q is the same as Q, that is, N×d. Step 4: Integration of global high-order enhancement modules. Global high-order enhancement modules (GHoE modules) are integrated at each key stage of the network. These modules enhance feature representation through high-order statistical information to capture the complex relationships between features. Step 5: Feature association and attention mechanism. By calculating the attention map A between feature F and query Q, we select the most relevant feature to the query and update the query Q to Q′ accordingly. First, according to formula 1, we calculate the attention map A i =Attn(F,Q) (1) in, represents the i-th attention map, N Q Represents the number of queries. N P The number of representative points. By calculating the attention through formula 2, the part with similar semantics to the query can be obtained. Therefore, the result of cross attention is: A=(A i F)W (2) in, Step 6: Calculate and update feature similarity. The updated Q′ is calculated according to Formula 3. Using the updated query Q′ and the original feature F, the similarity matrix M is calculated according to Formula 4, and the feature representation is further enhanced through sampling and residual connection. Finally, M′ is obtained according to Formula 5 Q′=Q+A (3) M=Norm(MLP)(FQ′ T ) (4) M′=MQ′+F (5) Where, Step 7: Feature fusion and further high-order enhancement: The sampled feature M′ is fused with the original feature F and further enhanced by another GHoE module to obtain a richer feature representation. Step 8: Posture estimation with multi-task learning. In the final stage of the network, three independent multi-layer perceptrons (MLPs) are used to predict the rotation, translation, and size of the object, respectively, to achieve multi-task learning. Step 9: Global feature integration. Global average pooling and attention mechanism are used to integrate global features, and the final feature transformation is performed through MLP to output accurate pose estimation results. Step 10: Model training and optimization. Configure the model training process, including selecting the AdamW optimizer, setting the learning rate and weight decay, and determining the training cycle and batch size. Step 11: Evaluation indicators and experimental settings. According to the NOCS evaluation scheme, the mean average precision (mAP) and n°m cm at different IoU thresholds are used as evaluation criteria to evaluate the performance of the model.