Autoencoder Pose Estimator Training Heterogeneous Data
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current pose estimation methods for autonomous systems are not robust enough in absolute 3D space, especially under heavy occlusion, and require large, heterogeneous datasets with different key point definitions, making them resource-intensive and difficult to combine, which limits their applicability in scenarios like robotics and automotive applications.
Innovation Solution
A method involving the use of an autoencoder to map different pose formats to a common format, allowing a pose estimator to be trained on multiple datasets with varying formats, using a shared backbone and prediction heads, and incorporating latent key point representations for dimensionality reduction and affine combinations to enhance consistency and accuracy.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Quantity of substance
If multiple heterogeneous datasets with different pose formats are combined for training, then the training data volume increases, but the complexity of data processing and format unification increases
Solution Approach 1:
The patent introduces an autoencoder as an intermediary component that learns to map different pose formats from multiple datasets into a unified latent space representation. This mediator automatically handles format unification without requiring manual preprocessing or complex data processing pipelines, thus increasing training data volume while avoiding proportional increases in processing complexity
Solution Approach 2:
The patent transforms pose data from various formats by changing their parameter representation into a common latent space through the autoencoder. Different pose formats (with varying key point definitions and skeleton structures) are converted into a unified parameter space, enabling seamless combination of multiple datasets without format mismatch issues
2Adaptability or versatility
If different pose formats with heterogeneous key point definitions are used, then the adaptability to different datasets improves, but the consistency of pose estimation results deteriorates
Solution Approach 1:
The patent moves pose representation from the original heterogeneous format space into a new latent space dimension through the autoencoder. This dimensional transformation allows the system to maintain adaptability to various input formats while achieving consistency in the target latent space, where all pose estimates are expressed in a unified coordinate system with consistent key point definitions
Solution Approach 2:
The autoencoder serves as a universal translator that can process multiple different pose formats (humans, animals, vehicles, objects) and convert them into a common representation. This multi-functional component enables the system to handle diverse datasets while producing consistent output formats, resolving the trade-off between adaptability and consistency
3Measurement precision
If large datasets are used for training, then the accuracy of pose estimation improves, but the resource consumption for data collection and processing increases
Solution Approach 1:
The patent merges multiple existing pose estimation datasets into a single unified training set through the autoencoder's format unification capability. Instead of collecting and processing one extremely large dataset, the system combines several smaller datasets that would otherwise be incompatible, achieving large-scale training benefits with reduced individual data collection efforts and lower overall processing resources
Solution Approach 2:
The autoencoder learns to copy and transform pose information from various source formats into the unified latent space representation. This copying mechanism allows the system to leverage existing annotated datasets without requiring expensive new data collection, maintaining high accuracy while reducing resource consumption for data acquisition
Data Source
AI summary
Training a pose estimator. The pose estimator may receive as input a sensor measurement representing an object and to produce a pose of the object as output. Training the pose estimator may include applying multiple trained initial pose estimators to a pool of sensor measurements to obtain multiple estimated poses for a sensor measurement. A further pose estimator may be trained on multiple training data sets using at least part of an autoencoder trained on the multiple estimated poses to map a pose from a first pose format to a second pose format.


