Autoencoder Pose Estimator Training Heterogeneous Data

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current pose estimation methods for autonomous systems are not robust enough in absolute 3D space, especially under heavy occlusion, and require large, heterogeneous datasets with different key point definitions, making them resource-intensive and difficult to combine, which limits their applicability in scenarios like robotics and automotive applications.

Innovation Solution

A method involving the use of an autoencoder to map different pose formats to a common format, allowing a pose estimator to be trained on multiple datasets with varying formats, using a shared backbone and prediction heads, and incorporating latent key point representations for dimensionality reduction and affine combinations to enhance consistency and accuracy.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Quantity of substance

If multiple heterogeneous datasets with different pose formats are combined for training, then the training data volume increases, but the complexity of data processing and format unification increases

Engineering Contradiction:
Improvetraining data volumeVSAvoiddata processing complexity
Core Design Contradiction:
Quantity of substanceVSDevice complexity

Solution Approach 1:

The patent introduces an autoencoder as an intermediary component that learns to map different pose formats from multiple datasets into a unified latent space representation. This mediator automatically handles format unification without requiring manual preprocessing or complex data processing pipelines, thus increasing training data volume while avoiding proportional increases in processing complexity

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The patent transforms pose data from various formats by changing their parameter representation into a common latent space through the autoencoder. Different pose formats (with varying key point definitions and skeleton structures) are converted into a unified parameter space, enabling seamless combination of multiple datasets without format mismatch issues

Inventive Principle:
Principle #35Parameter changes

2Adaptability or versatility

If different pose formats with heterogeneous key point definitions are used, then the adaptability to different datasets improves, but the consistency of pose estimation results deteriorates

Engineering Contradiction:
Improvedataset compatibilityVSAvoidpose estimation consistency
Core Design Contradiction:
Adaptability or versatilityVSReliability

Solution Approach 1:

The patent moves pose representation from the original heterogeneous format space into a new latent space dimension through the autoencoder. This dimensional transformation allows the system to maintain adaptability to various input formats while achieving consistency in the target latent space, where all pose estimates are expressed in a unified coordinate system with consistent key point definitions

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

Solution Approach 2:

The autoencoder serves as a universal translator that can process multiple different pose formats (humans, animals, vehicles, objects) and convert them into a common representation. This multi-functional component enables the system to handle diverse datasets while producing consistent output formats, resolving the trade-off between adaptability and consistency

Inventive Principle:
Principle #6Universality (Multi-functionality)

3Measurement precision

If large datasets are used for training, then the accuracy of pose estimation improves, but the resource consumption for data collection and processing increases

Engineering Contradiction:
Improvepose estimation accuracyVSAvoidresource consumption
Core Design Contradiction:
Measurement precisionVSUse of energy by moving object

Solution Approach 1:

The patent merges multiple existing pose estimation datasets into a single unified training set through the autoencoder's format unification capability. Instead of collecting and processing one extremely large dataset, the system combines several smaller datasets that would otherwise be incompatible, achieving large-scale training benefits with reduced individual data collection efforts and lower overall processing resources

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The autoencoder learns to copy and transform pose information from various source formats into the unified latent space representation. This copying mechanism allows the system to leverage existing annotated datasets without requiring expensive new data collection, maintaining high accuracy while reducing resource consumption for data acquisition

Inventive Principle:
Principle #26Copying

Data Source

PatentUS20240212195A1Method for training a pose estimator
Publication Date: 2024.06.27 ROBERT BOSCH GMBH
  • US20240212195A1 patent drawing
  • US20240212195A1 patent drawing
  • US20240212195A1 patent drawing

AI summary

Training a pose estimator. The pose estimator may receive as input a sensor measurement representing an object and to produce a pose of the object as output. Training the pose estimator may include applying multiple trained initial pose estimators to a pool of sensor measurements to obtain multiple estimated poses for a sensor measurement. A further pose estimator may be trained on multiple training data sets using at least part of an autoencoder trained on the multiple estimated poses to map a pose from a first pose format to a second pose format.