Pose Estimation Model Synthetic Data Pretraining

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing pose estimation techniques using machine learning models face challenges in generating a large and diverse training dataset, which can lead to poor generalization and accuracy due to the need for extensive manual labeling and limited data coverage of human appearances, poses, and environments.

Innovation Solution

A technique that pretrains pose estimation models using synthetic data and further trains them with unlabeled real-world images, allowing the model to distinguish between left and right sides of objects and improve prediction accuracy.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If a large and diverse training dataset is collected manually, then pose estimation accuracy can be improved, but time consumption and resource requirements increase significantly

Engineering Contradiction:
Improvepose estimation accuracyVSAvoidtime consumption for data collection and labeling
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The patent uses synthetic data generation to create copies of training samples through rendering engines that simulate various poses, appearances, and environments. This allows the model to learn from numerous synthetic examples without manual data collection, significantly reducing time consumption while maintaining diversity and coverage for accurate pose estimation.

Inventive Principle:
Principle #26Copying

2Adaptability or versatility

If manually labeled training data is used, then model training can proceed, but the dataset lacks sufficient coverage of human appearances, poses, and environments

Engineering Contradiction:
Improvecoverage of human appearances, poses, and environmentsVSAvoidamount of training data
Core Design Contradiction:
Adaptability or versatilityVSQuantity of substance

Solution Approach 1:

The patent employs rendering engines that systematically vary parameters such as human appearance characteristics, pose configurations, and environmental conditions to generate diverse synthetic training data. This approach enables comprehensive coverage of the parameter space without requiring proportional increases in manual data collection efforts.

Inventive Principle:
Principle #35Parameter changes

3Productivity

If synthetic data is used for training, then data generation efficiency improves, but model generalization to real-world data deteriorates

Engineering Contradiction:
Improvedata generation efficiencyVSAvoidgeneralization to real-world data
Core Design Contradiction:
ProductivityVSReliability

Solution Approach 1:

The patent implements a two-stage training approach where the model is first pre-trained on efficiently generated synthetic data to learn fundamental pose estimation patterns, then fine-tuned on a smaller set of real-world data to adapt to actual variations. This preliminary training on synthetic data followed by real-world adaptation resolves the generalization problem while maintaining data generation efficiency.

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS20220392099A1Stable pose estimation with analysis by synthesis
Publication Date: 2022.12.08 DISNEY ENTERPRISES INC
  • US20220392099A1 patent drawing
  • US20220392099A1 patent drawing
  • US20220392099A1 patent drawing

AI summary

One embodiment of the present invention sets forth a technique for generating a pose estimation model. The technique includes generating one or more trained components included in the pose estimation model based on a first set of training images and a first set of labeled poses associated with the first set of training images, wherein each labeled pose includes a first set of positions on a left side of an object and a second set of positions on a right side of the object. The technique also includes training the pose estimation model based on a set of reconstructions of a second set of training images, wherein the set of reconstructions is generated by the pose estimation model from a set of predicted poses outputted by the one or more trained components.