CycleGAN Lighting Translation for Hand Pose Estimation Data Augmentation

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing hand pose estimation methods using deep neural networks require a large amount of training data, which is often not available in practice, and face challenges with inconsistent lighting conditions between synthetic and real-world images.

Innovation Solution

The method employs a Cycle-Consistent Adversarial Network (CycleGAN) to translate lighting conditions between synthetic hand pose images and real-world background images, generating augmented training data that aligns with real-world lighting conditions, thereby improving the accuracy and reliability of hand pose estimation models.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If deep neural networks are used for hand pose estimation, then estimation accuracy is improved, but the requirement for large amount of training data increases

Engineering Contradiction:
Improvehand pose estimation accuracyVSAvoidtraining data quantity
Core Design Contradiction:
Measurement precisionVSQuantity of substance

Solution Approach 1:

The patent uses CycleGAN to generate synthetic training images by copying and transforming hand pose data across different lighting conditions. The model learns to translate source domain images (with one lighting condition) to target domain images (with different lighting conditions), creating artificial training data that expands the effective training dataset without requiring additional real-world data collection

Inventive Principle:
Principle #26Copying

Solution Approach 2:

The patent changes the lighting condition parameter of training images using CycleGAN translation. By transforming images under different lighting conditions, the system creates diverse training samples from limited source data, effectively increasing the quantity and variability of training data while maintaining the underlying hand pose information

Inventive Principle:
Principle #35Parameter changes

2Quantity of substance

If synthetic training data is used to augment data quantity, then training data quantity is improved, but lighting condition consistency deteriorates

Engineering Contradiction:
Improvetraining data quantityVSAvoidlighting condition consistency
Core Design Contradiction:
Quantity of substanceVSReliability

Solution Approach 1:

The patent implements cycle-consistency feedback where images are translated from source to target domain and then back to source domain. The CycleGAN model is trained to minimize the reconstruction error, ensuring that the translated images maintain structural consistency with the original images while adapting to target lighting conditions. This feedback mechanism guarantees that synthetic data remains reliable and consistent

Inventive Principle:
Principle #23Feedback

Solution Approach 2:

The patent uses CycleGAN as an intermediary translation layer between source and target domains. Rather than directly combining synthetic and real data with potentially conflicting lighting conditions, the model mediates the transformation, learning the mapping between different lighting conditions and producing synthetic data that is consistent with the target domain's lighting characteristics

Inventive Principle:
Principle #24Intermediary (Mediator)

Data Source

PatentUS11176699B2Augmenting reliable training data with CycleGAN for hand pose estimation
Publication Date: 2021.11.16 TENCENT AMERICA LLC
  • US11176699B2 patent drawing
  • US11176699B2 patent drawing
  • US11176699B2 patent drawing

AI summary

A method and apparatus for generating augmented training data for hand pose estimation include receiving source data that is associated with a first lighting condition. Target data that is associated with a second lighting condition is received. A lighting condition translation between the first lighting condition and the second lighting condition is determined. Lighting translated data is generated based on the lighting condition translation and the source data. Augmented training data for hand pose estimation is generated based on the target data and the lighting translated data.