Neural Network Keypoint Disentanglement for Unsupervised Pose Prediction

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Pose prediction in computer vision requires expensive pixel-level annotations for object parts, making the training process costly and labor-intensive due to the need for human supervision.

Innovation Solution

The use of neural networks that learn semantic keypoint-like representations without expensive keypoint-level annotations by disentangling input data into foreground and background components, allowing for unsupervised training and reconstruction of keypoints.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If pixel-level annotations are used for training pose prediction models, then measurement precision is improved, but loss of time and manufacturing cost increase due to human supervision requirements

Engineering Contradiction:
Improvekeypoint annotation precisionVSAvoidtraining time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The system uses self-supervised learning where the neural network automatically generates its own training data by predicting keypoint annotations from images without requiring manual annotations. The model trains by reconstructing images from predicted keypoints, allowing the system to serve itself rather than requiring human annotators.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The patent creates synthetic annotated data by copying and transforming unlabeled images. The neural network generates fake annotations by predicting keypoint locations and using these to create synthetic labeled images that mimic real annotated data, allowing the model to learn from synthetic copies rather than real manual annotations.

Inventive Principle:
Principle #26Copying

2Manufacturing precision

If pixel-level annotations are provided for training data, then manufacturing precision is improved, but device complexity and cost increase due to human supervision requirements

Engineering Contradiction:
Improveannotation precisionVSAvoidtraining system complexity
Core Design Contradiction:
Manufacturing precisionVSDevice complexity

Solution Approach 1:

The system eliminates the need for complex manual annotation processes by implementing self-supervised learning. The neural network automatically generates annotations and training data through image reconstruction tasks, simplifying the overall training system by removing human supervision infrastructure.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The patent replaces the mechanical process of manual human annotation with an automated neural network-based system. Instead of requiring human experts to manually mark keypoints, the system uses deep learning models to automatically predict and generate annotations through computational processes.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

3Loss of time

If unsupervised learning without annotations is used, then loss of time is reduced, but measurement precision deteriorates due to lack of ground truth data

Engineering Contradiction:
Improveannotation timeVSAvoidkeypoint prediction accuracy
Core Design Contradiction:
Loss of timeVSMeasurement precision

Solution Approach 1:

The system implements feedback loops where the neural network predicts keypoints, reconstructs images from these predictions, and uses the reconstruction error as feedback to refine future predictions. This iterative feedback process gradually improves prediction accuracy without requiring initial ground truth annotations.

Inventive Principle:
Principle #23Feedback

Solution Approach 2:

The patent creates synthetic annotated data by copying and transforming unlabeled images. The neural network generates fake annotations by predicting keypoint locations and using these to create synthetic labeled images that mimic real annotated data, allowing the model to learn from synthetic copies rather than real manual annotations.

Inventive Principle:
Principle #26Copying

4Manufacturing precision

If manual annotation of object parts is performed, then manufacturing precision is improved, but productivity decreases due to labor-intensive process

Engineering Contradiction:
Improvekeypoint annotation precisionVSAvoidtraining data preparation efficiency
Core Design Contradiction:
Manufacturing precisionVSProductivity

Solution Approach 1:

The system uses self-supervised learning where the neural network automatically generates its own training data by predicting keypoint annotations from images without requiring manual annotations. The model trains by reconstructing images from predicted keypoints, allowing the system to serve itself rather than requiring human annotators.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The patent replaces the mechanical process of manual human annotation with an automated neural network-based system. Instead of requiring human experts to manually mark keypoints, the system uses deep learning models to automatically predict and generate annotations through computational processes.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Data Source

PatentUS20220180528A1Disentanglement of image attributes using a neural network
Publication Date: 2022.06.09 NVIDIA CORP
  • US20220180528A1 patent drawing
  • US20220180528A1 patent drawing
  • US20220180528A1 patent drawing

AI summary

Apparatuses, systems, and techniques to perform unsupervised keypoint or landmark learning using one or more neural networks. In at least one embodiment, one or more neural networks use pose and appearance information to construct a foreground and a background, which are then used to reconstruct an input image and determine loss values to train the one or more neural networks.