Viewpoint-Agnostic 3D Pose Estimation from Monocular Images

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current occupant monitoring systems in vehicles face challenges in accurately predicting three-dimensional pose estimates of occupants from two-dimensional image data without knowledge of the image sensor's location, rotation, and translation, and require expensive depth sensors for reliable 3D data, which limits their deployment in production vehicles.

Innovation Solution

A viewpoint-agnostic occupant pose detection model is trained using multi-view image sensor data from synchronized optical image sensors to predict relative scale-normalized 3D pose estimates from a single monocular image frame, employing a combination of pose alignment, pose depth, and ground truth kinematic losses for improved accuracy, without the need for bulky motion capture suits or expensive depth sensors.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If expensive depth sensors are used to obtain reliable 3D data, then measurement precision is improved, but device complexity and cost increase

Engineering Contradiction:
Improve3D pose estimation accuracyVSAvoidsensor deployment complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent creates a virtual 3D copy of the occupant pose from 2D image data using a trained neural network model. Instead of using expensive depth sensors to directly measure 3D position, the system copies 3D pose information from 2D images through learned transformations, achieving accurate 3D estimation without additional hardware complexity

Inventive Principle:
Principle #26Copying

Solution Approach 2:

The patent replaces the mechanical/optical depth sensing system with a computational approach. A neural network model processes 2D image data and predicts 3D pose estimates, substituting physical depth sensors with an information-processing system that achieves the same measurement goal through software rather than hardware

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

2Adaptability or versatility

If viewpoint-agnostic 3D pose estimation is achieved without sensor location knowledge, then adaptability is improved, but measurement precision deteriorates

Engineering Contradiction:
Improveviewpoint independenceVSAvoid3D pose estimation accuracy
Core Design Contradiction:
Adaptability or versatilityVSMeasurement precision

Solution Approach 1:

The patent changes the parameter representation by training the neural network to output relative scale-normalized 3D pose estimates rather than absolute measurements. This parameter transformation allows the model to be viewpoint-agnostic while maintaining accuracy, as the normalized parameters are invariant to camera position and rotation

Inventive Principle:
Principle #35Parameter changes

Solution Approach 2:

The patent performs preliminary training of the neural network model using multi-view training data that includes various sensor locations, rotations, and translations. This preliminary action embeds viewpoint invariance into the model during training, enabling it to generalize to unseen viewpoints while maintaining measurement precision

Inventive Principle:
Principle #10Preliminary action

3Measurement precision

If multi-view training data from multiple synchronized sensors is used, then measurement precision is improved, but device complexity increases

Engineering Contradiction:
Improvepose detection accuracyVSAvoidsensor synchronization complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent creates a virtual multi-view training dataset by synthesizing images from different viewpoints using a single camera and 3D pose information. Instead of requiring multiple physical synchronized sensors, the system copies and transforms a single viewpoint image into multiple virtual viewpoints through geometric transformations and the trained pose estimation model

Inventive Principle:
Principle #26Copying

Solution Approach 2:

The patent introduces a single camera as an intermediary that captures 2D images, which then serve as input to the neural network model. This intermediary approach eliminates the need for multiple synchronized sensors, as the single camera data is transformed into multi-view equivalent information through computational methods

Inventive Principle:
Principle #24Intermediary (Mediator)

Data Source

PatentUS20250022155A1Three-dimensional pose estimation using two-dimensional images
Publication Date: 2025.01.16 NVIDIA CORP
  • US20250022155A1 patent drawing
  • US20250022155A1 patent drawing
  • US20250022155A1 patent drawing

AI summary

In various examples, systems and methods for pose detection model training for predicting three-dimensional pose estimates using two-dimensional image data are provided. The occupant pose detection model may be trained using multi-view image sensor training data that includes image frames that capture a pose of a training subject within a machine interior using multiple synchronized optical image sensors placed around the machine interior that produce a set of captured image frames of the training subject from different viewpoints. Based on the multi-view image sensor training data, the occupant pose detection model may generate a set of individual, predicted 3D pose estimates for the training subject from a captured image frame from each of the respective optical image sensors. To adjust the occupant pose detection model during training, a loss feedback may be generated that comprises a pose alignment loss, a pose depth loss, and/or a ground truth kinematic loss.