Viewpoint-Agnostic 3D Pose Estimation from Monocular Images
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current occupant monitoring systems in vehicles face challenges in accurately predicting three-dimensional pose estimates of occupants from two-dimensional image data without knowledge of the image sensor's location, rotation, and translation, and require expensive depth sensors for reliable 3D data, which limits their deployment in production vehicles.
Innovation Solution
A viewpoint-agnostic occupant pose detection model is trained using multi-view image sensor data from synchronized optical image sensors to predict relative scale-normalized 3D pose estimates from a single monocular image frame, employing a combination of pose alignment, pose depth, and ground truth kinematic losses for improved accuracy, without the need for bulky motion capture suits or expensive depth sensors.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If expensive depth sensors are used to obtain reliable 3D data, then measurement precision is improved, but device complexity and cost increase
Solution Approach 1:
The patent creates a virtual 3D copy of the occupant pose from 2D image data using a trained neural network model. Instead of using expensive depth sensors to directly measure 3D position, the system copies 3D pose information from 2D images through learned transformations, achieving accurate 3D estimation without additional hardware complexity
Solution Approach 2:
The patent replaces the mechanical/optical depth sensing system with a computational approach. A neural network model processes 2D image data and predicts 3D pose estimates, substituting physical depth sensors with an information-processing system that achieves the same measurement goal through software rather than hardware
2Adaptability or versatility
If viewpoint-agnostic 3D pose estimation is achieved without sensor location knowledge, then adaptability is improved, but measurement precision deteriorates
Solution Approach 1:
The patent changes the parameter representation by training the neural network to output relative scale-normalized 3D pose estimates rather than absolute measurements. This parameter transformation allows the model to be viewpoint-agnostic while maintaining accuracy, as the normalized parameters are invariant to camera position and rotation
Solution Approach 2:
The patent performs preliminary training of the neural network model using multi-view training data that includes various sensor locations, rotations, and translations. This preliminary action embeds viewpoint invariance into the model during training, enabling it to generalize to unseen viewpoints while maintaining measurement precision
3Measurement precision
If multi-view training data from multiple synchronized sensors is used, then measurement precision is improved, but device complexity increases
Solution Approach 1:
The patent creates a virtual multi-view training dataset by synthesizing images from different viewpoints using a single camera and 3D pose information. Instead of requiring multiple physical synchronized sensors, the system copies and transforms a single viewpoint image into multiple virtual viewpoints through geometric transformations and the trained pose estimation model
Solution Approach 2:
The patent introduces a single camera as an intermediary that captures 2D images, which then serve as input to the neural network model. This intermediary approach eliminates the need for multiple synchronized sensors, as the single camera data is transformed into multi-view equivalent information through computational methods
Data Source
AI summary
In various examples, systems and methods for pose detection model training for predicting three-dimensional pose estimates using two-dimensional image data are provided. The occupant pose detection model may be trained using multi-view image sensor training data that includes image frames that capture a pose of a training subject within a machine interior using multiple synchronized optical image sensors placed around the machine interior that produce a set of captured image frames of the training subject from different viewpoints. Based on the multi-view image sensor training data, the occupant pose detection model may generate a set of individual, predicted 3D pose estimates for the training subject from a captured image frame from each of the respective optical image sensors. To adjust the occupant pose detection model during training, a loss feedback may be generated that comprises a pose alignment loss, a pose depth loss, and/or a ground truth kinematic loss.


