Avatar Rendering from IR Headset Cameras via Domain Transfer
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
The challenge lies in constructing an avatar that accurately mimics facial expressions based on partial, close-up, and oblique views of the face captured by IR cameras in AR/VR headsets, due to a modality gap between IR and visible-light spectrums, and the lack of clear correspondence between captured IR images and actual facial expressions.
Innovation Solution
A method involving a domain-transfer machine learning model to transfer IR images to rendered avatar images, followed by a parameter-extraction model to identify avatar parameters, which are then used to train a real-time tracking model for non-intrusive cameras to accurately map IR images to avatar parameters, enabling accurate facial expression and head pose rendering.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If IR cameras are used to capture facial images in AR/VR headsets, then the headset can provide artificial reality content, but the IR cameras only provide partial, close-up, oblique views of the face instead of a complete view
Solution Approach 1:
The patent introduces an intermediary system consisting of multiple machine learning models (domain transfer model, parameter extraction model, real-time tracking model) that acts as a mediator between the limited IR camera inputs and the complete avatar representation. This intermediary processing chain transforms the partial oblique views into comprehensive facial expression data, resolving the information loss problem while maintaining AR/VR functionality
Solution Approach 2:
The patent creates a virtual copy (avatar) of the user's face that replicates facial expressions. Instead of directly observing the complete real face, the system generates a virtual representation that copies the essential expressive characteristics from the limited IR camera views, allowing full facial expression capture without requiring complete physical visibility
2Measurement precision
If a training headset with additional intrusive IR cameras is used, then more facial views can be captured for training, but the cameras still generate only patchwork close-up oblique views and are more intrusive
Solution Approach 1:
The patent performs preliminary action by using an intrusive training headset with additional cameras to collect comprehensive training data, then uses this data to train machine learning models that can subsequently operate with fewer, less-intrusive cameras. The heavy data collection and model training is done in advance, allowing the final system to achieve high measurement precision with minimal user intrusion
Solution Approach 2:
The patent segments the camera system into two distinct phases: a training phase with multiple intrusive cameras for comprehensive data collection, and an operational phase with fewer non-intrusive cameras. This segmentation allows the system to achieve high measurement precision during training while maintaining user comfort during actual use
3Loss of information
If visible-light cameras are used to supplement IR cameras, then more facial views can be obtained, but visible-light cameras still do not provide views of portions of the face occluded by the headset
Solution Approach 1:
The patent introduces machine learning models as intermediaries that process and fuse data from both IR and visible-light cameras. This intermediary processing layer synthesizes the complementary information from both camera types, reconstructing occluded facial regions by intelligently combining the partial views, thereby reducing information loss without simply adding more physical cameras
Solution Approach 2:
The patent creates a composite sensing system that combines IR and visible-light camera data streams. By fusing these different spectral modalities through machine learning, the system achieves comprehensive facial visibility that neither camera type could achieve alone, effectively creating a composite view that overcomes the limitations of individual sensor types
4Measurement precision
If machine learning models are trained to transfer IR images to avatar parameters, then correspondence can be established, but the process requires complex multi-model training
Solution Approach 1:
The patent segments the complex mapping problem into three distinct sequential machine learning models: (1) domain transfer model that converts IR images to visible-light domain, (2) parameter extraction model that extracts avatar parameters from the transferred images, and (3) real-time tracking model that performs final tracking. This segmentation allows each model to specialize in a specific transformation step, improving overall mapping accuracy while making the training process more manageable through modular design
Data Source
AI summary
In one embodiment, a computing system may access a plurality of first captured images that are captured in a first spectral domain, generate, using a first machine-learning model, a plurality of first domain-transferred images based on the first captured images, wherein the first domain-transferred images are in a second spectral domain, render, based on a first avatar, a plurality of first rendered images comprising views of the first avatar, and update the first machine-learning model based on comparisons between the first domain-transferred images and the first rendered images, wherein the first machine-learning model is configured to translate images in the first spectral domain to the second spectral domain. The system may also generate, using a second machine-learning model, the first avatar based on the first captured images. The first avatar may be rendered using a parametric face model based on a plurality of avatar parameters.


