Avatar Rendering from IR Headset Cameras via Domain Transfer

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

The challenge lies in constructing an avatar that accurately mimics facial expressions based on partial, close-up, and oblique views of the face captured by IR cameras in AR/VR headsets, due to a modality gap between IR and visible-light spectrums, and the lack of clear correspondence between captured IR images and actual facial expressions.

Innovation Solution

A method involving a domain-transfer machine learning model to transfer IR images to rendered avatar images, followed by a parameter-extraction model to identify avatar parameters, which are then used to train a real-time tracking model for non-intrusive cameras to accurately map IR images to avatar parameters, enabling accurate facial expression and head pose rendering.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If IR cameras are used to capture facial images in AR/VR headsets, then the headset can provide artificial reality content, but the IR cameras only provide partial, close-up, oblique views of the face instead of a complete view

Engineering Contradiction:
ImproveAR/VR functionalityVSAvoidfacial expression information
Core Design Contradiction:
Adaptability or versatilityVSLoss of information

Solution Approach 1:

The patent introduces an intermediary system consisting of multiple machine learning models (domain transfer model, parameter extraction model, real-time tracking model) that acts as a mediator between the limited IR camera inputs and the complete avatar representation. This intermediary processing chain transforms the partial oblique views into comprehensive facial expression data, resolving the information loss problem while maintaining AR/VR functionality

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The patent creates a virtual copy (avatar) of the user's face that replicates facial expressions. Instead of directly observing the complete real face, the system generates a virtual representation that copies the essential expressive characteristics from the limited IR camera views, allowing full facial expression capture without requiring complete physical visibility

Inventive Principle:
Principle #26Copying

2Measurement precision

If a training headset with additional intrusive IR cameras is used, then more facial views can be captured for training, but the cameras still generate only patchwork close-up oblique views and are more intrusive

Engineering Contradiction:
Improvefacial expression capture accuracyVSAvoiduser comfort
Core Design Contradiction:
Measurement precisionVSEase of operation

Solution Approach 1:

The patent performs preliminary action by using an intrusive training headset with additional cameras to collect comprehensive training data, then uses this data to train machine learning models that can subsequently operate with fewer, less-intrusive cameras. The heavy data collection and model training is done in advance, allowing the final system to achieve high measurement precision with minimal user intrusion

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent segments the camera system into two distinct phases: a training phase with multiple intrusive cameras for comprehensive data collection, and an operational phase with fewer non-intrusive cameras. This segmentation allows the system to achieve high measurement precision during training while maintaining user comfort during actual use

Inventive Principle:
Principle #1Segmentation

3Loss of information

If visible-light cameras are used to supplement IR cameras, then more facial views can be obtained, but visible-light cameras still do not provide views of portions of the face occluded by the headset

Engineering Contradiction:
Improvefacial visibilityVSAvoidcamera system complexity
Core Design Contradiction:
Loss of informationVSDevice complexity

Solution Approach 1:

The patent introduces machine learning models as intermediaries that process and fuse data from both IR and visible-light cameras. This intermediary processing layer synthesizes the complementary information from both camera types, reconstructing occluded facial regions by intelligently combining the partial views, thereby reducing information loss without simply adding more physical cameras

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The patent creates a composite sensing system that combines IR and visible-light camera data streams. By fusing these different spectral modalities through machine learning, the system achieves comprehensive facial visibility that neither camera type could achieve alone, effectively creating a composite view that overcomes the limitations of individual sensor types

Inventive Principle:
Principle #40Composite materials

4Measurement precision

If machine learning models are trained to transfer IR images to avatar parameters, then correspondence can be established, but the process requires complex multi-model training

Engineering Contradiction:
ImproveIR to avatar parameter mapping accuracyVSAvoidmodel training complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent segments the complex mapping problem into three distinct sequential machine learning models: (1) domain transfer model that converts IR images to visible-light domain, (2) parameter extraction model that extracts avatar parameters from the transferred images, and (3) real-time tracking model that performs final tracking. This segmentation allows each model to specialize in a specific transformation step, improving overall mapping accuracy while making the training process more manageable through modular design

Inventive Principle:
Principle #1Segmentation

Data Source

PatentUS10885693B1Animating avatars from headset cameras
Publication Date: 2021.01.05 META PLATFORMS TECHNOLOGIES LLC
  • US10885693B1 patent drawing
  • US10885693B1 patent drawing
  • US10885693B1 patent drawing

AI summary

In one embodiment, a computing system may access a plurality of first captured images that are captured in a first spectral domain, generate, using a first machine-learning model, a plurality of first domain-transferred images based on the first captured images, wherein the first domain-transferred images are in a second spectral domain, render, based on a first avatar, a plurality of first rendered images comprising views of the first avatar, and update the first machine-learning model based on comparisons between the first domain-transferred images and the first rendered images, wherein the first machine-learning model is configured to translate images in the first spectral domain to the second spectral domain. The system may also generate, using a second machine-learning model, the first avatar based on the first captured images. The first avatar may be rendered using a parametric face model based on a plurality of avatar parameters.