Periocular and Audio Synthesis for Full Face Image Generation

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

In mixed reality environments, head-mounted devices struggle to update a 3D virtual avatar of a user's face with facial expressions, especially when the periocular region is occluded, as they cannot directly image the lower face region, leading to difficulties in synthesizing a complete and dynamic face image.

Innovation Solution

The system generates a mapping between the observed periocular region and the unobserved lower face using audio inputs, such as phonemes, and visemes, combined with periocular images, to deduce the conformations of the lower face, allowing the head-mounted device to synthesize a full face image by combining observed and unobserved portions.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If a head-mounted device is used to capture face images in mixed reality environments, then the device can provide immersive VR/AR/MR experiences, but the device cannot directly image the lower face region due to occlusion by the device itself

Engineering Contradiction:
Improveimmersive experience capabilityVSAvoidlower face region visibility
Core Design Contradiction:
Adaptability or versatilityVSLoss of information

Solution Approach 1:

The patent uses an intermediary computational model (mapping from periocular region to lower face) to bridge the gap between observable and unobservable regions. Audio inputs and visemes serve as additional intermediaries to infer lower face conformation when directly imaging is impossible.

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The system creates a virtual copy of the lower face region by generating synthetic image data based on mappings from the periocular region and audio inputs, allowing the complete face to be reconstructed without directly imaging the occluded area.

Inventive Principle:
Principle #26Copying

2Loss of information

If the head-mounted device tries to capture the complete face image, then the full face information is obtained, but the device structure prevents direct observation of the lower face region

Engineering Contradiction:
Improveface image completenessVSAvoidimaging system structure
Core Design Contradiction:
Loss of informationVSDevice complexity

Solution Approach 1:

The face imaging problem is segmented into observable (periocular) and unobservable (lower face) regions. The system processes these segments separately, using the observable region and audio inputs to infer the unobservable region, then combines them to reconstruct the complete face image.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The system transitions from a purely spatial imaging approach to incorporating temporal and acoustic dimensions. By using audio inputs and visemes over time, the system infers lower face conformation that cannot be obtained through spatial imaging alone.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

3Device complexity

If the system uses only periocular images to infer lower face conformation, then the imaging process is simple, but the accuracy of facial expression synthesis is insufficient

Engineering Contradiction:
Improveimaging process simplicityVSAvoidfacial expression accuracy
Core Design Contradiction:
Device complexityVSMeasurement precision

Solution Approach 1:

The system merges multiple data sources including periocular images, audio inputs, and viseme information to infer lower face conformation. This combination of heterogeneous data sources improves the accuracy of facial expression synthesis beyond what any single source could provide.

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The system uses feedback from audio inputs and viseme recognition to continuously refine the inferred lower face conformation. The mapping model is trained and adjusted based on the relationship between audio/viseme data and actual facial expressions, improving synthesis accuracy over time.

Inventive Principle:
Principle #23Feedback

Data Source

PatentEP4202840A1Periocular and audio synthesis of a full face image
Publication Date: 2023.06.28 MAGIC LEAP INC
  • EP4202840A1 patent drawingFigure 1
  • EP4202840A1 patent drawingFigure 2
  • EP4202840A1 patent drawingFigure 3

AI summary

Systems and methods for training a machine learning derived model using a first plurality of images of a first region of a user's face and a second plurality of images of a second region, to generate a mapping from a first conformation of the first region to a second conformation of the second region.