Mixed Reality Avatar Eye Inpainting via Generative Model

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current mixed reality technologies struggle to accurately capture and display a user's eye regions, especially when they are partially or fully occluded, leading to a less immersive experience due to inadequate representation in virtual environments.

Innovation Solution

A method and system that utilize a trained eye landmark generative model to generate accurate 3D eye landmarks by receiving 3D non-eye landmarks and voice audio, applying random noise, and performing iterative refinement to render a realistic 3D face mesh, ensuring accurate eye region representation in mixed reality environments.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If external depth camera is used to capture eye regions, then 3D face mesh can be obtained, but eye regions may be partially or fully occluded and cannot be accurately captured

Engineering Contradiction:
Improveeye region capture accuracyVSAvoideye region detection reliability
Core Design Contradiction:
Measurement precisionVSReliability

Solution Approach 1:

The patent introduces an eye landmark generative model as an intermediary that takes voice audio and non-eye facial landmarks as input to generate missing eye landmarks. This mediator bridges the gap when direct camera capture fails due to occlusion, allowing the system to reconstruct eye regions that cannot be directly observed.

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The patent replaces the mechanical camera-based eye tracking system with an AI-based generative model that uses voice audio signals and facial geometry to infer eye landmarks. This substitution allows eye region capture to work reliably even when the physical camera cannot directly observe the eyes.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

2Adaptability or versatility

If voice audio is used as input for eye landmark generation, then occlusion issues are addressed, but system complexity increases

Engineering Contradiction:
Improveocclusion handling capabilityVSAvoidsystem architecture complexity
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The patent makes the eye landmark generative model multi-functional by enabling it to process both voice audio inputs and non-eye facial landmark inputs to produce eye landmarks. This universal approach allows a single system to handle multiple input modalities and various occlusion scenarios without requiring separate specialized systems.

Inventive Principle:
Principle #6Universality (Multi-functionality)

3Measurement precision

If iterative refinement is performed on generated eye landmarks, then accuracy is improved, but processing time increases

Engineering Contradiction:
Improveeye landmark accuracyVSAvoidprocessing time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The patent implements periodic action through iterative refinement, where the eye landmark generative model processes the input data multiple times in successive iterations. Each iteration refines the eye landmark estimates, progressively improving accuracy. The system performs a fixed number of iterations to balance accuracy improvement with processing time constraints.

Inventive Principle:
Principle #19Periodic action

Data Source

PatentUS20240320926A1Mixed reality avatar eye inpainting based on user speech
Publication Date: 2024.09.26 INTERNATIONAL BUSINESS MACHINE CORPORATION
  • US20240320926A1 patent drawing
  • US20240320926A1 patent drawing
  • US20240320926A1 patent drawing

AI summary

According to one embodiment, a method, computer system, and computer program product for mixed reality is provided. The present invention may include receiving one or more 3D non-eye landmarks of a user; receiving at least one voice audio of the user; using random noise sampled with a unit normal distribution as one or more noised 3D eye landmarks for the user; inputting the received one or more 3D non-eye landmarks of the user, the at least one voice audio of the user, and the one or more noised 3D eye landmarks for the user, into a trained eye landmark generative model; generating one or more 3D eye landmarks for the user using the trained eye landmark generative model; performing iterative refinement of the one or more generated 3D eye landmarks using the trained eye landmark generative model; and rendering the user's generated face model using a formed 3D face mesh.