Avatar Expression Mimicry via Neural Texture Synthesis

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current video conference systems lack the ability to effectively generate avatars that mimic the expressions and gaze directions of participants in a virtual 3D environment, leading to a less immersive and less natural interaction experience.

Innovation Solution

A system and method for generating avatars in a 3D video conference environment that includes receiving direction of gaze information, determining updated 3D participant representation, and generating texture maps and 3D models to reflect the participant's expressions and gaze, using neural networks and generative adversarial networks to create realistic and dynamic avatars.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Manufacturing precision

If traditional video conference systems are used, then the system is simple and easy to operate, but the avatar cannot effectively mimic expressions and gaze directions of participants

Engineering Contradiction:
Improveexpression mimicry accuracyVSAvoidsystem complexity
Core Design Contradiction:
Manufacturing precisionVSDevice complexity

Solution Approach 1:

The patent introduces an intermediary system comprising neural networks and generative adversarial networks that act as a mediator between the participant's real-time video feed and the avatar representation. This intermediary processes video frames to extract facial landmarks, generates expression parameters, and synthesizes realistic avatar expressions, thereby achieving accurate expression mimicry without requiring direct complex hardware modifications to the video conference system.

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The patent replaces traditional mechanical or rule-based expression mapping systems with machine learning-based neural networks. Instead of using predefined rules to map facial movements to avatar expressions, the system employs trained neural networks that learn the complex non-linear relationships between real facial expressions and corresponding avatar expressions, achieving higher accuracy while maintaining system efficiency.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

2Reliability

If neural networks and generative adversarial networks are used to generate realistic avatars, then the realism and immersion of the virtual 3D environment is improved, but the computational resources and processing time increase

Engineering Contradiction:
Improveavatar realismVSAvoidcomputational energy consumption
Core Design Contradiction:
ReliabilityVSUse of energy by moving object

Solution Approach 1:

The patent applies preliminary action by pre-training the neural networks and generative adversarial networks offline using large datasets of facial expressions and video data. This pre-training phase captures the complex relationships between real and virtual expressions, allowing the trained models to perform efficient real-time inference during actual video conferences with minimal computational overhead, thus achieving high realism without excessive energy consumption during operation.

Inventive Principle:
Principle #10Preliminary action

3Manufacturing precision

If the camera is located close to the display for video conference calls, then the setup is simple, but the avatar representation and expression capture quality deteriorates

Engineering Contradiction:
Improveexpression capture qualityVSAvoidcamera setup complexity
Core Design Contradiction:
Manufacturing precisionVSEase of operation

Solution Approach 1:

The patent employs copying by creating a virtual 3D replica (avatar) of the participant based on their video feed, rather than requiring direct high-quality camera capture. The neural networks extract essential facial expression information from standard video feeds and synthesize a photorealistic 3D copy that mimics the participant's expressions and gaze directions, achieving high expression capture quality without requiring specialized camera equipment or complex physical setups.

Inventive Principle:
Principle #26Copying

Data Source

PatentUS20230085339A1Generating an avatar having expressions that mimics expressions of a person
Publication Date: 2023.03.16 CAVENDISH CAPITAL LLC
  • US20230085339A1 patent drawing
  • US20230085339A1 patent drawing
  • US20230085339A1 patent drawing

AI summary

A method for generating an avatar having expressions that mimics expressions of a person, the method may include obtaining expression parameters that represent an expression of the person; generating, in real time, a texture map of a face of the person, the texture map of the face of the person represents the expression of the person, wherein the generating is based on the expression parameters, wherein the generating comprises (i) determining, using a neural network and based on the expression parameters, weights, and (ii) calculating a weighted sum of a set of base texture maps, wherein the set of base texture maps belongs to a group of base texture maps that spans different acquired texture maps, the acquired texture maps are calculated based on images of arbitrary expressions made by the person; and rendering the avatar using the texture map.