Avatar Expression Mimicry via Neural Texture Synthesis
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current video conference systems lack the ability to effectively generate avatars that mimic the expressions and gaze directions of participants in a virtual 3D environment, leading to a less immersive and less natural interaction experience.
Innovation Solution
A system and method for generating avatars in a 3D video conference environment that includes receiving direction of gaze information, determining updated 3D participant representation, and generating texture maps and 3D models to reflect the participant's expressions and gaze, using neural networks and generative adversarial networks to create realistic and dynamic avatars.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Manufacturing precision
If traditional video conference systems are used, then the system is simple and easy to operate, but the avatar cannot effectively mimic expressions and gaze directions of participants
Solution Approach 1:
The patent introduces an intermediary system comprising neural networks and generative adversarial networks that act as a mediator between the participant's real-time video feed and the avatar representation. This intermediary processes video frames to extract facial landmarks, generates expression parameters, and synthesizes realistic avatar expressions, thereby achieving accurate expression mimicry without requiring direct complex hardware modifications to the video conference system.
Solution Approach 2:
The patent replaces traditional mechanical or rule-based expression mapping systems with machine learning-based neural networks. Instead of using predefined rules to map facial movements to avatar expressions, the system employs trained neural networks that learn the complex non-linear relationships between real facial expressions and corresponding avatar expressions, achieving higher accuracy while maintaining system efficiency.
2Reliability
If neural networks and generative adversarial networks are used to generate realistic avatars, then the realism and immersion of the virtual 3D environment is improved, but the computational resources and processing time increase
Solution Approach 1:
The patent applies preliminary action by pre-training the neural networks and generative adversarial networks offline using large datasets of facial expressions and video data. This pre-training phase captures the complex relationships between real and virtual expressions, allowing the trained models to perform efficient real-time inference during actual video conferences with minimal computational overhead, thus achieving high realism without excessive energy consumption during operation.
3Manufacturing precision
If the camera is located close to the display for video conference calls, then the setup is simple, but the avatar representation and expression capture quality deteriorates
Solution Approach 1:
The patent employs copying by creating a virtual 3D replica (avatar) of the participant based on their video feed, rather than requiring direct high-quality camera capture. The neural networks extract essential facial expression information from standard video feeds and synthesize a photorealistic 3D copy that mimics the participant's expressions and gaze directions, achieving high expression capture quality without requiring specialized camera equipment or complex physical setups.
Data Source
AI summary
A method for generating an avatar having expressions that mimics expressions of a person, the method may include obtaining expression parameters that represent an expression of the person; generating, in real time, a texture map of a face of the person, the texture map of the face of the person represents the expression of the person, wherein the generating is based on the expression parameters, wherein the generating comprises (i) determining, using a neural network and based on the expression parameters, weights, and (ii) calculating a weighted sum of a set of base texture maps, wherein the set of base texture maps belongs to a group of base texture maps that spans different acquired texture maps, the acquired texture maps are calculated based on images of arbitrary expressions made by the person; and rendering the avatar using the texture map.


