Real-time Avatar Generation Using Dynamic Texture Blending
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing methods for generating realistic facial animations from a single image often result in artifacts such as uncanny eyes, unusual mouth shapes, and texture deformation when applied to three-dimensional models, leading to less than excellent results.
Innovation Solution
A system utilizing a generative adversarial network (GAN) with dynamic textures, which includes training data of two-dimensional and fully modeled face images, generates key expression meshes and blendshape texture maps to create photorealistic avatars capable of real-time animation, addressing the challenges of facial texture generation and animation from a single image.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Ease of manufacture
If traditional single-image avatar generation methods are used, then the avatar can be created from minimal input, but the facial animations contain artifacts and deformations
Solution Approach 1:
The patent segments the facial animation problem into multiple key expressions (e.g., neutral, smile, frown, surprise) and generates specific texture maps for each expression. This segmentation allows the system to handle complex facial animations by breaking them down into manageable, pre-rendered components that can be blended together, avoiding the artifacts that occur with traditional single-image generation methods.
Solution Approach 2:
The patent performs preliminary action by pre-generating texture maps for key facial expressions during an offline training phase. These pre-computed texture maps are then stored and rapidly blended during real-time animation, eliminating the need to generate textures on-the-fly and avoiding the deformation artifacts that plague real-time single-image generation.
2Manufacturing precision
If complex rigs and multi-camera systems are used, then facial texture accuracy is improved, but the system complexity and cost increase
Solution Approach 1:
The patent uses a trained neural network to copy realistic facial textures from a training dataset into the generated avatar. The neural network learns the statistical relationships between facial structures and textures from diverse examples, then synthesizes photorealistic textures for the target avatar without requiring complex physical capture equipment.
Solution Approach 2:
The patent replaces the mechanical multi-camera rig system with a computational approach using a trained neural network. Instead of using physical cameras and light rigs to capture facial textures, the system uses machine learning to synthesize realistic textures from a single input image, dramatically reducing hardware complexity while maintaining or improving texture quality.
3Speed
If real-time animation is implemented, then the avatar responsiveness is improved, but texture deformation and artifacts increase
Solution Approach 1:
The patent performs preliminary action by pre-generating texture maps for key facial expressions during an offline training phase. These pre-computed texture maps are then stored and rapidly blended during real-time animation, eliminating the need to generate textures on-the-fly and avoiding the deformation artifacts that plague real-time single-image generation.
Solution Approach 2:
The patent implements a dynamic blending system that smoothly transitions between pre-generated texture maps based on the current facial expression. The system dynamically adjusts the blending weights of different key expression textures to match the target expression, enabling realistic real-time animation without the artifacts associated with simple texture stretching.
Data Source
AI summary
A system and method for generating real-time facial animation is disclosed. The system relies upon pre-generating a series of key expression images from a single neutral image using a pre-trained generative adversarial neural network. The key expression images are used to generate a set of FACS expressions and associated textures which may be applied to a three-dimensional model to generate facial animation. The FACS expressions and textures may be provided to a mobile device to enable that mobile device to generate convincing three-dimensional avatars in real-time with convincing animation in a processor non-intensive way through a blending process using the pre-determined FACS expressions and textures.


