Real-time Avatar Generation Using Dynamic Texture Blending

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing methods for generating realistic facial animations from a single image often result in artifacts such as uncanny eyes, unusual mouth shapes, and texture deformation when applied to three-dimensional models, leading to less than excellent results.

Innovation Solution

A system utilizing a generative adversarial network (GAN) with dynamic textures, which includes training data of two-dimensional and fully modeled face images, generates key expression meshes and blendshape texture maps to create photorealistic avatars capable of real-time animation, addressing the challenges of facial texture generation and animation from a single image.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Ease of manufacture

If traditional single-image avatar generation methods are used, then the avatar can be created from minimal input, but the facial animations contain artifacts and deformations

Engineering Contradiction:
Improveavatar creation from single imageVSAvoidfacial texture accuracy
Core Design Contradiction:
Ease of manufactureVSManufacturing precision

Solution Approach 1:

The patent segments the facial animation problem into multiple key expressions (e.g., neutral, smile, frown, surprise) and generates specific texture maps for each expression. This segmentation allows the system to handle complex facial animations by breaking them down into manageable, pre-rendered components that can be blended together, avoiding the artifacts that occur with traditional single-image generation methods.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent performs preliminary action by pre-generating texture maps for key facial expressions during an offline training phase. These pre-computed texture maps are then stored and rapidly blended during real-time animation, eliminating the need to generate textures on-the-fly and avoiding the deformation artifacts that plague real-time single-image generation.

Inventive Principle:
Principle #10Preliminary action

2Manufacturing precision

If complex rigs and multi-camera systems are used, then facial texture accuracy is improved, but the system complexity and cost increase

Engineering Contradiction:
Improvefacial texture accuracyVSAvoidcapture system complexity
Core Design Contradiction:
Manufacturing precisionVSDevice complexity

Solution Approach 1:

The patent uses a trained neural network to copy realistic facial textures from a training dataset into the generated avatar. The neural network learns the statistical relationships between facial structures and textures from diverse examples, then synthesizes photorealistic textures for the target avatar without requiring complex physical capture equipment.

Inventive Principle:
Principle #26Copying

Solution Approach 2:

The patent replaces the mechanical multi-camera rig system with a computational approach using a trained neural network. Instead of using physical cameras and light rigs to capture facial textures, the system uses machine learning to synthesize realistic textures from a single input image, dramatically reducing hardware complexity while maintaining or improving texture quality.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

3Speed

If real-time animation is implemented, then the avatar responsiveness is improved, but texture deformation and artifacts increase

Engineering Contradiction:
Improveanimation responsivenessVSAvoidtexture quality
Core Design Contradiction:
SpeedVSManufacturing precision

Solution Approach 1:

The patent performs preliminary action by pre-generating texture maps for key facial expressions during an offline training phase. These pre-computed texture maps are then stored and rapidly blended during real-time animation, eliminating the need to generate textures on-the-fly and avoiding the deformation artifacts that plague real-time single-image generation.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent implements a dynamic blending system that smoothly transitions between pre-generated texture maps based on the current facial expression. The system dynamically adjusts the blending weights of different key expression textures to match the target expression, enabling realistic real-time animation without the artifacts associated with simple texture stretching.

Inventive Principle:
Principle #15Dynamics

Data Source

PatentUS10896535B2Real-time avatars using dynamic textures
Publication Date: 2021.01.19 PINSCREEN INC
  • US10896535B2 patent drawing
  • US10896535B2 patent drawing
  • US10896535B2 patent drawing

AI summary

A system and method for generating real-time facial animation is disclosed. The system relies upon pre-generating a series of key expression images from a single neutral image using a pre-trained generative adversarial neural network. The key expression images are used to generate a set of FACS expressions and associated textures which may be applied to a three-dimensional model to generate facial animation. The FACS expressions and textures may be provided to a mobile device to enable that mobile device to generate convincing three-dimensional avatars in real-time with convincing animation in a processor non-intensive way through a blending process using the pre-determined FACS expressions and textures.