Real-Time Digital Avatar Expression Transfer via Depth Sensing
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current technologies for transferring facial expressions from a live actor to a digital avatar in real-time either require extensive actor training, use physical markings, or produce unrealistic results, failing to accurately portray expressions and maintain actor anonymity.
Innovation Solution
A system that uses depth-sensing cameras to capture and process facial landmarks, transforming them into expression blendshape coefficients which are then applied to a digital avatar's face in real-time without revealing the actor's image, using a combination of machine learning and geometric transformations.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If motion capture technology is used to transfer facial expressions from actor to avatar, then expression accuracy is improved, but the process becomes time-intensive and requires actor training and physical markings
Solution Approach 1:
The patent extracts only the essential facial expression data from the actor's face using depth-sensing cameras and machine learning algorithms, rather than capturing the entire face in high resolution. This extraction approach achieves accurate expression transfer without requiring full motion capture procedures, physical markings, or extensive actor training, thus resolving the contradiction between accuracy and time efficiency
Solution Approach 2:
The patent replaces the mechanical motion capture system with a combination of depth-sensing cameras and machine learning algorithms. Instead of using physical tracking dots and head-mounted cameras, the system uses automated computer vision to detect facial landmarks and generate blendshape coefficients, eliminating the need for actor training and physical markings while maintaining real-time performance
2Productivity
If deep-fake face swapping is used for real-time expression transfer, then processing speed is improved, but visual artifacts and unrealistic results occur
Solution Approach 1:
The patent changes the parameter representation from direct pixel manipulation (deep-fake approach) to geometric parameter space using 3D facial landmarks and blendshape coefficients. This parameter transformation enables mathematically precise expression transfer that maintains visual realism while achieving real-time processing speeds, resolving the contradiction between speed and reliability
Solution Approach 2:
The patent introduces an intermediary mathematical representation (blendshape coefficients) between the actor's facial data and the avatar's final appearance. This intermediary layer ensures that expressions are transferred through controlled geometric transformations rather than direct image manipulation, preventing visual artifacts while maintaining real-time performance
3Device complexity
If 3D face reconstruction from RGB images is used, then setup complexity is reduced, but the ability to drive avatars of different faces is lost
Solution Approach 1:
The patent creates a universal expression representation system using blendshape coefficients that can drive any 3D avatar face regardless of its specific geometry. The system extracts facial expressions into a standardized mathematical form that is compatible with different avatar models, achieving both low setup complexity and high avatar compatibility through this universal intermediary representation
4Measurement precision
If high-resolution offline rendering is used for motion capture, then expression fidelity is improved, but real-time performance is lost
Solution Approach 1:
The patent applies partial action by capturing only the essential facial expression data needed for avatar animation rather than rendering the entire high-resolution actor face. Using depth-sensing cameras and landmark detection, the system extracts minimal necessary information (facial landmarks and blendshape coefficients) that drives the avatar's expression, achieving both high fidelity and real-time performance by doing only what is necessary
Data Source
AI summary
Embodiments of the invention are directed toward methods and computer graphics systems that can capture from a plurality of depth-sensing digital cameras, a series of images of an actor's face, and, in real time, without giving the actor special training and without using physical face markings, (1) extract from each of the captured images a set of facial landmarks of the actor's face; (2) transform those facial landmarks into a vector of expression blendshape coefficients that each correspond to a component expression identified in the actor's face; and then (3) output the vector of expression blendshape coefficients to a graphics rendering engine, where the vector can be applied to corresponding expression blendshapes associated with a digital avatar's face, thereby enabling the live actor's facial expressions to be transferred mathematically and rendered to the digital avatar's face in real time without revealing an image of the actor's face.


