3D Avatar Rendering via Audio-Driven Viseme Synchronization
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current technologies face challenges in generating realistic 3D avatar representations in artificial reality applications, as they struggle to accurately capture complex facial expressions and synchronize them with vocal outputs in real-time, often resulting in computationally exhaustive processes and stylized avatar shapes that do not accurately reflect human facial movements.
Innovation Solution
A method and system that predict phonemes from an audio stream, translate them into visemes, and combine these with 3D model blendshapes and weights determined from image data to create a synchronized 3D avatar representation, using iterative landmark determination and synchronization techniques to enhance facial expression realism.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If current technologies are used to generate 3D avatar representations, then the process can be implemented, but the computational complexity becomes exhaustive and the facial expressions fail to accurately reflect human movements
Solution Approach 1:
The patent transforms the complex continuous problem of facial expression capture into a discrete parameter-based system by identifying key facial landmarks and representing expressions through specific parameters. This reduces computational complexity while maintaining accuracy by focusing on critical facial movement parameters rather than processing all pixel data.
Solution Approach 2:
The patent replaces traditional mechanical computer vision approaches with a neural network-based system that learns facial expression patterns directly from data. This substitution enables more accurate capture of human facial movements while optimizing computational efficiency through learned representations rather than exhaustive algorithmic processing.
2Reliability
If phonemes are predicted from audio stream and combined with 3D model blendshapes, then realistic facial movements can be generated, but the processing time increases
Solution Approach 1:
The patent performs preliminary processing by pre-identifying facial landmarks and pre-processing audio data into phoneme predictions before the actual avatar rendering is needed. This advance preparation reduces the critical path processing time during real-time avatar generation while maintaining synchronization accuracy between facial movements and vocal output.
Data Source
AI summary
Disclosed herein includes a system, a method, and a non-transitory computer readable medium for rendering a three-dimensional (3D) model of an avatar according to an audio stream including a vocal output of a person and image data capturing a face of the person. In one aspect, phonemes of the vocal output are predicted according to the audio stream, and the predicted phonemes of the vocal output are translated into visemes. In one aspect, a plurality of blendshapes and corresponding weights are determined, according to the corresponding image data of the face, to form the 3D model of the avatar of the person. The visemes may be combined with the 3D model of the avatar to form a 3D representation of the avatar, by synchronizing the visemes with the 3D model of the avatar in time.


