Audio-Driven Synthetic Avatar for Video Calls
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
In video communication sessions, users who are speaking but do not want to be on camera can lead to a negative user experience for listeners, as their video feed is often replaced with a static image.
Innovation Solution
A computing system that receives audio from a client device, uses a facial feature model to estimate facial movement, generates a synthetic video of an avatar moving according to the estimated facial movement, and synchronizes this video with the audio to be presented to the other client device.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Ease of operation
If a user's video feed is replaced with a static image when they do not want to be on camera, then the user's privacy and comfort are protected, but the listener's user experience deteriorates due to lack of visual engagement
Solution Approach 1:
The system creates a synthetic video copy of the user's face using audio-driven facial movement estimation. Instead of showing the actual user video or a static image, the system generates a synthetic avatar that replicates facial expressions and movements based on audio input, providing a dynamic visual representation that engages listeners while protecting user privacy
Solution Approach 2:
The system introduces a synthetic video intermediary between the user's actual video feed and the listener's display. This intermediary processes audio input and generates corresponding facial movements, serving as a mediator that transforms audio signals into visual representations without exposing the user's actual video feed
2Reliability
If the user's actual video feed is transmitted, then visual engagement for listeners is improved, but the user's privacy and comfort are compromised when they do not want to be on camera
Solution Approach 1:
The system creates a synthetic copy of the user's facial movements driven by audio rather than transmitting the actual video feed. This copy captures essential visual information for engagement while completely separating the source video from the output, eliminating privacy exposure
Solution Approach 2:
The system replaces the mechanical video transmission system with an audio-driven synthesis system. Instead of transmitting actual video signals, the system uses audio processing and facial movement estimation to generate synthetic visual representations, substituting one transmission mechanism for another that inherently protects privacy
3Loss of energy
If a static image is used instead of video, then bandwidth consumption is reduced, but visual engagement and face-to-face interaction quality deteriorate
Solution Approach 1:
The system transforms the static image into a dynamic synthetic video that responds to audio input in real-time. The synthetic avatar's facial movements change dynamically based on speech content, providing visual engagement comparable to live video while maintaining lower bandwidth requirements than full video transmission
Solution Approach 2:
The system changes the parameter of visual representation from static to dynamically generated. By using audio-driven facial movement estimation, the system creates video-like engagement with controlled bandwidth consumption, adjusting the balance between visual quality and data transmission
Data Source
AI summary
In one aspect, an example method includes (i) receiving audio from a first client device engaged in a communication session with a second client device, the audio comprising one or more words spoken by a user of the first client device; (ii) using the audio and a facial feature model to estimate facial movement that corresponds to the one or more words spoken by the user; (iii) generating a synthetic video depicting an avatar of the user moving according to the estimated facial movement; and (iv) in response to generating the synthetic video, causing the second client device to present the synthetic video synchronized with the audio.


