Audio-Driven Synthetic Avatar for Video Calls

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

In video communication sessions, users who are speaking but do not want to be on camera can lead to a negative user experience for listeners, as their video feed is often replaced with a static image.

Innovation Solution

A computing system that receives audio from a client device, uses a facial feature model to estimate facial movement, generates a synthetic video of an avatar moving according to the estimated facial movement, and synchronizes this video with the audio to be presented to the other client device.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Ease of operation

If a user's video feed is replaced with a static image when they do not want to be on camera, then the user's privacy and comfort are protected, but the listener's user experience deteriorates due to lack of visual engagement

Engineering Contradiction:
Improveuser comfortVSAvoiduser experience
Core Design Contradiction:
Ease of operationVSReliability

Solution Approach 1:

The system creates a synthetic video copy of the user's face using audio-driven facial movement estimation. Instead of showing the actual user video or a static image, the system generates a synthetic avatar that replicates facial expressions and movements based on audio input, providing a dynamic visual representation that engages listeners while protecting user privacy

Inventive Principle:
Principle #26Copying

Solution Approach 2:

The system introduces a synthetic video intermediary between the user's actual video feed and the listener's display. This intermediary processes audio input and generates corresponding facial movements, serving as a mediator that transforms audio signals into visual representations without exposing the user's actual video feed

Inventive Principle:
Principle #24Intermediary (Mediator)

2Reliability

If the user's actual video feed is transmitted, then visual engagement for listeners is improved, but the user's privacy and comfort are compromised when they do not want to be on camera

Engineering Contradiction:
Improveuser experienceVSAvoidprivacy exposure
Core Design Contradiction:
ReliabilityVSObject-affected harmful factors

Solution Approach 1:

The system creates a synthetic copy of the user's facial movements driven by audio rather than transmitting the actual video feed. This copy captures essential visual information for engagement while completely separating the source video from the output, eliminating privacy exposure

Inventive Principle:
Principle #26Copying

Solution Approach 2:

The system replaces the mechanical video transmission system with an audio-driven synthesis system. Instead of transmitting actual video signals, the system uses audio processing and facial movement estimation to generate synthetic visual representations, substituting one transmission mechanism for another that inherently protects privacy

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

3Loss of energy

If a static image is used instead of video, then bandwidth consumption is reduced, but visual engagement and face-to-face interaction quality deteriorate

Engineering Contradiction:
Improvebandwidth consumptionVSAvoidvisual engagement
Core Design Contradiction:
Loss of energyVSReliability

Solution Approach 1:

The system transforms the static image into a dynamic synthetic video that responds to audio input in real-time. The synthetic avatar's facial movements change dynamically based on speech content, providing visual engagement comparable to live video while maintaining lower bandwidth requirements than full video transmission

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

The system changes the parameter of visual representation from static to dynamically generated. By using audio-driven facial movement estimation, the system creates video-like engagement with controlled bandwidth consumption, adjusting the balance between visual quality and data transmission

Inventive Principle:
Principle #35Parameter changes

Data Source

PatentUS20250184450A1Generating a User Avatar for Video Communications
Publication Date: 2025.06.05 ROKU INC
  • US20250184450A1 patent drawing
  • US20250184450A1 patent drawing
  • US20250184450A1 patent drawing

AI summary

In one aspect, an example method includes (i) receiving audio from a first client device engaged in a communication session with a second client device, the audio comprising one or more words spoken by a user of the first client device; (ii) using the audio and a facial feature model to estimate facial movement that corresponds to the one or more words spoken by the user; (iii) generating a synthetic video depicting an avatar of the user moving according to the estimated facial movement; and (iv) in response to generating the synthetic video, causing the second client device to present the synthetic video synchronized with the audio.