Dynamic Avatar Generation for Audio-Only Meetings
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Online collaborative communication environments lack effective representation of meeting participants without a video stream, leading to less engaging and less personal interactions, as conventional systems rely on static images that do not animate or move in sync with speakers.
Innovation Solution
A system generates an avatar that mimics the facial features and lip movements of a meeting participant using machine learning, allowing the avatar to be injected into the data stream when the participant speaks, even without a video feed, to provide a dynamic and engaging representation.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Device complexity
If a static image is used to represent a meeting participant without video, then bandwidth is saved and device complexity is reduced, but engagement and personalization are worsened
Solution Approach 1:
The patent applies the dynamics principle by transforming the static image into a dynamic avatar that animates facial features and lip movements in sync with the speaker's audio. This allows the representation to adapt and respond to real-time speech without requiring actual video transmission, thus maintaining engagement while avoiding the complexity of video conferencing infrastructure.
2Quantity of substance
If audio only transmission is used, then bandwidth consumption is reduced, but the ability to put a face to the name is worsened
Solution Approach 1:
The patent applies the copying principle by creating a synthetic avatar copy that replicates the speaker's facial appearance and lip movements. This copy is generated from a static image and animated to match the audio stream, providing visual information that mimics the real person without transmitting actual video data, thus preserving visual presence while maintaining low bandwidth usage.
3Use of energy by moving object
If a static image is attached to represent a speaker, then system resources are saved, but the avatar does not animate or move in sync with the speaker
Solution Approach 1:
The patent applies the feedback principle by synchronizing the avatar's lip movements with the audio stream in real-time. The system processes the audio signal and uses it as feedback to drive the animation of the avatar's facial features, ensuring that the avatar moves in sync with the speaker's speech. This feedback mechanism maintains animation accuracy while using computational resources efficiently.
Data Source
AI summary
A method includes receiving a facial data associated with a participant user; generating an avatar data of the participant user based on the facial data using a first machine learning (ML) model; receiving data associated with facial movement; and training the generated avatar based on the data associated with facial movement using a second ML model to generate a trained avatar data, wherein the trained avatar data mimics appropriate facial movements associated with audio data.


