Dynamic Avatar Generation for Audio-Only Meetings

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Online collaborative communication environments lack effective representation of meeting participants without a video stream, leading to less engaging and less personal interactions, as conventional systems rely on static images that do not animate or move in sync with speakers.

Innovation Solution

A system generates an avatar that mimics the facial features and lip movements of a meeting participant using machine learning, allowing the avatar to be injected into the data stream when the participant speaks, even without a video feed, to provide a dynamic and engaging representation.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Device complexity

If a static image is used to represent a meeting participant without video, then bandwidth is saved and device complexity is reduced, but engagement and personalization are worsened

Engineering Contradiction:
Improvesystem complexityVSAvoidengagement
Core Design Contradiction:
Device complexityVSAdaptability or versatility

Solution Approach 1:

The patent applies the dynamics principle by transforming the static image into a dynamic avatar that animates facial features and lip movements in sync with the speaker's audio. This allows the representation to adapt and respond to real-time speech without requiring actual video transmission, thus maintaining engagement while avoiding the complexity of video conferencing infrastructure.

Inventive Principle:
Principle #15Dynamics

2Quantity of substance

If audio only transmission is used, then bandwidth consumption is reduced, but the ability to put a face to the name is worsened

Engineering Contradiction:
ImprovebandwidthVSAvoidvisual information
Core Design Contradiction:
Quantity of substanceVSLoss of information

Solution Approach 1:

The patent applies the copying principle by creating a synthetic avatar copy that replicates the speaker's facial appearance and lip movements. This copy is generated from a static image and animated to match the audio stream, providing visual information that mimics the real person without transmitting actual video data, thus preserving visual presence while maintaining low bandwidth usage.

Inventive Principle:
Principle #26Copying

3Use of energy by moving object

If a static image is attached to represent a speaker, then system resources are saved, but the avatar does not animate or move in sync with the speaker

Engineering Contradiction:
Improveprocessing resourcesVSAvoidanimation accuracy
Core Design Contradiction:
Use of energy by moving objectVSReliability

Solution Approach 1:

The patent applies the feedback principle by synchronizing the avatar's lip movements with the audio stream in real-time. The system processes the audio signal and uses it as feedback to drive the animation of the avatar's facial features, ensuring that the avatar moves in sync with the speaker's speech. This feedback mechanism maintains animation accuracy while using computational resources efficiently.

Inventive Principle:
Principle #23Feedback

Data Source

PatentUS20240212248A1System and method for generating avatar of an active speaker in a meeting
Publication Date: 2024.06.27 RINGCENTRAL INC
  • US20240212248A1 patent drawing
  • US20240212248A1 patent drawing
  • US20240212248A1 patent drawing

AI summary

A method includes receiving a facial data associated with a participant user; generating an avatar data of the participant user based on the facial data using a first machine learning (ML) model; receiving data associated with facial movement; and training the generated avatar based on the data associated with facial movement using a second ML model to generate a trained avatar data, wherein the trained avatar data mimics appropriate facial movements associated with audio data.