Speech-Driven Avatar Facial Features for Realistic Videoconferencing

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Facial features of avatars in videoconferences often do not correspond to the speech of users, leading to unrealistic representations.

Innovation Solution

A computing system modifies facial features of avatars based on speech transitions and vowel transitions within a speech signal, including changes in sound and acoustic resonances, to create a realistic representation of the user's speech.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If avatars are used in videoconferences to represent users, then data transmission requirements are reduced and user privacy is protected, but facial features of avatars do not correspond to speech of users resulting in unrealistic representation

Engineering Contradiction:
Improverealism of avatar representationVSAvoidcomplexity of speech-to-facial-feature mapping system
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The patent replaces traditional mechanical video transmission with a computational approach that processes speech signals and generates corresponding facial features. Instead of transmitting actual video frames, the system converts audio signals into facial movement parameters, substituting a complex video processing pipeline with a signal processing and synthesis system.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Solution Approach 2:

The system transforms the speech signal into modified facial feature parameters by analyzing acoustic characteristics and mapping them to corresponding mouth movements. The patent applies parameter transformation to convert raw audio data into synthesized facial animation parameters, enabling realistic avatar representation without actual video transmission.

Inventive Principle:
Principle #35Parameter changes

2Reliability

If facial features of avatars are modified based on speech transitions and vowel transitions, then realism of avatar representation is improved, but processing time and computational resources increase

Engineering Contradiction:
Improverealism of avatar representationVSAvoidprocessing time for speech signal analysis
Core Design Contradiction:
ReliabilityVSLoss of time

Solution Approach 1:

The patent segments the speech signal into distinct phonetic components including speech transitions and vowel transitions. By dividing the continuous speech stream into discrete phonetic units, the system can process and analyze each segment independently, reducing overall computational complexity while maintaining realistic facial feature synchronization.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The system performs preliminary analysis of speech transitions and vowel characteristics before generating final facial features. By pre-processing the speech signal to identify transition patterns and acoustic features, the system reduces real-time computational burden and enables more efficient avatar animation generation.

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS20250322843A1Modifying facial feature based on speech signal
Publication Date: 2025.10.16 GOOGLE LLC
  • US20250322843A1 patent drawing
  • US20250322843A1 patent drawing
  • US20250322843A1 patent drawing

AI summary

A computer-implemented method can include determining a speech transition within a speech signal, the speech transition including a change of sound; determining a mouth state based on the speech transition; determining a vowel transition during a vowel sound within the speech signal; and modifying a facial feature of an avatar based on the mouth state and the vowel transition.