Emotion-Aware Virtual Facial Expression Generation From Audio

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing software applications lack the ability to generate facial images with emotionally-aware facial expressions, limiting their interaction capabilities, particularly in applications like gaming and automated interactive responses.

Innovation Solution

An emotionally-aware digital content engine that utilizes a speech-to-text converter, emotion classifier, phoneme generator, and emotionally-aware content generator to identify emotions from audio input and generate corresponding facial expressions, incorporating them into digital content.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If traditional facial image generation methods are used, then the generation process is simple, but the facial expressions lack emotional awareness and interaction capability

Engineering Contradiction:
Improveemotional awareness capabilityVSAvoidsystem complexity
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The system segments the facial expression generation process into distinct functional modules: audio input processing, emotion classification, phoneme generation, and facial expression synthesis. Each module handles a specific aspect of emotional awareness, allowing the complex task to be divided into manageable components that can be developed and optimized independently while collectively achieving emotionally-aware facial expressions

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent introduces intermediate processing layers including emotion classifiers that translate audio into emotional states, and phoneme generators that bridge emotions to facial expressions. These intermediary components act as mediators that transform raw audio input into structured emotional representations, which then guide the generation of appropriate facial expressions, enabling emotional awareness without direct complex mapping

Inventive Principle:
Principle #24Intermediary (Mediator)

2Ease of operation

If emotional awareness is added to facial expressions, then user interaction capability is enhanced, but the processing time and computational resources increase

Engineering Contradiction:
Improveuser interaction capabilityVSAvoidprocessing time
Core Design Contradiction:
Ease of operationVSLoss of time

Solution Approach 1:

The system performs preliminary processing of audio input by extracting emotional features and generating phoneme representations in advance of actual facial expression generation. By pre-processing the audio signal to identify emotional states and key phonemic elements, the system prepares the necessary data structures and emotional context beforehand, reducing the computational burden during real-time expression synthesis and minimizing processing delays

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent implements dynamic adjustment of processing depth based on application requirements. The emotion classifier and phoneme generator can operate at different levels of detail depending on the specific interaction context, allowing the system to balance between comprehensive emotional analysis and faster processing when appropriate. This dynamic approach enables the system to adapt its computational intensity to match the actual needs of user interaction scenarios

Inventive Principle:
Principle #15Dynamics

Data Source

PatentUS12450806B2System and method for generating emotionally-aware virtual facial expressions
Publication Date: 2025.10.21 VERIZON PATENT & LICENSING INC
  • US12450806B2 patent drawing
  • US12450806B2 patent drawing
  • US12450806B2 patent drawing

AI summary

Techniques for generating emotionally-aware digital content are disclosed. In one embodiment, a method is disclosed comprising obtaining audio input, obtaining a textual representation of the audio input; using the textual representation of the audio input to identify an emotion corresponding to the audio input; generating an emotionally-aware facial representation in accordance with the textual representation and the identified emotion; using the emotionally-aware facial representation to generate one or more images comprising at least one facial expression corresponding to the identified emotion; and providing digital content comprising the one or more images.