Digital Avatar Speech Emotion Sequencing for Realistic Animation

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional systems determine a single emotional state for speech, failing to accurately represent changes in emotion throughout speech, leading to less realistic character or avatar animations.

Innovation Solution

Utilize machine learning models trained with multiple processes to determine probabilities for distributions of emotional states and sequences of emotional states, incorporating user feedback to optimize the models.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Device complexity

If a machine learning model determines only a single emotional state for speech, then the system complexity is reduced and processing is simplified, but the accuracy of representing actual emotional changes throughout speech deteriorates

Engineering Contradiction:
Improvesystem complexityVSAvoidaccuracy of emotional state representation
Core Design Contradiction:
Device complexityVSMeasurement precision

Solution Approach 1:

The speech is divided into multiple temporal segments or frames, and each segment is assigned its own emotional state determination. This allows the system to capture emotional changes over time without requiring a single complex model to predict all emotions simultaneously, thus maintaining manageable system complexity while improving accuracy.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The system transitions from determining a static single emotional state to determining dynamic sequences of emotional states that change over time. This is achieved by processing speech in temporal segments and generating a sequence of emotional state predictions that reflect the dynamic nature of human emotion during speech.

Inventive Principle:
Principle #15Dynamics

2Measurement precision

If a machine learning model determines multiple emotional states for speech, then the accuracy of representing actual emotions improves, but the device complexity and processing requirements increase

Engineering Contradiction:
Improveaccuracy of emotional state representationVSAvoidsystem complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

By segmenting speech into temporal frames and determining emotional states for each frame independently or with limited context, the system avoids the exponential complexity of modeling all possible emotional combinations across the entire speech at once. This segmentation enables accurate multi-emotion representation while keeping computational complexity manageable.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The system determines emotional states for each temporal segment with sufficient detail to capture emotional changes, even if this means performing more computations than a single overall prediction would require. This partial action approach (processing segment by segment) achieves high accuracy without requiring the system to handle the full complexity of entire speech emotions simultaneously.

Inventive Principle:
Principle #16Partial or excessive action

3Speed

If conventional systems use a single emotional state determination, then processing speed is maintained, but the realism of character animations deteriorates

Engineering Contradiction:
Improveprocessing speedVSAvoidrealism of character animation
Core Design Contradiction:
SpeedVSReliability

Solution Approach 1:

The speech is processed in temporal segments that can be determined in parallel or in a streamlined sequential manner, maintaining processing speed. Each segment's emotional state is determined independently or with minimal inter-segment dependency, avoiding the need for complex global optimization that would slow processing while still producing realistic animated character expressions that change over time.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The system determines emotional states at periodic temporal intervals (frames) throughout the speech, creating a rhythm of emotional state updates that matches the natural pacing of speech and animation. This periodic determination maintains processing efficiency while generating realistic sequences of emotional expressions for character animation.

Inventive Principle:
Principle #19Periodic action

Data Source

PatentUS20250272901A1Determining emotional states for speech in digital avatar systems and applications
Publication Date: 2025.08.28 NVIDIA CORP
  • US20250272901A1 patent drawing
  • US20250272901A1 patent drawing
  • US20250272901A1 patent drawing

AI summary

In various examples, determining emotional states for speech in conversational artificial intelligence (AI) and/or digital avatar systems and applications is descried herein. Systems and methods are disclosed that use one or more machine learning models to determine one or more emotional states associated with speech, where the machine learning model(s) may be trained using various processes. For instance, in some examples, the machine learning model(s) may be trained during a first training process to determine probabilities for distributions of values, where the distributions model different emotional states. For example, a distribution may include a first value for angry, a second value for happy, a third value for sad, and/or so forth. Additionally, or alternatively, in some examples, the machine learning model(s) may be trained during a second training process to more precisely determine the actual emotional states (and/or the probabilities) based on training data representing human feedback.