Digital Avatar Speech Emotion Sequencing for Realistic Animation
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional systems determine a single emotional state for speech, failing to accurately represent changes in emotion throughout speech, leading to less realistic character or avatar animations.
Innovation Solution
Utilize machine learning models trained with multiple processes to determine probabilities for distributions of emotional states and sequences of emotional states, incorporating user feedback to optimize the models.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Device complexity
If a machine learning model determines only a single emotional state for speech, then the system complexity is reduced and processing is simplified, but the accuracy of representing actual emotional changes throughout speech deteriorates
Solution Approach 1:
The speech is divided into multiple temporal segments or frames, and each segment is assigned its own emotional state determination. This allows the system to capture emotional changes over time without requiring a single complex model to predict all emotions simultaneously, thus maintaining manageable system complexity while improving accuracy.
Solution Approach 2:
The system transitions from determining a static single emotional state to determining dynamic sequences of emotional states that change over time. This is achieved by processing speech in temporal segments and generating a sequence of emotional state predictions that reflect the dynamic nature of human emotion during speech.
2Measurement precision
If a machine learning model determines multiple emotional states for speech, then the accuracy of representing actual emotions improves, but the device complexity and processing requirements increase
Solution Approach 1:
By segmenting speech into temporal frames and determining emotional states for each frame independently or with limited context, the system avoids the exponential complexity of modeling all possible emotional combinations across the entire speech at once. This segmentation enables accurate multi-emotion representation while keeping computational complexity manageable.
Solution Approach 2:
The system determines emotional states for each temporal segment with sufficient detail to capture emotional changes, even if this means performing more computations than a single overall prediction would require. This partial action approach (processing segment by segment) achieves high accuracy without requiring the system to handle the full complexity of entire speech emotions simultaneously.
3Speed
If conventional systems use a single emotional state determination, then processing speed is maintained, but the realism of character animations deteriorates
Solution Approach 1:
The speech is processed in temporal segments that can be determined in parallel or in a streamlined sequential manner, maintaining processing speed. Each segment's emotional state is determined independently or with minimal inter-segment dependency, avoiding the need for complex global optimization that would slow processing while still producing realistic animated character expressions that change over time.
Solution Approach 2:
The system determines emotional states at periodic temporal intervals (frames) throughout the speech, creating a rhythm of emotional state updates that matches the natural pacing of speech and animation. This periodic determination maintains processing efficiency while generating realistic sequences of emotional expressions for character animation.
Data Source
AI summary
In various examples, determining emotional states for speech in conversational artificial intelligence (AI) and/or digital avatar systems and applications is descried herein. Systems and methods are disclosed that use one or more machine learning models to determine one or more emotional states associated with speech, where the machine learning model(s) may be trained using various processes. For instance, in some examples, the machine learning model(s) may be trained during a first training process to determine probabilities for distributions of values, where the distributions model different emotional states. For example, a distribution may include a first value for angry, a second value for happy, a third value for sad, and/or so forth. Additionally, or alternatively, in some examples, the machine learning model(s) may be trained during a second training process to more precisely determine the actual emotional states (and/or the probabilities) based on training data representing human feedback.


