Multimodal Emotion-Aware Speech Generation for Conversational AI
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional systems for animating characters in conversational AI applications inaccurately determine emotional states based on text alone, failing to account for variations in emotional expression and voice characteristics, leading to unnatural user experiences.
Innovation Solution
Utilizing machine learning models that incorporate user information, character information, visual information, audio information, and dialogue history to determine emotion attributes and generate speech that accurately reflects emotional states.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Ease of operation
If only text of speech is used to determine emotional states, then the system is simple to operate, but the emotional state determination accuracy deteriorates
Solution Approach 1:
The patent combines multiple input modalities (text, audio characteristics, visual information) into a unified emotional state determination system. The machine learning model integrates these diverse data sources to accurately determine emotional states, resolving the contradiction between system simplicity and determination accuracy by merging multiple information streams into a single comprehensive analysis process.
Solution Approach 2:
The system employs a multi-functional machine learning model that processes various types of input data (textual, auditory, visual) through a single unified framework. This universal model handles multiple functions including emotion detection, voice characteristic analysis, and speech generation, eliminating the need for separate systems for each function while maintaining high accuracy.
2Ease of manufacture
If only set emotional states are used for animated characters, then the system is easy to implement, but the range of emotional expression deteriorates
Solution Approach 1:
The patent transitions from static, predefined emotional states to dynamic, continuously variable emotional expressions. The machine learning model generates emotional states as continuous outputs that can vary in intensity and combination, allowing characters to express nuanced emotions like 'somewhat happy' or 'very happy' rather than being limited to discrete categories. This dynamic approach maintains ease of implementation through automated generation while dramatically expanding emotional expression range.
Solution Approach 2:
The system changes the parameter representation of emotional states from discrete categorical values to continuous multi-dimensional parameters. Instead of selecting from fixed emotional categories, the model outputs continuous values that can be combined and weighted to create any emotional state along a spectrum, enabling fine-grained emotional control while keeping the system implementation straightforward through parameter-based control.
3Measurement precision
If additional inputs (user information, character information, visual information, audio information, dialogue history) are used to determine emotional states, then the emotional state determination accuracy is improved, but the device complexity increases
Solution Approach 1:
The patent employs a universal machine learning model that handles multiple input types (user information, character information, visual information, audio information, dialogue history) through a single integrated framework. This multi-functional model processes diverse data sources without requiring separate processing systems for each input type, thereby improving emotional state determination accuracy while minimizing the increase in device complexity through consolidated architecture.
Solution Approach 2:
The machine learning model performs self-service by automatically integrating and processing multiple input modalities without requiring complex external coordination. The model autonomously weighs and combines different information sources (text, audio, visual, dialogue history) to determine emotional states, reducing the need for complex system orchestration and management infrastructure while maintaining high determination accuracy.
Data Source
AI summary
In various examples, expressing emotion in speech for conversational AI systems and applications is described herein. Systems and methods are disclosed that use one or more machine learning models (e.g., one or more language models) to determine one or more attributes associated with speech, such as one or more emotion attributes and/or one or more response attributes, and then use the attribute(s) to generate the speech that expresses emotion. In some examples, the machine learning model(s) may use various types of information to determine the attribute(s), such as user information, character information, a dialogue history, a current prompt, and/or so forth. For instance, using the information, the machine learning model(s) may determine one or more tags associated with emotional states and/or voice characteristics, where the tag(s) is then used to generate the speech in a voice that relates to the emotion.


