Emotion-Targeted Speech Synthesis via Candidate Selection
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current computer speech synthesis systems lack the ability to generate voice with emotional content, resulting in unnatural sounding speech that degrades the user experience.
Innovation Solution
The system generates a plurality of emotionally diverse candidate speech segments for a given message using crowd-sourcing and machine learning algorithms, which are then selected based on predetermined emotion types, allowing for the inclusion of emotional content in speech output.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Ease of manufacture
If straight text-to-speech conversion is used, then the speech generation is simple and efficient, but the emotional content is absent making the voice sound unnatural
Solution Approach 1:
The speech generation process is segmented into multiple candidate speech segments, each representing different emotional interpretations of the same text. This segmentation allows the system to explore multiple emotional possibilities without complicating the overall generation pipeline, resolving the contradiction between simplicity and naturalness.
Solution Approach 2:
Multiple candidate speech segments with different emotional contents are generated in advance as preliminary options. This preliminary action enables the system to pre-compute diverse emotional expressions, which are then selected based on the target emotion type, maintaining efficiency while improving naturalness.
2Reliability
If multiple candidate speech segments are generated to include emotional content, then the naturalness and emotional expression are improved, but the system complexity increases
Solution Approach 1:
Multiple candidate speech segments are generated in advance and stored, eliminating the need for complex real-time emotional synthesis. This preliminary preparation reduces online computational complexity while maintaining high emotional content accuracy during actual speech generation.
Solution Approach 2:
Instead of generating entirely new speech segments with emotional content, the system creates multiple candidate versions (copies) of the same text with different emotional characteristics. This copying approach maintains system simplicity while providing diverse emotional expressions for selection.
3Adaptability or versatility
If crowd-sourcing is used to generate candidate speech segments, then the diversity of emotional content is enhanced, but the time and resources required for generation increase
Solution Approach 1:
Candidate speech segments with diverse emotional content are generated in advance through crowd-sourcing and stored for future use. This preliminary generation eliminates the need for real-time crowd-sourcing, reducing time loss during actual speech generation while maintaining emotional diversity.
Solution Approach 2:
The system creates multiple candidate copies of speech segments with different emotional interpretations, collected through crowd-sourcing. These pre-generated copies provide diverse emotional content without requiring ongoing crowd-sourcing efforts, balancing versatility with time efficiency.
Data Source
AI summary
Techniques for generating voice with predetermined emotion type. In an aspect, semantic content and emotion type are separately specified for a speech segment to be generated. A candidate generation module generates a plurality of emotionally diverse candidate speech segments, wherein each candidate has the specified semantic content. A candidate selection module identifies an optimal candidate from amongst the plurality of candidate speech segments, wherein the optimal candidate most closely corresponds to the predetermined emotion type. In further aspects, crowd-sourcing techniques may be applied to generate the plurality of speech output candidates associated with a given semantic content, and machine-learning techniques may be applied to derive parameters for a real-time algorithm for the candidate selection module.


