Emotion-Aware Speech Synthesis Using Style Encoder and Token Sets
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current speech synthesis technologies struggle to produce synthesized speech that accurately reflects emotions, particularly in multi-language environments, resulting in unnatural and mechanical outputs.
Innovation Solution
An electronic device equipped with memory storing token sets for various emotions and processors that identify emotions in reference speech, generate style information, and synthesize speech using a style encoder and decoder, ensuring emotional reflection and naturalness.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Manufacturing precision
If speech synthesis technology uses deep learning to improve quality, then synthesis quality is improved, but emotional naturalness deteriorates
Solution Approach 1:
The speech synthesis system is segmented into multiple independent components: a style encoder that processes reference speech to extract emotional characteristics, a token set database that stores emotion-specific representations, and a decoder that generates synthesized speech. This segmentation allows each component to specialize in specific aspects (quality, emotion, generation) and work together to resolve the contradiction between synthesis quality and emotional naturalness.
Solution Approach 2:
A style encoder is introduced as an intermediary component that bridges the gap between reference speech and the synthesis process. The style encoder extracts and encodes emotional characteristics from reference speech into style information, which then guides the decoder to generate emotionally natural synthesized speech while maintaining high synthesis quality through the deep learning framework.
2Adaptability or versatility
If speech synthesis reflects strong emotions like anger or pleasantness, then emotional expression is improved, but naturalness deteriorates becoming mechanical
Solution Approach 1:
The system changes key parameters of speech synthesis by introducing style information that controls multiple acoustic parameters simultaneously (pitch, energy, timing). By adjusting these parameters based on emotion-specific token sets and style encoding from reference speech, the system achieves natural emotional expression without mechanical artifacts, resolving the contradiction between emotional expression and naturalness.
3Adaptability or versatility
If speech synthesis technology is applied in multi-language environments, then versatility is improved, but training data availability deteriorates
Solution Approach 1:
The style encoder and token set framework are designed with universality to handle multiple languages. The style encoder processes reference speech in any language to extract emotional characteristics, and the token sets are organized to support multiple languages. This universal design allows the system to achieve multi-language capability without requiring separate training data for each language, resolving the contradiction between versatility and training data availability.
Data Source
AI summary
An electronic device is provided. The electronic device includes memory storing one or more computer programs and a plurality of token sets corresponding to respective multiple emotions, and one or more processors communicatively coupled to the memory, wherein the one or more computer programs include computer-executable instructions that, when executed by the one or more processors individually or collectively, cause the electronic device to, based on receiving a reference speech, identify an emotion corresponding to the reference speech among the plurality of emotions, obtain a token set corresponding to the identified emotion from among the plurality of token sets stored in the memory, input information on the reference speech and the obtained token set into a style encoder and obtain style information for outputting a synthesized speech of the identified emotion, based on a text being input, input the text into a decoder obtained on the basis of the style information and obtain a synthesized speech corresponding to the text, and output the synthesized speech corresponding to the text.


