Emotion-Aware Speech Synthesis Using Style Encoder and Token Sets

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current speech synthesis technologies struggle to produce synthesized speech that accurately reflects emotions, particularly in multi-language environments, resulting in unnatural and mechanical outputs.

Innovation Solution

An electronic device equipped with memory storing token sets for various emotions and processors that identify emotions in reference speech, generate style information, and synthesize speech using a style encoder and decoder, ensuring emotional reflection and naturalness.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Manufacturing precision

If speech synthesis technology uses deep learning to improve quality, then synthesis quality is improved, but emotional naturalness deteriorates

Engineering Contradiction:
Improvesynthesis qualityVSAvoidemotional naturalness
Core Design Contradiction:
Manufacturing precisionVSReliability

Solution Approach 1:

The speech synthesis system is segmented into multiple independent components: a style encoder that processes reference speech to extract emotional characteristics, a token set database that stores emotion-specific representations, and a decoder that generates synthesized speech. This segmentation allows each component to specialize in specific aspects (quality, emotion, generation) and work together to resolve the contradiction between synthesis quality and emotional naturalness.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

A style encoder is introduced as an intermediary component that bridges the gap between reference speech and the synthesis process. The style encoder extracts and encodes emotional characteristics from reference speech into style information, which then guides the decoder to generate emotionally natural synthesized speech while maintaining high synthesis quality through the deep learning framework.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Adaptability or versatility

If speech synthesis reflects strong emotions like anger or pleasantness, then emotional expression is improved, but naturalness deteriorates becoming mechanical

Engineering Contradiction:
Improveemotional expressionVSAvoidnaturalness
Core Design Contradiction:
Adaptability or versatilityVSReliability

Solution Approach 1:

The system changes key parameters of speech synthesis by introducing style information that controls multiple acoustic parameters simultaneously (pitch, energy, timing). By adjusting these parameters based on emotion-specific token sets and style encoding from reference speech, the system achieves natural emotional expression without mechanical artifacts, resolving the contradiction between emotional expression and naturalness.

Inventive Principle:
Principle #35Parameter changes

3Adaptability or versatility

If speech synthesis technology is applied in multi-language environments, then versatility is improved, but training data availability deteriorates

Engineering Contradiction:
Improvemulti-language capabilityVSAvoidtraining data availability
Core Design Contradiction:
Adaptability or versatilityVSQuantity of substance

Solution Approach 1:

The style encoder and token set framework are designed with universality to handle multiple languages. The style encoder processes reference speech in any language to extract emotional characteristics, and the token sets are organized to support multiple languages. This universal design allows the system to achieve multi-language capability without requiring separate training data for each language, resolving the contradiction between versatility and training data availability.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Data Source

PatentUS20250191572A1Electronic device for obtaining synthesized speech by considering emotion and control method therefor
Publication Date: 2025.06.12 SAMSUNG ELECTRONICS CO LTD
  • US20250191572A1 patent drawing
  • US20250191572A1 patent drawing
  • US20250191572A1 patent drawing

AI summary

An electronic device is provided. The electronic device includes memory storing one or more computer programs and a plurality of token sets corresponding to respective multiple emotions, and one or more processors communicatively coupled to the memory, wherein the one or more computer programs include computer-executable instructions that, when executed by the one or more processors individually or collectively, cause the electronic device to, based on receiving a reference speech, identify an emotion corresponding to the reference speech among the plurality of emotions, obtain a token set corresponding to the identified emotion from among the plurality of token sets stored in the memory, input information on the reference speech and the obtained token set into a style encoder and obtain style information for outputting a synthesized speech of the identified emotion, based on a text being input, input the text into a decoder obtained on the basis of the style information and obtain a synthesized speech corresponding to the text, and output the synthesized speech corresponding to the text.