Emotion Encoder Voice Conversion for TTS Dataset Generation

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing text-to-speech (TTS) models struggle to generate diverse styles of speech, including multiple genders and emotions, due to the high cost and complexity of acquiring large multi-speaker multi-emotion datasets.

Innovation Solution

The implementation of a voice conversion-based method that uses an emotion encoder to generate a multi-speaker multi-emotion dataset by converting the gender style of input speech while retaining its emotional style and linguistic content.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If a multi-speaker multi-emotion dataset is used to train a TTS model to generate diverse speech styles, then the speech style diversity is improved, but the data acquisition cost and complexity increase

Engineering Contradiction:
Improvespeech style diversityVSAvoiddata acquisition complexity
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The patent uses voice conversion technology to copy and transform speech styles from a single-speaker dataset into multiple speaker identities. The system learns speaker-specific characteristics from target speeches and applies them to convert the single speaker's emotional expressions into multiple diverse speech styles, eliminating the need to collect actual multi-speaker data.

Inventive Principle:
Principle #26Copying

Solution Approach 2:

The patent changes the parameters of speech generation by introducing speaker embedding vectors that control speaker identity and emotion embeddings that control emotional expression. By adjusting these parameters during inference, the system can generate diverse speech styles from a single trained model without requiring complex multi-speaker training data.

Inventive Principle:
Principle #35Parameter changes

2Manufacturing precision

If a large multi-speaker multi-emotion dataset is acquired to train a high-quality TTS model, then the speech quality is improved, but the data generation cost increases

Engineering Contradiction:
Improvespeech qualityVSAvoiddata volume
Core Design Contradiction:
Manufacturing precisionVSQuantity of substance

Solution Approach 1:

The system creates high-quality multi-speaker speech data by copying and transforming a single speaker's high-quality emotional expressions through learned speaker characteristics. This approach maintains speech quality while avoiding the costly data collection process for multiple speakers and emotions.

Inventive Principle:
Principle #26Copying

Solution Approach 2:

The patent extracts speaker-specific characteristics and emotional expressions separately from the input speech. The style encoder extracts speaker identity features while the emotion encoder extracts emotional features, allowing independent control and combination to generate diverse high-quality speech without needing large datasets.

Inventive Principle:
Principle #2Taking out (Extraction)

3Device complexity

If existing TTS methods are used to generate speech, then the model simplicity is maintained, but the emotional expression capability is limited

Engineering Contradiction:
Improvemodel simplicityVSAvoidemotional expression capability
Core Design Contradiction:
Device complexityVSAdaptability or versatility

Solution Approach 1:

The patent segments the speech generation process into distinct components: a style encoder for speaker identity, an emotion encoder for emotional expression, and a speech generator that combines them. This modular architecture maintains relative model simplicity while significantly enhancing emotional expression capability through separate specialized encoders.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent creates a universal TTS model that can generate multiple speaker identities and emotional expressions through a single architecture. The model uses speaker embeddings and emotion embeddings as controllable parameters, allowing one model to perform multiple speech style generations without requiring separate models for each speaker or emotion.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Data Source

PatentUS20250166602A1Systems and methods for speech generation by emotional voice conversion
Publication Date: 2025.05.22 DATUM POINT LABS INC
  • US20250166602A1 patent drawing
  • US20250166602A1 patent drawing
  • US20250166602A1 patent drawing

AI summary

Embodiments described herein include voice conversion (VC) based Emotion Data generation. Embodiments described herein may generate a multi-speaker multi-emotion dataset by changing the gender style of the input speech while retaining its emotion style and linguistic contents. For example, a single-speaker multi-emotion dataset may be used as the input speech and a multi-speaker single emotion dataset may be the target speech. The generated data may be used as the training data for a text to speech (TTS) model so that it can generate speeches with diverse styles of emotions and speakers. To generate a multi-speaker multi-emotion dataset, embodiments herein add an emotion encoder to a VC model and use acoustic properties to preserve the emotion speech style of the input speech while changing just the gender style to the target gender style.