Expressive Neural TTS Training Using Expressivity-Scored Data

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing text-to-speech (TTS) synthesis systems lack the ability to generate expressive speech that conveys emotional information and sounds natural and human-like, necessitating improved training methods to enhance vocal expressiveness.

Innovation Solution

A neural network-based TTS system is trained using a two-phase training process, utilizing a first sub-dataset with lower expressivity scores followed by a second sub-dataset with higher expressivity scores to generate expressive speech, and an expressivity score calculation method to evaluate audio samples accurately.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If traditional TTS training methods are used, then the system can generate basic speech, but the speech lacks vocal expressiveness and emotional information

Engineering Contradiction:
Improvevocal expressivenessVSAvoidtraining process complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The training dataset is segmented into multiple sub-datasets with different expressivity score ranges. The training process is divided into phases, where each phase uses a specific sub-dataset. This segmentation allows the system to progressively learn expressive speech patterns without overwhelming complexity in a single training run.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The expressivity score serves as a key parameter for selecting and weighting training samples. By changing the parameter range of expressivity scores used in different training phases (from lower to higher ranges), the system adapts its learning focus to progressively improve vocal expressiveness while managing training complexity.

Inventive Principle:
Principle #35Parameter changes

2Reliability

If a single training dataset is used, then the training process is simple, but the generated speech does not achieve high expressiveness

Engineering Contradiction:
Improvespeech expressivenessVSAvoidtraining efficiency
Core Design Contradiction:
ReliabilityVSProductivity

Solution Approach 1:

The expressivity score is calculated for all training samples beforehand, and samples are pre-organized into sub-datasets based on their expressivity score ranges. This preliminary action enables efficient sampling during training without requiring complex real-time evaluations, thus maintaining training efficiency while achieving high expressiveness.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

Instead of using the entire training dataset uniformly, the method selectively uses specific portions (sub-datasets with higher expressivity scores) in later training phases. This partial action focuses computational resources on the most expressive samples, improving speech expressiveness without proportionally increasing training cost.

Inventive Principle:
Principle #16Partial or excessive action

3Reliability

If uniform sampling from training data is used, then training is computationally efficient, but the system cannot capture highly expressive speech patterns

Engineering Contradiction:
Improveemotional information conveyanceVSAvoiddata selection process
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The method changes the sampling parameter from uniform distribution to expressivity-score-based distribution. By adjusting the expressivity score range parameter across training phases (from lower to higher ranges), the system captures increasingly expressive speech patterns while using a relatively simple scoring mechanism rather than complex data selection processes.

Inventive Principle:
Principle #35Parameter changes

Data Source

PatentUS12586561B2Text-to-speech synthesis method and system, a method of training a text-to-speech synthesis system, and a method of calculating an expressivity score
Publication Date: 2026.03.24 SPOTIFY
  • US12586561B2 patent drawing
  • US12586561B2 patent drawing
  • US12586561B2 patent drawing

AI summary

A method includes receiving text and inputting the received text in a prediction network. The method further includes generating, using the prediction network, speech data. The prediction network comprises a neural network that is trained to generate expressive speech data from text. The neural network is trained by: receiving a first training dataset comprising audio data and corresponding text data; acquiring a respective expressivity score for each audio sample of the audio data; selecting, from the first training dataset, a first subset of training data based on the respective expressivity scores of the audio data in the first training dataset; generating, for the first subset of training data, prediction audio data for the corresponding text data; and comparing the prediction audio data to the audio data of the first subset of training data.