TTS Vocal Characteristic Transfer via Neural Alignment

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current text-to-speech (TTS) systems face challenges in efficiently generating synthesized speech with diverse vocal characteristics without requiring extensive recording times and professional performers, limiting the ability to produce high-quality speech with varied emotional and accentual traits in real-time.

Innovation Solution

The system transfers vocal characteristics from a source speaker's audio data to a neural-network model trained on a target speaker's data, allowing the generation of synthesized speech with the target speaker's voice and the source speaker's vocal characteristics, using alignment and aggregation of linguistic units and vocal features to modify the audio output.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Manufacturing precision

If traditional TTS systems use extensive recording by professional performers to achieve diverse vocal characteristics, then speech quality is improved, but recording time and production complexity increase significantly

Engineering Contradiction:
Improvespeech qualityVSAvoidrecording time
Core Design Contradiction:
Manufacturing precisionVSLoss of time

Solution Approach 1:

The system creates vocal characteristic representations from source speech and applies them to target speaker models, effectively copying vocal traits without requiring extensive new recordings. This allows high-quality diverse speech generation by synthesizing vocal characteristics rather than recording them extensively.

Inventive Principle:
Principle #26Copying

Solution Approach 2:

The system modifies speech parameters by adjusting vocal characteristic representations (such as pitch, timbre, and prosody parameters) to generate diverse vocal styles from a single speaker model, eliminating the need for multiple professional performers and extensive recording sessions.

Inventive Principle:
Principle #35Parameter changes

2Adaptability or versatility

If traditional TTS systems record extensive speech data to capture varied emotional and accentual traits, then vocal diversity is improved, but data storage requirements and processing complexity increase

Engineering Contradiction:
Improvevocal diversityVSAvoiddata storage
Core Design Contradiction:
Adaptability or versatilityVSQuantity of substance

Solution Approach 1:

The system extracts essential vocal characteristics from source speech into compact representations (vocal characteristic representations) that capture emotional and accentual traits without storing the entire original speech corpus. This extraction process reduces data storage requirements while maintaining vocal diversity.

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

A single target speaker model can generate multiple vocal characteristics by applying different vocal characteristic representations, making the system universal and adaptable to diverse emotional and accentual requirements without requiring separate recordings for each scenario.

Inventive Principle:
Principle #6Universality (Multi-functionality)

3Reliability

If traditional TTS systems use multiple professional performers to achieve vocal variety, then emotional expression is improved, but production cost and complexity increase

Engineering Contradiction:
Improveemotional expressionVSAvoidproduction complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The system copies vocal characteristics from source speakers into structured representations that can be applied to any target speaker model, enabling reliable emotional expression without requiring multiple professional performers. This copying mechanism preserves emotional traits while simplifying production.

Inventive Principle:
Principle #26Copying

Solution Approach 2:

The system achieves different emotional expressions by modifying vocal characteristic parameters (pitch contours, energy levels, timing parameters) within a unified model framework, eliminating the need for multiple performers and reducing production complexity while maintaining emotional authenticity.

Inventive Principle:
Principle #35Parameter changes

Data Source

PatentUS11410684B1Text-to-speech (TTS) processing with transfer of vocal characteristics
Publication Date: 2022.08.09 AMAZON TECH INC
  • US11410684B1 patent drawing
  • US11410684B1 patent drawing
  • US11410684B1 patent drawing

AI summary

Audio data from a first, source speaker is received and processed to determine linguistic units and vocal characteristics corresponding to those linguistic units. The linguistic units may either be determined from received text data or may be determined from the audio data using automatic speech recognition. A model is trained using training data from a second, target speaker. The trained model concatenates the linguistic units with the vocal characteristics to produce output speech that has the “voice” of the target speaker and the vocal characteristics of the source speaker.