Multi-Speaker Neural TTS with Trainable Speaker Embeddings
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Developing text-to-speech (TTS) systems that support multiple voices is labor-intensive and requires significantly more data and development effort compared to single-voice systems, as most existing systems rely on distinct speech databases or model parameters for each speaker.
Innovation Solution
The development of all-neural multi-speaker TTS systems that share a vast majority of parameters across different voices, utilizing trainable speaker embeddings to enable speech synthesis from multiple voices with significantly less data per speaker, by incorporating these embeddings into architectures like Deep Voice 2 and Tacotron.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If distinct speech databases or model parameters are used for each speaker, then multiple speaker voices are supported, but development effort and data requirements increase significantly
Solution Approach 1:
The patent merges multiple speaker-specific databases and model parameters into a unified neural network model with shared parameters. The system uses a single speech database that contains recordings from multiple speakers, and a unified acoustic model that learns speaker-specific characteristics through trainable speaker embeddings rather than maintaining separate models for each speaker. This consolidation directly reduces development effort while maintaining multi-speaker support.
Solution Approach 2:
The unified neural network model is designed to handle multiple speakers universally through speaker embedding vectors that can represent any speaker. The model architecture includes shared acoustic model parameters that work across all speakers, making the system multi-functional rather than requiring separate specialized models. This universal approach eliminates the need for separate development processes for each speaker.
2Adaptability or versatility
If distinct speech databases are used for each speaker, then multiple speaker voices are supported, but data requirements increase significantly
Solution Approach 1:
The patent combines multiple speaker-specific datasets into a single unified speech database containing recordings from hundreds of different speakers. This merged database is then processed by the unified neural network model using speaker embedding vectors that encode speaker identity. By merging the data rather than maintaining separate databases, the system reduces overall data requirements while supporting multiple speakers.
Solution Approach 2:
The system transforms speaker-specific data requirements into speaker-agnostic learning through parameter changes. Instead of requiring extensive data per speaker, the model uses trainable speaker embedding vectors that learn compact representations of speaker characteristics from limited data. This parameter transformation allows the system to generalize across speakers with much smaller data requirements per speaker.
3Device complexity
If a unified neural network model with shared parameters is used, then development effort is reduced, but the system must learn effectively from small amounts of data spread among many speakers
Solution Approach 1:
The patent introduces speaker embedding vectors as intermediary representations that bridge the gap between unified model parameters and speaker-specific characteristics. These embedding vectors act as mediators that allow the shared acoustic model to adapt to different speakers without requiring speaker-specific training. The speaker embeddings encode speaker identity information in a compact form that the unified model can process, enabling effective learning from limited multi-speaker data.
Solution Approach 2:
The system performs preliminary encoding of speaker characteristics into speaker embedding vectors before the main speech synthesis process. This preliminary action of creating compact speaker representations allows the unified neural network to quickly adapt to new speakers with minimal data. By pre-processing speaker identity information into embeddings, the system prepares the data in a form that maximizes learning effectiveness from small datasets.
Data Source
AI summary
Described herein are systems and methods for augmenting neural speech synthesis networks with low-dimensional trainable speaker embeddings in order to generate speech from different voices from a single model. As a starting point for multi-speaker experiments, improved single-speaker model embodiments, which may be referred to generally as Deep Voice 2 embodiments, were developed, as well as a post-processing neural vocoder for Tacotron (a neural character-to-spectrogram model). New techniques for multi-speaker speech synthesis were performed for both Deep Voice 2 and Tacotron embodiments on two multi-speaker TTS datasets—showing that neural text-to-speech systems can learn hundreds of unique voices from twenty-five minutes of audio per speaker.


