Multi-Speaker Neural TTS Synthesis with Speaker Latent Space
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Neural Text-to-Speech (TTS) systems face challenges in generalization, particularly when synthesizing out-of-domain texts, leading to unnatural speech with issues like wrong pronunciation, strange prosody, and limited speaker similarity.
Innovation Solution
A multi-speaker neural TTS system is developed, utilizing a well-designed multi-speaker corpus set with wide content coverage and speaker variety, and incorporating speaker latent space information to enhance speech generation and adaptation to target speakers.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Manufacturing precision
If traditional TTS techniques are used, then pronunciation and prosody can be controlled through well-designed frontend linguistic features, but the synthesized speech lacks naturalness
Solution Approach 1:
The patent replaces traditional mechanical rule-based TTS systems with a neural network-based system. The neural TTS model learns pronunciation and prosody patterns automatically from data, substituting the manual rule-based approach with a data-driven neural network that generates more natural speech while maintaining pronunciation accuracy.
Solution Approach 2:
The patent transforms the TTS system from using discrete linguistic features to using continuous acoustic features as intermediate representation. This parameter change allows the system to maintain pronunciation control while achieving naturalness, as the neural network can learn complex mappings from text to acoustic features that capture both articulation and prosody.
2Productivity
If neural TTS systems are trained on limited domain data, then training efficiency is improved, but generalization ability deteriorates when synthesizing out-of-domain texts
Solution Approach 1:
The patent makes the TTS system universal by enabling it to handle both in-domain and out-of-domain texts effectively. The neural TTS model is trained to learn general speech patterns that can be applied across different domains, allowing it to maintain high quality synthesis even when encountering texts outside its training domain, thus achieving multi-functionality.
Solution Approach 2:
The patent performs preliminary training on diverse multi-domain corpus data before deployment. By pre-training on wide-coverage data that includes various domains, styles, and speakers, the system acquires robust generalization capabilities upfront, enabling it to handle out-of-domain texts without requiring additional domain-specific training.
3Reliability
If speaker-specific training data is collected for each target speaker, then speaker similarity is improved, but data collection time and system complexity increase
Solution Approach 1:
The patent enables the system to adapt to new speakers without requiring extensive manual data collection. The neural TTS model can perform self-adaptation using minimal speaker data or even without additional training, by leveraging its pre-learned speech generation capabilities and adjusting to the target speaker's characteristics through efficient fine-tuning or speaker embedding techniques.
Solution Approach 2:
The patent uses speaker embeddings to create a compact representation of speaker characteristics that can be copied and applied across different synthesis tasks. Instead of collecting and storing large amounts of speaker-specific data, the system extracts essential speaker features into compact embeddings that capture speaker identity and style, enabling efficient speaker adaptation.
4Manufacturing precision
If end-to-end neural TTS structure is adopted, then joint optimization of pronunciation and prosody is achieved, but the system requires large amounts of text-speech training data pairs
Solution Approach 1:
The patent uses a composite training approach that combines multiple types of training data and loss functions. The system integrates pronunciation-focused losses, prosody-focused losses, and naturalness-focused losses in a composite objective function, allowing joint optimization while being more data-efficient by leveraging multiple supervisory signals from diverse data sources.
Data Source
AI summary
A method for generating speech through multi-speaker neural text-to-speech (TTS) synthesis is provided. A text input may be received (1410). Speaker latent space information of a target speaker may be provided through at least one speaker model (1420). At least one acoustic feature may be predicted through an acoustic feature predictor based on the text input and the speaker latent space information (1430). A speech waveform corresponding to the text input may be generated through a neural vocoder based on the at least one acoustic feature and the speaker latent space information (1440).


