Parallel Wave Generation for Speech Synthesis
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Traditional text-to-speech systems are hindered by the autoregressive nature of WaveNet, which makes them prohibitively slow at inference, requiring highly engineered inference kernels to generate high-fidelity speech in real time, and existing end-to-end speech synthesis methods rely on separate waveform synthesizers, leading to unstable training processes and suboptimal performance.
Innovation Solution
A novel parallel wave generation method using Gaussian inverse autoregressive flow for speech synthesis, which simplifies the distillation algorithm by minimizing a regularized KL divergence and enables fast end-to-end training with a fully convolutional text-to-wave neural architecture that conditions a parallel neural vocoder on learned hidden representations.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Manufacturing precision
If autoregressive generative models are used for waveform synthesis, then high-fidelity speech can be generated, but inference becomes prohibitively slow because each sample must be drawn sequentially
Solution Approach 1:
The speech generation process is segmented into two independent parallel streams: a mel-spectrogram generator and a waveform synthesizer. The mel-spectrogram is generated in parallel without autoregressive constraints, then fed to the waveform synthesizer which also operates in parallel using the mel-spectrogram as conditional input. This segmentation eliminates the sequential bottleneck while preserving high-fidelity output.
Solution Approach 2:
The mel-spectrogram serves as an intermediary representation between the text input and the final waveform output. Instead of generating waveforms directly in an autoregressive manner, the system first generates mel-spectrograms (which can be done in parallel) and then uses these as conditional inputs for waveform synthesis. This intermediary approach enables parallel processing while maintaining speech quality.
2Productivity
If end-to-end speech synthesis methods are used, then training efficiency is improved, but the training process becomes unstable and performance is suboptimal
Solution Approach 1:
The end-to-end system is segmented into two trainable modules: the mel-spectrogram generator and the waveform synthesizer. Each module can be trained independently with its own loss function and optimization strategy, allowing for stable gradient flow and avoiding the training instability that plishes monolithic end-to-end systems. The modular architecture enables progressive training and better convergence.
Solution Approach 2:
The mel-spectrogram generator is trained first to learn the mapping from text to mel-spectrogram representations. Once trained, it provides stable conditional inputs to the waveform synthesizer, which is then trained. This preliminary training of the first module establishes a reliable foundation for the second module, improving overall training stability and performance.
Data Source
AI summary
Described herein are embodiments of an end-to-end text-to-speech (TTS) system with parallel wave generation. In one or more embodiments, a Gaussian inverse autoregressive flow is distilled from an autoregressive WaveNet by minimizing a novel regularized Kullback-Leibler (KL) divergence between their highly-peaked output distributions. Embodiments of the methodology computes the KL divergence in a closed-form, which simplifies the training process and provides very efficient distillation. Embodiments of a novel text-to-wave neural architecture for speech synthesis are also described, which are fully convolutional and enable fast end-to-end training from scratch. These embodiments significantly outperform the previous pipeline that connects a text-to-spectrogram model to a separately trained WaveNet. Also, a parallel waveform synthesizer embodiment conditioned on the hidden representation in an embodiment of this end-to-end model were successfully distilled.


