Parallel Wave Generation for Speech Synthesis

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Traditional text-to-speech systems are hindered by the autoregressive nature of WaveNet, which makes them prohibitively slow at inference, requiring highly engineered inference kernels to generate high-fidelity speech in real time, and existing end-to-end speech synthesis methods rely on separate waveform synthesizers, leading to unstable training processes and suboptimal performance.

Innovation Solution

A novel parallel wave generation method using Gaussian inverse autoregressive flow for speech synthesis, which simplifies the distillation algorithm by minimizing a regularized KL divergence and enables fast end-to-end training with a fully convolutional text-to-wave neural architecture that conditions a parallel neural vocoder on learned hidden representations.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Manufacturing precision

If autoregressive generative models are used for waveform synthesis, then high-fidelity speech can be generated, but inference becomes prohibitively slow because each sample must be drawn sequentially

Engineering Contradiction:
Improvespeech fidelityVSAvoidinference speed
Core Design Contradiction:
Manufacturing precisionVSProductivity

Solution Approach 1:

The speech generation process is segmented into two independent parallel streams: a mel-spectrogram generator and a waveform synthesizer. The mel-spectrogram is generated in parallel without autoregressive constraints, then fed to the waveform synthesizer which also operates in parallel using the mel-spectrogram as conditional input. This segmentation eliminates the sequential bottleneck while preserving high-fidelity output.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The mel-spectrogram serves as an intermediary representation between the text input and the final waveform output. Instead of generating waveforms directly in an autoregressive manner, the system first generates mel-spectrograms (which can be done in parallel) and then uses these as conditional inputs for waveform synthesis. This intermediary approach enables parallel processing while maintaining speech quality.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Productivity

If end-to-end speech synthesis methods are used, then training efficiency is improved, but the training process becomes unstable and performance is suboptimal

Engineering Contradiction:
Improvetraining efficiencyVSAvoidtraining stability
Core Design Contradiction:
ProductivityVSReliability

Solution Approach 1:

The end-to-end system is segmented into two trainable modules: the mel-spectrogram generator and the waveform synthesizer. Each module can be trained independently with its own loss function and optimization strategy, allowing for stable gradient flow and avoiding the training instability that plishes monolithic end-to-end systems. The modular architecture enables progressive training and better convergence.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The mel-spectrogram generator is trained first to learn the mapping from text to mel-spectrogram representations. Once trained, it provides stable conditional inputs to the waveform synthesizer, which is then trained. This preliminary training of the first module establishes a reliable foundation for the second module, improving overall training stability and performance.

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS11482207B2Waveform generation using end-to-end text-to-waveform system
Publication Date: 2022.10.25 BAIDU USA LLC
  • US11482207B2 patent drawing
  • US11482207B2 patent drawing
  • US11482207B2 patent drawing

AI summary

Described herein are embodiments of an end-to-end text-to-speech (TTS) system with parallel wave generation. In one or more embodiments, a Gaussian inverse autoregressive flow is distilled from an autoregressive WaveNet by minimizing a novel regularized Kullback-Leibler (KL) divergence between their highly-peaked output distributions. Embodiments of the methodology computes the KL divergence in a closed-form, which simplifies the training process and provides very efficient distillation. Embodiments of a novel text-to-wave neural architecture for speech synthesis are also described, which are fully convolutional and enable fast end-to-end training from scratch. These embodiments significantly outperform the previous pipeline that connects a text-to-spectrogram model to a separately trained WaveNet. Also, a parallel waveform synthesizer embodiment conditioned on the hidden representation in an embodiment of this end-to-end model were successfully distilled.