Neural Vocoder Speech Waveform Generation
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Traditional parametric vocoders suffer from irreversible quantization and reconstruction losses, leading to 'muffled' or 'buzzy' synthetic speech, and advanced autoregressive generative models are computationally expensive, making them unsuitable for real-time synthesis on devices.
Innovation Solution
A neural network-based vocoder is designed to mimic source and filter models, using glottal and vocal tract features to generate high-quality speech waveforms with low computational and memory costs, incorporating knowledge from speech signal processing to improve synthesis efficiency.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Loss of energy
If traditional parametric vocoders are used for speech synthesis, then computational efficiency is maintained, but speech quality deteriorates due to quantization and reconstruction losses
Solution Approach 1:
The patent replaces traditional mechanical vocoding systems with a neural network-based system. The neural vocoder uses deep learning models to directly map spectral features to waveforms, eliminating the need for parametric modeling and reconstruction stages that cause quality loss, while maintaining computational efficiency through optimized network architecture
Solution Approach 2:
The patent changes the fundamental parameters of the speech synthesis system by transitioning from parametric representation (spectral coefficients) to direct waveform generation. The neural network learns optimal waveform representations and generation strategies, fundamentally changing how speech is synthesized while improving both quality and efficiency
2Manufacturing precision
If advanced autoregressive generative models are used for speech synthesis, then speech quality is improved, but computational cost increases making real-time synthesis infeasible
Solution Approach 1:
The patent segments the speech synthesis process into distinct neural network components: a spectral extractor that processes input text to obtain spectral features, and a neural vocoder that generates waveforms from these features. This segmentation allows each component to be optimized independently, achieving high quality while maintaining real-time performance
Solution Approach 2:
The patent uses a streamlined neural network architecture that performs only the essential waveform generation task without the excessive computational overhead of autoregressive models. By focusing on the critical path from spectral features to waveform output, the system achieves real-time synthesis capability
Data Source
AI summary
A method and apparatus for generating a speech waveform. Fundamental frequency information, glottal features and vocal tract features associated with an input may be received, wherein the glottal features include a phase feature, a shape feature, and an energy feature (1310). A glottal waveform is generated based on the fundamental frequency information and the glottal features through a first neural network model (1320). A speech waveform is generated based on the glottal waveform and the vocal tract features through a second neural network model (1330).


