Neural Vocoder Speech Waveform Generation

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Traditional parametric vocoders suffer from irreversible quantization and reconstruction losses, leading to 'muffled' or 'buzzy' synthetic speech, and advanced autoregressive generative models are computationally expensive, making them unsuitable for real-time synthesis on devices.

Innovation Solution

A neural network-based vocoder is designed to mimic source and filter models, using glottal and vocal tract features to generate high-quality speech waveforms with low computational and memory costs, incorporating knowledge from speech signal processing to improve synthesis efficiency.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Loss of energy

If traditional parametric vocoders are used for speech synthesis, then computational efficiency is maintained, but speech quality deteriorates due to quantization and reconstruction losses

Engineering Contradiction:
Improvecomputational efficiencyVSAvoidspeech synthesis quality
Core Design Contradiction:
Loss of energyVSManufacturing precision

Solution Approach 1:

The patent replaces traditional mechanical vocoding systems with a neural network-based system. The neural vocoder uses deep learning models to directly map spectral features to waveforms, eliminating the need for parametric modeling and reconstruction stages that cause quality loss, while maintaining computational efficiency through optimized network architecture

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Solution Approach 2:

The patent changes the fundamental parameters of the speech synthesis system by transitioning from parametric representation (spectral coefficients) to direct waveform generation. The neural network learns optimal waveform representations and generation strategies, fundamentally changing how speech is synthesized while improving both quality and efficiency

Inventive Principle:
Principle #35Parameter changes

2Manufacturing precision

If advanced autoregressive generative models are used for speech synthesis, then speech quality is improved, but computational cost increases making real-time synthesis infeasible

Engineering Contradiction:
Improvespeech synthesis qualityVSAvoidreal-time synthesis capability
Core Design Contradiction:
Manufacturing precisionVSProductivity

Solution Approach 1:

The patent segments the speech synthesis process into distinct neural network components: a spectral extractor that processes input text to obtain spectral features, and a neural vocoder that generates waveforms from these features. This segmentation allows each component to be optimized independently, achieving high quality while maintaining real-time performance

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent uses a streamlined neural network architecture that performs only the essential waveform generation task without the excessive computational overhead of autoregressive models. By focusing on the critical path from spectral features to waveform output, the system achieves real-time synthesis capability

Inventive Principle:
Principle #16Partial or excessive action

Data Source

PatentUS11869482B2Speech waveform generation
Publication Date: 2024.01.09 MICROSOFT TECHNOLOGY LICENSING LLC
  • US11869482B2 patent drawing
  • US11869482B2 patent drawing
  • US11869482B2 patent drawing

AI summary

A method and apparatus for generating a speech waveform. Fundamental frequency information, glottal features and vocal tract features associated with an input may be received, wherein the glottal features include a phase feature, a shape feature, and an energy feature (1310). A glottal waveform is generated based on the fundamental frequency information and the glottal features through a first neural network model (1320). A speech waveform is generated based on the glottal waveform and the vocal tract features through a second neural network model (1330).