Neural Homomorphic Vocoder for Speech Synthesis

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing speech synthesis methods require significant computational resources and result in low-quality synthesized speech due to inefficient processing of acoustic features.

Innovation Solution

A neural homomorphic vocoder framework that processes acoustic features using a neural network filter estimator to obtain impulse response information, with harmonic and noise components modeled by time-varying filters, reducing computational complexity while improving speech quality.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Manufacturing precision

If neural vocoders are used for speech synthesis, then synthesis quality is improved, but computational complexity increases significantly

Engineering Contradiction:
Improvesynthesis qualityVSAvoidcomputational complexity
Core Design Contradiction:
Manufacturing precisionVSDevice complexity

Solution Approach 1:

The patent segments the speech signal into harmonic components and noise components, processing each separately through dedicated time-varying filters. This segmentation allows the system to focus computational resources on specific signal characteristics rather than processing the entire signal uniformly, thereby maintaining high synthesis quality while reducing overall computational complexity.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent transforms the speech synthesis problem from direct waveform generation to parameter-based synthesis by estimating impulse response information and using time-varying filter parameters. This parameter change approach allows for more efficient computation while preserving the quality characteristics of the original speech signal.

Inventive Principle:
Principle #35Parameter changes

2Manufacturing precision

If source-filter models with neural networks are used, then synthesis quality is improved, but computation time increases

Engineering Contradiction:
Improvesynthesis qualityVSAvoidcomputation time
Core Design Contradiction:
Manufacturing precisionVSLoss of time

Solution Approach 1:

The patent performs preliminary action by pre-estimating impulse response information from acoustic features before actual speech synthesis. This preprocessing step creates reusable filter parameters that can be applied during synthesis without requiring intensive real-time computation, thereby reducing computation time while maintaining quality.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent substitutes traditional mechanical speech synthesis mechanisms with a neural network-based filter estimation system. By replacing direct waveform generation with neural network-based parameter estimation and time-varying filtering, the system achieves faster computation while preserving synthesis quality.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

3Productivity

If parallel generation models are used, then generation speed is improved, but synthesis quality decreases

Engineering Contradiction:
Improvegeneration speedVSAvoidsynthesis quality
Core Design Contradiction:
ProductivityVSManufacturing precision

Solution Approach 1:

The patent introduces dynamics by using time-varying filters that adapt their parameters over time based on the input acoustic features. This dynamic adaptation allows the system to capture temporal variations in speech signals more effectively than static parallel models, improving synthesis quality while maintaining efficient parallel generation capabilities.

Inventive Principle:
Principle #15Dynamics

Data Source

PatentEP4099316B1Speech synthesis method and system
Publication Date: 2024.09.25 AISPEECH CO LTD
  • EP4099316B1 patent drawingFigure 1
  • EP4099316B1 patent drawingFigure 2~4
  • EP4099316B1 patent drawingFigure 5~8

AI summary

Disclosed is a speech synthesis method including: acquiring fundamental frequency information and acoustic feature information from original speech; generating an impulse train from the fundamental frequency information, and inputting it to a harmonic time-varying filter; inputting the acoustic feature information into a neural network filter estimator to obtain corresponding impulse response information; generating noise signal by a noise generator; determining, by the harmonic time-varying filter, harmonic component information through filtering processing on the impulse train and the impulse response information; determining, by a noise time-varying filter, noise component information based on the impulse response information and the noise; and generating a synthesized speech from the harmonic component information and the noise component information. Acoustic features are processed to obtain corresponding impulse response information, and harmonic component information and noise component information are modeled respectively, thereby reducing computation of speech synthesis and improving the quality of the synthesized speech.