Deep Generation Model for Voice Signal Parameter Estimation

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing voice signal analysis technologies face challenges in accurately estimating parameters from fundamental frequency patterns and reconstructing voice signals with high accuracy, particularly due to high calculation costs and limited estimation accuracy in handling voice generation and inverse problems.

Innovation Solution

A voice signal analysis apparatus and method utilizing a deep generation model with an encoder and decoder, where the encoder estimates latent variables from fundamental frequency patterns and the decoder reconstructs the patterns, allowing for efficient and accurate parameter estimation and pattern reconstruction using parallel data.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If conventional methods (HMM, VAE) are used for parameter estimation and pattern reconstruction, then the system can handle voice generation and inverse problems, but the calculation cost is high and estimation accuracy is limited

Engineering Contradiction:
Improveparameter estimation accuracyVSAvoidcalculation time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The patent segments the voice signal analysis into distinct components: fundamental frequency pattern extraction, parameter estimation through encoder, and pattern reconstruction through decoder. This segmentation allows each component to be optimized independently, reducing overall calculation time while maintaining accuracy.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent replaces conventional statistical methods (HMM) and basic neural network approaches with a deep generative model framework. This substitution enables more efficient parameter estimation and pattern reconstruction by leveraging the generative capabilities of deep learning models, thereby reducing calculation time while improving accuracy.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

2Device complexity

If conventional VAE is used with normal distribution assumption, then the model structure is simple, but it cannot effectively associate observation data with interpretable parameters

Engineering Contradiction:
Improvemodel structure complexityVSAvoidinterpretable parameter association
Core Design Contradiction:
Device complexityVSLoss of information

Solution Approach 1:

The patent changes the distributional assumptions and structural parameters of the VAE to better suit voice signal analysis. By modifying how latent variables are distributed and how they relate to interpretable parameters like fundamental frequency patterns, the model achieves better association between observation data and parameters without excessive complexity.

Inventive Principle:
Principle #35Parameter changes

3Adaptability or versatility

If deep generation model is trained simultaneously for encoder and decoder, then voice generation and inverse problems can be solved together, but training complexity increases

Engineering Contradiction:
Improvedual problem solving capabilityVSAvoidtraining process complexity
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The patent merges the voice generation process and the inverse problem solving into a single unified deep generative model training process. The encoder and decoder are trained simultaneously with a combined objective function that optimizes both capabilities, allowing the system to handle both forward and inverse problems without separate training procedures.

Inventive Principle:
Principle #5Merging (Combining)

Data Source

PatentUS11798579B2Device, method, and program for analyzing speech signal
Publication Date: 2023.10.24 NIPPON TELEGRAPH & TELEPHONE CORP
  • US11798579B2 patent drawing
  • US11798579B2 patent drawing
  • US11798579B2 patent drawing

AI summary

A parameter included in a fundamental frequency pattern of a voice can be estimated from the fundamental frequency pattern with high accuracy and the fundamental frequency pattern of the voice can be reconstructed from the parameter included in the fundamental frequency pattern. A learning unit 30 learns a deep generation model including an encoder which regards a parameter included in a fundamental frequency pattern in a voice signal as a latent variable of the deep generation model and estimates the latent variable from the fundamental frequency pattern in the voice signal on the basis of parallel data of the fundamental frequency pattern in the voice signal and the parameter included in the fundamental frequency pattern in the voice signal, and a decoder which reconstructs the fundamental frequency pattern in the voice signal from the latent variable.