Non-Parametric Voice Conversion via HMM State Vector Replacement

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional methods for voice conversion in speech synthesis using hidden Markov models (HMMs) often result in over-smoothed spectral envelopes, leading to lower-quality speech with a muffled character, as they aim to adapt models to target speakers based on likelihood rather than capturing voice characteristics effectively.

Innovation Solution

A non-parametric voice conversion method is employed, where a source HMM is adapted by replacing its generator-model functions with target-speaker vectors and applying a fundamental frequency (F0) transform to match the target speaker's F0 statistics, ensuring the converted HMM generates speech with the voice characteristics of the target speaker.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If conventional parametric methods are used to adapt HMM models to target speakers based on likelihood, then model adaptation is achieved, but the spectral envelopes become over-smoothed resulting in lower-quality speech with a muffled character

Engineering Contradiction:
Improvemodel adaptation accuracyVSAvoidspectral envelope quality
Core Design Contradiction:
ReliabilityVSManufacturing precision

Solution Approach 1:

The patent changes the fundamental parameter representation from parametric (Gaussian parameters) to non-parametric (spectral vectors). By representing speaker characteristics as spectral vectors derived from actual speech signals rather than statistical parameters, the method preserves fine spectral details while adapting to target speakers, avoiding the over-smoothing problem inherent in parametric approaches.

Inventive Principle:
Principle #35Parameter changes

Solution Approach 2:

The patent uses spectral vectors that directly copy the spectral characteristics of target speaker speech. Instead of modeling speaker traits through statistical parameters, the system extracts and utilizes actual spectral snapshots from target speech, thereby preserving authentic spectral envelope details and avoiding artificial smoothing.

Inventive Principle:
Principle #26Copying

2Manufacturing precision

If non-parametric voice conversion is used to preserve spectral details, then speech quality is improved, but computational complexity increases due to vector matching and F0 transform operations

Engineering Contradiction:
Improvespectral envelope qualityVSAvoidprocessing complexity
Core Design Contradiction:
Manufacturing precisionVSDevice complexity

Solution Approach 1:

The patent segments the spectral matching process into discrete vector operations. By representing speech as sequences of spectral vectors and performing matching at the vector level rather than continuous function level, the system simplifies computational operations while preserving spectral fidelity, making the complex non-parametric approach more tractable.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent transforms the computational problem from continuous parameter optimization to discrete vector matching. By changing from parametric model adaptation to non-parametric vector substitution, the system reduces computational complexity while maintaining or improving spectral quality, as vector matching is more direct and less computationally intensive than iterative parametric optimization.

Inventive Principle:
Principle #35Parameter changes

3Measurement precision

If target-speaker vectors are used to replace generator-model functions, then voice characteristics are accurately captured, but the conversion process becomes more complex

Engineering Contradiction:
Improvevoice characteristics accuracyVSAvoidconversion process complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent directly copies spectral characteristics from target speaker speech into the HMM generator functions. By substituting generator-model functions with target-speaker spectral vectors, the system accurately captures voice characteristics without requiring complex transformation models, as the spectral vectors themselves contain the authentic voice print of the target speaker.

Inventive Principle:
Principle #26Copying

Solution Approach 2:

The patent inverts the conventional approach by instead of adapting the model to match target speaker statistics, it directly replaces model functions with target speaker spectral representations. This inversion simplifies the conversion process by eliminating the need for complex adaptation algorithms, as the target spectral vectors directly provide the desired voice characteristics.

Inventive Principle:
Principle #13The other way round (Inversion)

Data Source

PatentUS9183830B2Method and system for non-parametric voice conversion
Publication Date: 2015.11.10 GOOGLE LLC
  • US9183830B2 patent drawing
  • US9183830B2 patent drawing
  • US9183830B2 patent drawing

AI summary

A method and system is disclosed for non-parametric speech conversion. A text-to-speech (TTS) synthesis system may include hidden Markov model (HMM) HMM based speech modeling for both synthesizing output speech. A converted HMM may be initially set to a source HMM trained with a voice of a source speaker. A parametric representation of speech may be extract from speech of a target speaker to generate a set of target-speaker vectors. A matching procedure, carried out under a transform that compensates for speaker differences, may be used to match each HMM state of the source HMM to a target-speaker vector. The HMM states of the converted HMM may be replaced with the matched target-speaker vectors. Transforms may be applied to further adapt the converted HMM to the voice of target speaker. The converted HMM may be used to synthesize speech with voice characteristics of the target speaker.