Voice Transformation Using Timbre Vectors and Glottal Segmentation

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current voice transformation methods in speech synthesis and automatic speech recognition suffer from degraded voice quality and accuracy due to reliance on overlapping frames and inadequate separation of prosody and timbre data, leading to unnatural speech and low recognition accuracy.

Innovation Solution

A novel mathematical representation using timbre vectors, formed from non-overlapping frames segmented by glottal closure moments, converted via Fourier analysis and Laguerre functions, allowing for accurate separation of prosody and timbre, enabling high-quality voice transformation and recognition.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Ease of manufacture

If overlapping frames with window functions are used in voice transformation, then speech synthesis can be achieved, but voice quality is severely degraded and speech sounds robotic

Engineering Contradiction:
Improvespeech synthesis capabilityVSAvoidvoice quality
Core Design Contradiction:
Ease of manufactureVSManufacturing precision

Solution Approach 1:

The patent segments the speech signal into non-overlapping frames based on glottal closure moments, eliminating the need for window functions and overlapping frames. This segmentation approach allows for precise separation of prosodic and timbral information without the artifacts introduced by windowing operations, thereby improving voice quality while maintaining speech synthesis capability

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent extracts and separates prosodic features (pitch, duration, intensity) from timbral features (spectral characteristics) by using non-overlapping frames segmented at glottal closure moments. This extraction allows independent manipulation of each feature type, enabling high-quality voice transformation without the degradation caused by conventional overlapping frame methods

Inventive Principle:
Principle #2Taking out (Extraction)

2Adaptability or versatility

If conventional parameterization methods (LPC, mel-cepstral coefficients) are used, then speech can be represented parametrically, but accuracy of speech parameterization is insufficient for high-quality synthesis and recognition

Engineering Contradiction:
Improveparametric representation capabilityVSAvoidspeech parameterization accuracy
Core Design Contradiction:
Adaptability or versatilityVSMeasurement precision

Solution Approach 1:

The patent employs dynamic time warping and non-linear parameter transformations to adapt the parametric representation to the specific characteristics of each speech segment. By using non-overlapping frames segmented at glottal closure moments, the system can dynamically adjust parameter extraction to capture transient and formant characteristics with higher precision than static conventional methods

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

The patent transforms the spectral characteristics of each non-overlapping frame into multiple parameter spaces (time-domain, frequency-domain, cepstral-domain) using non-linear mappings. This multi-parameter representation captures both prosodic and timbral information with high accuracy, enabling superior speech synthesis and recognition performance

Inventive Principle:
Principle #35Parameter changes

3Quantity of substance

If a large speech database is used in concatenative TTS systems, then more speech units are available for selection, but mismatches at segment borders increase and modifications are required

Engineering Contradiction:
Improvespeech database sizeVSAvoidsegment boundary continuity
Core Design Contradiction:
Quantity of substanceVSReliability

Solution Approach 1:

The patent extracts and separates prosodic parameters (pitch contour, duration, intensity) from the speech signal at the frame level using non-overlapping segments. This extraction allows for precise control of segment boundaries and enables smooth transitions between segments by independently adjusting prosodic parameters, eliminating mismatches without requiring large databases or complex modifications

Inventive Principle:
Principle #2Taking out (Extraction)

Applied Scientific Principles

This section explains which scientific principles are used to turn an abstract innovation direction into a practical engineering solution.

Function Achieved in This Case

This approach enables natural voice modification with varied speaker identity, pitch, and prosodic variations, improving speech synthesis quality and automatic speech recognition accuracy by using a compact database and eliminating the need for window functions.

Implementation Method 1

Using Fourier analysis, the speech signal in each frame is converted into amplitude spectrum

Methodology Applied
Scientific EffectFourier analysis:

Implementation Method 2

Laguerre functions (based on a set of orthogonal polynomials) are used to convert the amplitude spectrum into a unit vector characteristic to the instantaneous timbre

Methodology Applied
Scientific EffectOrthogonal function transformation:

Data Source

PatentUS8744854B1System and method for voice transformation
Publication Date: 2014.06.03 THE TRUSTEES OF COLUMBIA UNIV IN THE CITY OF NEW YORK
  • US8744854B1 patent drawing
  • US8744854B1 patent drawing
  • US8744854B1 patent drawing

AI summary

The present invention is a method and system to convert speech signal into a parametric representation in terms of timbre vectors, and to recover the speech signal thereof. The speech signal is first segmented into non-overlapping frames using the glottal closure instant information, each frame is converted into an amplitude spectrum using a Fourier analyzer, and then using Laguerre functions to generate a set of coefficients which constitute a timbre vector. A sequence of timbre vectors can be subject to a variety of manipulations. The new timbre vectors are converted back into voice signals by first transforming into amplitude spectra using Laguerre functions, then generating phase spectra from the amplitude spectra using Kramers-Knonig relations. A Fourier transformer converts the amplitude spectra and phase spectra into elementary waveforms, then superposed to become the output voice. The method and system can be used for voice transformation, speech synthesis, and automatic speech recognition.