Voice Transformation Using Timbre Vectors and Glottal Segmentation
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current voice transformation methods in speech synthesis and automatic speech recognition suffer from degraded voice quality and accuracy due to reliance on overlapping frames and inadequate separation of prosody and timbre data, leading to unnatural speech and low recognition accuracy.
Innovation Solution
A novel mathematical representation using timbre vectors, formed from non-overlapping frames segmented by glottal closure moments, converted via Fourier analysis and Laguerre functions, allowing for accurate separation of prosody and timbre, enabling high-quality voice transformation and recognition.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Ease of manufacture
If overlapping frames with window functions are used in voice transformation, then speech synthesis can be achieved, but voice quality is severely degraded and speech sounds robotic
Solution Approach 1:
The patent segments the speech signal into non-overlapping frames based on glottal closure moments, eliminating the need for window functions and overlapping frames. This segmentation approach allows for precise separation of prosodic and timbral information without the artifacts introduced by windowing operations, thereby improving voice quality while maintaining speech synthesis capability
Solution Approach 2:
The patent extracts and separates prosodic features (pitch, duration, intensity) from timbral features (spectral characteristics) by using non-overlapping frames segmented at glottal closure moments. This extraction allows independent manipulation of each feature type, enabling high-quality voice transformation without the degradation caused by conventional overlapping frame methods
2Adaptability or versatility
If conventional parameterization methods (LPC, mel-cepstral coefficients) are used, then speech can be represented parametrically, but accuracy of speech parameterization is insufficient for high-quality synthesis and recognition
Solution Approach 1:
The patent employs dynamic time warping and non-linear parameter transformations to adapt the parametric representation to the specific characteristics of each speech segment. By using non-overlapping frames segmented at glottal closure moments, the system can dynamically adjust parameter extraction to capture transient and formant characteristics with higher precision than static conventional methods
Solution Approach 2:
The patent transforms the spectral characteristics of each non-overlapping frame into multiple parameter spaces (time-domain, frequency-domain, cepstral-domain) using non-linear mappings. This multi-parameter representation captures both prosodic and timbral information with high accuracy, enabling superior speech synthesis and recognition performance
3Quantity of substance
If a large speech database is used in concatenative TTS systems, then more speech units are available for selection, but mismatches at segment borders increase and modifications are required
Solution Approach 1:
The patent extracts and separates prosodic parameters (pitch contour, duration, intensity) from the speech signal at the frame level using non-overlapping segments. This extraction allows for precise control of segment boundaries and enables smooth transitions between segments by independently adjusting prosodic parameters, eliminating mismatches without requiring large databases or complex modifications
Applied Scientific Principles
This section explains which scientific principles are used to turn an abstract innovation direction into a practical engineering solution.
Function Achieved in This Case
This approach enables natural voice modification with varied speaker identity, pitch, and prosodic variations, improving speech synthesis quality and automatic speech recognition accuracy by using a compact database and eliminating the need for window functions.
Implementation Method 1
Using Fourier analysis, the speech signal in each frame is converted into amplitude spectrum
Implementation Method 2
Laguerre functions (based on a set of orthogonal polynomials) are used to convert the amplitude spectrum into a unit vector characteristic to the instantaneous timbre
Data Source
AI summary
The present invention is a method and system to convert speech signal into a parametric representation in terms of timbre vectors, and to recover the speech signal thereof. The speech signal is first segmented into non-overlapping frames using the glottal closure instant information, each frame is converted into an amplitude spectrum using a Fourier analyzer, and then using Laguerre functions to generate a set of coefficients which constitute a timbre vector. A sequence of timbre vectors can be subject to a variety of manipulations. The new timbre vectors are converted back into voice signals by first transforming into amplitude spectra using Laguerre functions, then generating phase spectra from the amplitude spectra using Kramers-Knonig relations. A Fourier transformer converts the amplitude spectra and phase spectra into elementary waveforms, then superposed to become the output voice. The method and system can be used for voice transformation, speech synthesis, and automatic speech recognition.


