Pitch-Synchronous Timbre Vector Speech Coding
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current speech coding technologies, particularly those based on linear predictive coding (LPC), are limited by their pitch-asynchronous nature and small number of parameters, resulting in poor quality speech representation, especially when compared to CD-quality standards, and fail to accurately capture the spectral details of wide-band speech signals.
Innovation Solution
The use of pitch-synchronous timbre vectors for speech coding, where pitch marks are identified and extended to segment speech signals into frames, and these frames are converted into unit vectors using FFT and Laguerre functions, allowing for scalar and vector quantization to encode speech signals with higher fidelity and lower bandwidth.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Ease of manufacture
If LPC-based speech coding is used, then the coding process is simple and widely applicable, but the speech quality is poor and cannot capture spectral details accurately
Solution Approach 1:
The patent changes the fundamental parameters from pitch-asynchronous LPC coefficients to pitch-synchronous timbre vectors. This parameter transformation enables accurate capture of spectral details while maintaining coding efficiency. The timbre vectors are derived from pitch-synchronous spectral analysis, allowing precise representation of frequency content without the convergence limitations of LPC.
Solution Approach 2:
The patent transitions from a low-dimensional LPC parameter space (10-16 coefficients) to a higher-dimensional timbre vector space that captures spectral details across multiple frequency bands. This dimensional expansion enables accurate representation of complex spectral patterns including fricatives and nasal sounds while maintaining manageable codebook sizes through vector quantization.
2Manufacturing precision
If the number of LPC coefficients is increased to improve quality, then spectral representation improves, but the coefficients become non-converging and the number of parameters becomes too large
Solution Approach 1:
The patent replaces LPC coefficients with timbre vectors that are derived from pitch-synchronous spectral analysis. This parameter substitution allows accurate spectral representation without the non-converging behavior of LPC. The timbre vectors are constructed from a manageable number of frequency band components, avoiding the need for excessive parameters while maintaining quality.
Solution Approach 2:
The patent uses codebooks that store representative timbre vectors for different speech categories. Instead of using a large number of parameters to describe all possible spectral patterns, the system copies and stores a manageable set of prototype vectors that can represent the full range of speech characteristics through vector quantization.
3Productivity
If pitch-asynchronous LPC is used, then the coding is computationally efficient, but the quality is limited compared to pitch-synchronous methods
Solution Approach 1:
The patent changes from pitch-asynchronous LPC parameters to pitch-synchronous timbre vectors. This parameter transformation enables accurate capture of pitch-related spectral variations while maintaining computational efficiency. The pitch-synchronous analysis is performed only at the pitch period boundaries, and the resulting timbre vectors are efficiently quantized using codebooks.
Solution Approach 2:
The patent segments the speech signal into pitch-synchronous frames based on pitch period detection. This segmentation allows the system to analyze and code only the relevant portions of the signal at the appropriate pitch intervals, maintaining computational efficiency while capturing pitch-related spectral details that are missed by fixed-frame LPC methods.
4Ease of manufacture
If fixed-duration frames are used for LPC coding, then the processing is straightforward, but the frame length must be longer than maximum pitch period and requires windowing and overlapping
Solution Approach 1:
The patent replaces fixed-duration frames with dynamic pitch-synchronous frames that adapt to the actual pitch period of the speech signal. This dynamic framing eliminates the need for fixed windowing and overlapping operations, as each frame is naturally aligned with the pitch periods. The frame duration automatically adjusts to match the speaker's pitch characteristics.
Solution Approach 2:
The pitch-synchronous framing system is self-adjusting, automatically detecting pitch periods and setting frame boundaries accordingly. This self-service mechanism eliminates the need for manual optimization of frame length, window function selection, and overlap parameters, simplifying the overall processing while maintaining accuracy.
Applied Scientific Principles
This section explains which scientific principles are used to turn an abstract innovation direction into a practical engineering solution.
Function Achieved in This Case
This approach enables the transmission of high-quality speech signals at low bandwidth, surpassing the limitations of traditional LPC-based methods by accurately capturing spectral details and reproducing CD-quality speech, including nuanced sounds like fricatives and nasal sounds, with improved naturalness and reduced encoding delay.
Implementation Method 1
The pitch marks are extended to unvoiced sections to generate a complete set of segmentation points. The speech signal is segmented into pitch-synchronous frames according to the said segmentation points.
Implementation Method 2
Using FFT (fast Fourier transform), the speech signal in each frame is converted into a pitch-synchronous amplitude spectrum
Implementation Method 3
use Laguerre functions to convert the said pitch-synchronous amplitude spectrum into a unit vector characteristic to the instantaneous timbre, referred to as the timbre vector
Implementation Method 4
Using scalar quantization, the pitch period and the intensity are converted into a pitch index and an intensity index using a pitch codebook and an intensity codebook
Implementation Method 5
Using vector quantization, each timbre vector is converted to a timbre index using a timbre codebook
Data Source
AI summary
A pitch-synchronous method and system for speech coding using timbre vectors is disclosed. On the encoder side, speech signal is segmented into pitch-synchronous frames without overlap, then converted into a pitch-synchronous amplitude spectrum using FFT. Using Laguerre functions, the amplitude spectrum is transformed into a timbre vector. Using vector quantization, each timbre vector is converted to a timbre index based on a timbre codebook. The intensity and pitch are also converted into indices respectively using scalar quantization. Those indices are transmitted as encoded speech. On the decoder side, by looking up the same codebooks, pitch, intensity and the timbre vector are recovered. Using Laguerre functions, the amplitude spectrum is recovered. Using Kramers-Kronig relations, the phase spectrum is recovered. Using FFT, the elementary waves are regenerated, and superposed to become the speech signal.


