Pitch-Synchronous Timbre Vector Speech Coding

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current speech coding technologies, particularly those based on linear predictive coding (LPC), are limited by their pitch-asynchronous nature and small number of parameters, resulting in poor quality speech representation, especially when compared to CD-quality standards, and fail to accurately capture the spectral details of wide-band speech signals.

Innovation Solution

The use of pitch-synchronous timbre vectors for speech coding, where pitch marks are identified and extended to segment speech signals into frames, and these frames are converted into unit vectors using FFT and Laguerre functions, allowing for scalar and vector quantization to encode speech signals with higher fidelity and lower bandwidth.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Ease of manufacture

If LPC-based speech coding is used, then the coding process is simple and widely applicable, but the speech quality is poor and cannot capture spectral details accurately

Engineering Contradiction:
Improvecoding process simplicityVSAvoidspectral detail accuracy
Core Design Contradiction:
Ease of manufactureVSManufacturing precision

Solution Approach 1:

The patent changes the fundamental parameters from pitch-asynchronous LPC coefficients to pitch-synchronous timbre vectors. This parameter transformation enables accurate capture of spectral details while maintaining coding efficiency. The timbre vectors are derived from pitch-synchronous spectral analysis, allowing precise representation of frequency content without the convergence limitations of LPC.

Inventive Principle:
Principle #35Parameter changes

Solution Approach 2:

The patent transitions from a low-dimensional LPC parameter space (10-16 coefficients) to a higher-dimensional timbre vector space that captures spectral details across multiple frequency bands. This dimensional expansion enables accurate representation of complex spectral patterns including fricatives and nasal sounds while maintaining manageable codebook sizes through vector quantization.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

2Manufacturing precision

If the number of LPC coefficients is increased to improve quality, then spectral representation improves, but the coefficients become non-converging and the number of parameters becomes too large

Engineering Contradiction:
Improvespectral representation qualityVSAvoidnumber of parameters
Core Design Contradiction:
Manufacturing precisionVSDevice complexity

Solution Approach 1:

The patent replaces LPC coefficients with timbre vectors that are derived from pitch-synchronous spectral analysis. This parameter substitution allows accurate spectral representation without the non-converging behavior of LPC. The timbre vectors are constructed from a manageable number of frequency band components, avoiding the need for excessive parameters while maintaining quality.

Inventive Principle:
Principle #35Parameter changes

Solution Approach 2:

The patent uses codebooks that store representative timbre vectors for different speech categories. Instead of using a large number of parameters to describe all possible spectral patterns, the system copies and stores a manageable set of prototype vectors that can represent the full range of speech characteristics through vector quantization.

Inventive Principle:
Principle #26Copying

3Productivity

If pitch-asynchronous LPC is used, then the coding is computationally efficient, but the quality is limited compared to pitch-synchronous methods

Engineering Contradiction:
Improvecomputational efficiencyVSAvoidspeech quality
Core Design Contradiction:
ProductivityVSManufacturing precision

Solution Approach 1:

The patent changes from pitch-asynchronous LPC parameters to pitch-synchronous timbre vectors. This parameter transformation enables accurate capture of pitch-related spectral variations while maintaining computational efficiency. The pitch-synchronous analysis is performed only at the pitch period boundaries, and the resulting timbre vectors are efficiently quantized using codebooks.

Inventive Principle:
Principle #35Parameter changes

Solution Approach 2:

The patent segments the speech signal into pitch-synchronous frames based on pitch period detection. This segmentation allows the system to analyze and code only the relevant portions of the signal at the appropriate pitch intervals, maintaining computational efficiency while capturing pitch-related spectral details that are missed by fixed-frame LPC methods.

Inventive Principle:
Principle #1Segmentation

4Ease of manufacture

If fixed-duration frames are used for LPC coding, then the processing is straightforward, but the frame length must be longer than maximum pitch period and requires windowing and overlapping

Engineering Contradiction:
Improveprocessing straightforwardnessVSAvoidframe processing requirements
Core Design Contradiction:
Ease of manufactureVSDevice complexity

Solution Approach 1:

The patent replaces fixed-duration frames with dynamic pitch-synchronous frames that adapt to the actual pitch period of the speech signal. This dynamic framing eliminates the need for fixed windowing and overlapping operations, as each frame is naturally aligned with the pitch periods. The frame duration automatically adjusts to match the speaker's pitch characteristics.

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

The pitch-synchronous framing system is self-adjusting, automatically detecting pitch periods and setting frame boundaries accordingly. This self-service mechanism eliminates the need for manual optimization of frame length, window function selection, and overlap parameters, simplifying the overall processing while maintaining accuracy.

Inventive Principle:
Principle #25Self-service

Applied Scientific Principles

This section explains which scientific principles are used to turn an abstract innovation direction into a practical engineering solution.

Function Achieved in This Case

This approach enables the transmission of high-quality speech signals at low bandwidth, surpassing the limitations of traditional LPC-based methods by accurately capturing spectral details and reproducing CD-quality speech, including nuanced sounds like fricatives and nasal sounds, with improved naturalness and reduced encoding delay.

Implementation Method 1

The pitch marks are extended to unvoiced sections to generate a complete set of segmentation points. The speech signal is segmented into pitch-synchronous frames according to the said segmentation points.

Methodology Applied
Scientific EffectSignal segmentation:

Implementation Method 2

Using FFT (fast Fourier transform), the speech signal in each frame is converted into a pitch-synchronous amplitude spectrum

Methodology Applied
Scientific EffectFast Fourier Transform:

Implementation Method 3

use Laguerre functions to convert the said pitch-synchronous amplitude spectrum into a unit vector characteristic to the instantaneous timbre, referred to as the timbre vector

Methodology Applied
Scientific EffectLaguerre transform:

Implementation Method 4

Using scalar quantization, the pitch period and the intensity are converted into a pitch index and an intensity index using a pitch codebook and an intensity codebook

Methodology Applied
Scientific EffectScalar quantization:

Implementation Method 5

Using vector quantization, each timbre vector is converted to a timbre index using a timbre codebook

Methodology Applied
Scientific EffectVector quantization:

Data Source

PatentUS20150262587A1Pitch Synchronous Speech Coding Based on Timbre Vectors
Publication Date: 2015.09.17 THE TRUSTEES OF COLUMBIA UNIV IN THE CITY OF NEW YORK
  • US20150262587A1 patent drawing
  • US20150262587A1 patent drawing
  • US20150262587A1 patent drawing

AI summary

A pitch-synchronous method and system for speech coding using timbre vectors is disclosed. On the encoder side, speech signal is segmented into pitch-synchronous frames without overlap, then converted into a pitch-synchronous amplitude spectrum using FFT. Using Laguerre functions, the amplitude spectrum is transformed into a timbre vector. Using vector quantization, each timbre vector is converted to a timbre index based on a timbre codebook. The intensity and pitch are also converted into indices respectively using scalar quantization. Those indices are transmitted as encoded speech. On the decoder side, by looking up the same codebooks, pitch, intensity and the timbre vector are recovered. Using Laguerre functions, the amplitude spectrum is recovered. Using Kramers-Kronig relations, the phase spectrum is recovered. Using FFT, the elementary waves are regenerated, and superposed to become the speech signal.