Speech Coding via Time-Varying Interpolation

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current speech coders, such as MBE vocoders, face challenges in reducing bit rate while maintaining speech quality, particularly in accurately conveying spectral formants and adapting to time-varying signal characteristics at low to medium bit rates.

Innovation Solution

The technique involves estimating log spectral magnitudes at fixed intervals, downsampling them data-dependently, and quantizing and reconstructing them, with omitted magnitudes interpolated to refine error measurement and optimize interpolation points for reduced bit rate without compromising time resolution or formant characteristics.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Quantity of substance

If spectral magnitudes are quantized at fixed intervals in MBE vocoders, then the system operates at medium to low data rates, but the ability to accurately convey spectral formants and adapt to time-varying signal characteristics deteriorates

Engineering Contradiction:
Improvedata rateVSAvoidspectral formant accuracy
Core Design Contradiction:
Quantity of substanceVSMeasurement precision

Solution Approach 1:

The patent applies dynamics by making the quantization interval variable rather than fixed. The quantization interval adapts to the local spectral characteristics and rate of change, allowing finer resolution when spectral formants are present or changing rapidly, and coarser resolution when the signal is stationary. This dynamic adjustment resolves the contradiction by enabling accurate spectral representation only when necessary, maintaining low data rates otherwise.

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

The patent changes the parameter of quantization interval from a fixed value to a variable parameter that adjusts based on signal characteristics. By monitoring spectral magnitude changes and adapting the quantization interval accordingly, the system achieves high spectral accuracy when needed while maintaining low overall data rates, thus resolving the contradiction between data rate and spectral formant accuracy.

Inventive Principle:
Principle #35Parameter changes

2Quantity of substance

If the quantization interval is increased to reduce bit rate, then the data rate decreases, but the time resolution and ability to capture rapid speech transitions deteriorates

Engineering Contradiction:
Improvebit rateVSAvoidtime resolution
Core Design Contradiction:
Quantity of substanceVSLoss of time

Solution Approach 1:

The patent uses dynamic adjustment of the quantization interval based on the local rate of change in spectral magnitudes. When rapid speech transitions are detected (high temporal variation), the quantization interval is reduced to preserve time resolution. When the signal is stationary (low temporal variation), the interval is increased to reduce bit rate. This dynamic behavior resolves the contradiction between bit rate and time resolution.

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

The patent applies local quality by allowing different quantization intervals for different time regions based on their specific characteristics. Regions with rapid speech transitions receive finer time resolution (smaller intervals), while stationary regions use coarser resolution (larger intervals). This localized adaptation resolves the contradiction by optimizing both bit rate and time resolution in their respective domains.

Inventive Principle:
Principle #3Local quality

3Quantity of substance

If fewer bits per frame are used to reduce bit rate, then the data rate decreases, but the ability to accurately represent spectral parameters deteriorates

Engineering Contradiction:
Improvebits per frameVSAvoidspectral parameter representation accuracy
Core Design Contradiction:
Quantity of substanceVSManufacturing precision

Solution Approach 1:

The patent changes the parameter of quantization interval from fixed to variable, allowing the system to achieve accurate spectral parameter representation with fewer bits per frame on average. By concentrating bits in regions requiring high precision (rapid transitions, spectral formants) and using coarser quantization elsewhere, the system maintains representation accuracy while reducing overall bit rate.

Inventive Principle:
Principle #35Parameter changes

Solution Approach 2:

The patent applies partial action by providing high-precision spectral parameter representation only where necessary (in regions with rapid changes or important spectral features) rather than uniformly across all frames. This selective precision maintains overall accuracy while significantly reducing the average bits per frame required.

Inventive Principle:
Principle #16Partial or excessive action

Data Source

PatentEP4088277B1Speech coding using time-varying interpolation
Publication Date: 2024.05.29 DIGITAL VOICE SYSTEMS INC
  • EP4088277B1 patent drawingFigure 1
  • EP4088277B1 patent drawingFigure 2
  • EP4088277B1 patent drawingFigure 3

AI summary

Encoding a sequence of digital speech samples into a bit stream includes dividing the digital speech samples into frames including N subframes (where N is greater than 1); computing subframe model parameters including spectral parameters; and generating a representation of the frame that includes information representing the spectral parameters of P subframes (where P < N) and information identifying the P subframes. The representation excludes information representing the spectral parameters of the N-P subframes not included in the P subframes. Generating the representation includes selecting the P subframes by, for multiple combinations of P subframes, determining an error induced by representing the frame using the spectral parameters for the P subframes and using interpolated spectral parameter values for the N-P subframes. A combination of P subframes is selected based on the determined error for the combination of P subframes.