Multi-tap LTP Filter with Sub-sample Delay for Speech Coding

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing speech coding technologies face challenges in efficiently modeling non-integral delay values and providing spectral shaping, especially in wideband speech coding systems, where the harmonic structure weakens at higher frequencies, and current methods often increase complexity or require additional quantization.

Innovation Solution

A novel multi-tap LTP filter is introduced, using sub-sample resolution delay to explicitly model fractional delay values and incorporate spectral shaping, allowing for efficient representation of both delay and spectral shaping without the need for additional filtering operations or increased complexity.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If conventional integer-sample resolution LTP filter is used, then device complexity is reduced, but measurement precision of delay values deteriorates

Engineering Contradiction:
Improvedelay value precisionVSAvoidfilter structure complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The delay value is segmented into an integer part and a fractional part. The integer part is handled by the conventional LTP filter structure, while the fractional part is modeled using a separate rational delay element. This segmentation allows precise delay representation without increasing the overall filter structure complexity.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

A rational delay element is introduced as an intermediary component between the integer-sample LTP filter and the input signal. This intermediary handles the fractional delay modeling, allowing the main LTP filter to maintain its simple integer-sample structure while achieving sub-sample precision through the rational delay mediator.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Reliability

If spectral shaping filter is added to LTP filter, then speech quality is improved, but device complexity increases

Engineering Contradiction:
Improvespeech qualityVSAvoidfilter structure complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The spectral shaping function is merged with the LTP filter by incorporating it into the rational delay element. Instead of adding a separate spectral shaping filter, the shaping characteristics are integrated into the existing rational delay structure, achieving speech quality improvement without increasing device complexity.

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The rational delay element is designed to serve multiple functions simultaneously: it provides fractional delay modeling, spectral shaping, and noise shaping. This multi-functionality eliminates the need for separate dedicated filters, maintaining device complexity at acceptable levels while achieving improved speech quality.

Inventive Principle:
Principle #6Universality (Multi-functionality)

3Reliability

If additional filtering operations are performed for spectral shaping, then speech quality is improved, but computational complexity increases

Engineering Contradiction:
Improvespeech qualityVSAvoidcomputational complexity
Core Design Contradiction:
ReliabilityVSPower

Solution Approach 1:

Spectral shaping is performed as a preliminary action within the rational delay element before the signal proceeds to subsequent processing stages. By pre-shaping the spectrum at the delay stage, additional filtering operations are eliminated, reducing computational complexity while maintaining speech quality improvement.

Inventive Principle:
Principle #10Preliminary action

4Measurement precision

If sub-sample resolution delay modeling is implemented, then measurement precision is improved, but ease of manufacture deteriorates

Engineering Contradiction:
Improvedelay value precisionVSAvoidimplementation difficulty
Core Design Contradiction:
Measurement precisionVSEase of manufacture

Solution Approach 1:

Instead of implementing complex physical sub-sample delay mechanisms, the patent uses a rational delay model that mathematically copies the effect of fractional delays. This approach achieves sub-sample precision through computational modeling rather than physical implementation, significantly easing the manufacturing and implementation process.

Inventive Principle:
Principle #26Copying

Data Source

PatentUS7792670B2Method and apparatus for speech coding
Publication Date: 2010.09.07 GOOGLE TECHNOLOGY HOLDINGS LLC
  • US7792670B2 patent drawing
  • US7792670B2 patent drawing
  • US7792670B2 patent drawing

AI summary

A method and apparatus for prediction in a speech-coding system is provided herein. The method of a 1st order long-term predictor (LTP) filter, using a sub-sample resolution delay, is extended to a multi-tap LTP filter, or, viewed from another vantage point, the conventional integer-sample resolution multi-tap LTP filter is extended to use sub-sample resolution delay. This novel formulation of a multi-tap LTP filter offers a number of advantages over the prior-art LTP filter configurations. Particularly, defining the lag with sub-sample resolution makes it possible to explicitly model the delay values that have a fractional component, within the limits of resolution of the over-sampling factor used by the interpolation filter. The coefficients of such a multi-tap LTP filter are thus largely freed from modeling the effect of delays that have a fractional component. Consequently their main function is to maximize the prediction gain of the LTP filter via modeling the degree of periodicity that is present and by imposing spectral shaping.