Audio Spectral Coefficient Coding With Shape-Adaptive Contexts

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing context-based arithmetic coding for spectral coefficients in audio signals is constrained by memory requirements, computational complexity, and robustness to channel errors, leading to lower coding efficiency, especially for tonal signals where the context has to be limited to exploit the harmonic structure.

Innovation Solution

A context-adaptive entropy coding method that adjusts the relative spectral distance between decoded/encoded spectral coefficients based on the shape of the audio signal's spectrum, using information such as pitch, inter-harmonic distance, and formant locations to adapt the spectral neighborhood for improved coding efficiency.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If context-based arithmetic coding is used to encode spectral coefficients, then coding efficiency is improved, but memory requirements and computational complexity increase

Engineering Contradiction:
Improvecoding efficiencyVSAvoidmemory requirements and computational complexity
Core Design Contradiction:
ProductivityVSDevice complexity

Solution Approach 1:

The patent applies local quality by using different context models for different spectral regions. Instead of using a single uniform context model for all spectral coefficients, the invention divides the spectrum into different regions (e.g., low-frequency and high-frequency regions) and applies specialized context models to each region. This allows the system to achieve high coding efficiency for tonal signals in specific frequency ranges while keeping the overall computational complexity and memory requirements manageable by not applying complex modeling to all regions uniformly.

Inventive Principle:
Principle #3Local quality

2Device complexity

If the context is limited to meet memory and complexity constraints, then device complexity is reduced, but coding gain decreases especially for tonal signals

Engineering Contradiction:
Improvememory and complexity constraintsVSAvoidcoding gain
Core Design Contradiction:
Device complexityVSProductivity

Solution Approach 1:

The patent implements local quality by creating region-specific context models that are tailored to the characteristics of different spectral regions. For tonal signals in low-frequency regions, the invention uses context models that exploit harmonic structures with appropriate spectral distances. For high-frequency regions, different context models are applied. This regional specialization enables the system to achieve high coding gain for tonal signals where it matters most while keeping the overall system complexity within acceptable bounds.

Inventive Principle:
Principle #3Local quality

Solution Approach 2:

The patent applies dynamics by making the context model adaptive and configurable. The system can dynamically select and adjust context model parameters based on the signal characteristics being encoded. For tonal signals, the context model can be configured to exploit harmonic structures by adjusting spectral distance parameters. This dynamic adaptability allows the system to optimize coding gain for different signal types without requiring a single overly complex fixed model.

Inventive Principle:
Principle #15Dynamics

3Speed

If low-overlap windows are used to decrease algorithmic delay, then processing speed is improved, but leakage in MDCT increases resulting in higher quantization noise

Engineering Contradiction:
Improvealgorithmic delayVSAvoidleakage and quantization noise
Core Design Contradiction:
SpeedVSObject-generated harmful factors

Solution Approach 1:

The patent applies the blessing in disguise principle by converting the harmful effects of spectral leakage into beneficial information for coding. Instead of trying to eliminate leakage through complex windowing functions, the invention exploits the structured patterns that arise from leakage in tonal signals. The context-based arithmetic coding is designed to model and exploit these leakage-induced patterns, particularly the harmonic structures that emerge even with low-overlap windows. This transforms what would normally be harmful quantization noise into predictable, compressible patterns that improve overall coding efficiency.

Inventive Principle:
Principle #22Blessing in disguise (Convert harm into benefit)

Data Source

PatentUS9892735B2Coding of spectral coefficients of a spectrum of an audio signal
Publication Date: 2018.02.13 FRAUNHOFER GESELLSCHAFT ZUR FORDERUNG DER ANGEWANDTEN FORSCHUNG EV
  • US9892735B2 patent drawing
  • US9892735B2 patent drawing
  • US9892735B2 patent drawing

AI summary

A coding efficiency of coding spectral coefficients of a spectrum of an audio signal is increased by en/decoding a currently to be en/decoded spectral coefficient by entropy en/decoding and, in doing so, performing the entropy en/decoding depending, in a context-adaptive manner, on a previously en/decoded spectral coefficient, while adjusting a relative spectral distance between the previously en/decoded spectral coefficient and the currently en/decoded spectral coefficient depending on an information concerning a shape of the spectrum. The information concerning the shape of the spectrum may have a measure of a pitch or periodicity of the audio signal, a measure of an inter-harmonic distance of the audio signal's spectrum and/or relative locations of formants and/or valleys of a spectral envelope of the spectrum, and on the basis of this knowledge, the spectral neighborhood which is exploited in order to form the context of the currently to be en/decoded spectral coefficients may be adapted to the thus determined shape of the spectrum, thereby enhancing the entropy coding efficiency.