Audio Bandwidth Extension Using IGF and Temporal Noise Shaping

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current audio codecs face limitations in bandwidth extension techniques, particularly in maintaining high-frequency detail and timbre, due to restricted transformation domains and limited temporal control, leading to pre- or post-echoes and increased computational complexity.

Innovation Solution

The implementation of Intelligent Gap Filling (IGF) technology, which combines Temporal Noise Shaping (TNS) or Temporal Tile Shaping (TTS) with high-frequency reconstruction, allowing bandwidth extension in the same spectral domain as the core decoder, and using frequency tiles for gap filling, thereby reducing echoes and computational complexity.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If bandwidth extension techniques are used to reduce bitrate, then coding efficiency is improved, but temporal continuity and audio quality deteriorate due to pre- or post-echoes

Engineering Contradiction:
Improvecoding efficiencyVSAvoidtemporal continuity
Core Design Contradiction:
ProductivityVSReliability

Solution Approach 1:

The patent applies temporal noise shaping by predicting and removing noise components before they manifest as audible artifacts. The predictor analyzes the spectral content and predicts noise in upcoming frames, allowing pre-echoes to be removed before they occur. Similarly, post-echoes are predicted and removed in advance, preventing temporal discontinuities that would degrade audio quality.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent implements a feedback mechanism where the predictor continuously monitors the decoded audio signal for noise components and feeds this information back to the noise shaper. The system analyzes the spectral content of previous frames and uses this feedback to adjust noise prediction and removal in current frames, maintaining temporal continuity while preserving coding efficiency.

Inventive Principle:
Principle #23Feedback

2Measurement precision

If transformation to second domain is applied for high-frequency reconstruction, then spectral accuracy is improved, but device complexity increases

Engineering Contradiction:
Improvespectral accuracyVSAvoidcomputational complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent merges the noise prediction and spectral shaping operations into a unified temporal noise shaping process that operates within the same transform domain as the audio coder. Instead of transforming to a second domain for high-frequency reconstruction, the system combines predictive noise removal with spectral envelope modification in the MDCT domain, reducing computational complexity while maintaining spectral accuracy.

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The patent introduces a spectral envelope as an intermediary that mediates between the time-domain audio signal and the frequency-domain representation. The temporal noise shaper modifies the spectral envelope to remove noise components without requiring transformation to a second domain. This intermediary approach maintains spectral accuracy while avoiding the computational overhead of multiple transform domains.

Inventive Principle:
Principle #24Intermediary (Mediator)

Data Source

PatentEP2883227B1Apparatus and method for encoding and decoding an encoded audio signal using temporal noise/patch shaping
Publication Date: 2016.08.17 FRAUNHOFER GESELLSCHAFT ZUR FORDERUNG DER ANGEWANDTEN FORSCHUNG EV
  • EP2883227B1 patent drawingFigure 1A~1B
  • EP2883227B1 patent drawingFigure 2A
  • EP2883227B1 patent drawingFigure 2B

AI summary

An apparatus for decoding an encoded audio signal, comprises: a spectral domain audio decoder (602) for generating a first decoded representation of a first set of first spectral portions being spectral prediction residual values; a frequency regenerator (604) for generating a reconstructed second spectral portion using a first spectral portion of the first set of first spectral portions, wherein the reconstructed second spectral portion additionally comprises spectral prediction residual values; and an inverse prediction filter (606) for performing an inverse prediction over frequency using the spectral residual values for the first set of first spectral portions and the reconstructed second spectral portion using prediction filter information (607) included in the encoded audio signal.