Audio Spectral Gap Filling With Temporal Noise Shaping

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current audio codecs face limitations in bandwidth extension techniques, leading to loss of high-frequency detail and timbre, as well as increased computational complexity and memory requirements, especially in mobile devices, due to the need for transformation into new domains and limited temporal control of bandwidth extension signals.

Innovation Solution

The implementation of Intelligent Gap Filling (IGF) technology, which performs bandwidth extension in the same spectral domain as the core decoder, using temporal noise shaping (TNS) or temporal tile shaping (TTS) to reduce echoes and artifacts, and parametrically encoding spectral portions with different resolutions to efficiently fill spectral gaps.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If bandwidth extension techniques transform audio signal into new domains, then high-frequency content can be reconstructed, but computational complexity and memory requirements increase

Engineering Contradiction:
Improvehigh-frequency content reconstructionVSAvoidcomputational complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent combines bandwidth extension processing with the existing MDCT spectral domain, eliminating the need for separate transform domains. Frequency tiles are generated and processed within the same spectral representation, merging reconstruction operations with the core decoding pipeline to reduce computational overhead and memory requirements.

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The patent segments the spectral content into frequency tiles that can be independently processed and filled. By dividing the spectrum into manageable tile units, the system can efficiently reconstruct high-frequency content without processing the entire spectrum at once, reducing computational complexity while maintaining reconstruction quality.

Inventive Principle:
Principle #1Segmentation

2Measurement precision

If bandwidth extension uses spectral patching from low-frequency regions, then high-frequency spectrum can be filled, but temporal continuity and timbre may be compromised

Engineering Contradiction:
Improvespectral content fillingVSAvoidtemporal continuity
Core Design Contradiction:
Measurement precisionVSStability of the object's composition

Solution Approach 1:

The patent applies temporal noise shaping filters that dynamically adapt to the signal characteristics. The filtering operation is applied in the temporal domain to shape quantization noise and maintain temporal envelope consistency, allowing the system to preserve temporal continuity while filling spectral gaps through parameter-driven processing.

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

The patent uses parameter-driven post-processing to adjust spectral shape, tilt, and temporal characteristics. By modifying parameters such as spectral envelope and temporal filtering coefficients, the system can maintain timbre and temporal continuity while reconstructing high-frequency content from lower-frequency source material.

Inventive Principle:
Principle #35Parameter changes

3Productivity

If coarse quantization is used to reduce bitrate, then coding efficiency improves, but spectral gaps and quantization noise increase

Engineering Contradiction:
Improvecoding efficiencyVSAvoidspectral gaps
Core Design Contradiction:
ProductivityVSLoss of information

Solution Approach 1:

The patent converts quantization noise and spectral gaps into opportunities for intelligent gap filling. By identifying spectral regions with quantization artifacts, the system applies frequency tile filling and temporal noise shaping to transform these harmful artifacts into perceptually acceptable reconstructed content, effectively converting information loss into beneficial spectral completion.

Inventive Principle:
Principle #22Blessing in disguise (Convert harm into benefit)

Solution Approach 2:

The patent introduces frequency tiles as intermediary elements that bridge spectral gaps caused by coarse quantization. These tiles serve as intermediate representations that can be selectively filled from adjacent frequency regions, acting as mediators between the quantized spectral data and the final reconstructed audio signal.

Inventive Principle:
Principle #24Intermediary (Mediator)

4Measurement precision

If traditional bandwidth extension methods are used, then high-frequency reconstruction is achieved, but echoes and artifacts increase perceptual annoyance

Engineering Contradiction:
Improvehigh-frequency reconstructionVSAvoidechoes and artifacts
Core Design Contradiction:
Measurement precisionVSObject-generated harmful factors

Solution Approach 1:

The patent converts potential echo and artifact problems into opportunities for improvement by applying temporal noise shaping. The filtering operation shapes quantization noise to follow the temporal envelope of the signal, turning what would be audible artifacts into perceptually masked noise that enhances rather than degrades the reconstructed high-frequency content.

Inventive Principle:
Principle #22Blessing in disguise (Convert harm into benefit)

Data Source

PatentUS10332539B2Apparatus and method for encoding and decoding an encoded audio signal using temporal noise/patch shaping
Publication Date: 2019.06.25 FRAUNHOFER GESELLSCHAFT ZUR FORDERUNG DER ANGEWANDTEN FORSCHUNG EV
  • US10332539B2 patent drawing
  • US10332539B2 patent drawing
  • US10332539B2 patent drawing

AI summary

An apparatus for decoding an encoded audio signal, includes: a spectral domain audio decoder for generating a first decoded representation of a first set of first spectral portions being spectral prediction residual values; a frequency regenerator for generating a reconstructed second spectral portion using a first spectral portion of the first set of first spectral portions, wherein the reconstructed second spectral portion additionally includes spectral prediction residual values; and an inverse prediction filter for performing an inverse prediction over frequency using the spectral residual values for the first set of first spectral portions and the reconstructed second spectral portion using prediction filter information included in the encoded audio signal.