Audio Spectral Decoding with IGF for Low-Bitrate HF Detail

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current audio codecs face limitations in bandwidth extension techniques, particularly in maintaining high-frequency detail and timbre at low bitrates, due to restricted spectral patching and transformation requirements, leading to pre- or post-echoes and increased computational complexity.

Innovation Solution

The implementation of Intelligent Gap Filling (IGF) technology, which performs bandwidth extension in the same spectral domain as the core decoder, using temporal noise shaping and tile shaping to reduce echoes and enhance coding efficiency, while allowing for full-rate core decoding and decoding without the need for downsampling or upsampling.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If spectral patching is used for bandwidth extension, then high-frequency content can be reconstructed from low-frequency regions, but pre- or post-echoes occur and temporal continuity is disrupted

Engineering Contradiction:
Improvehigh-frequency detailVSAvoidpre- or post-echoes
Core Design Contradiction:
Measurement precisionVSObject-generated harmful factors

Solution Approach 1:

The patent applies temporal noise shaping by preprocessing the spectral patches before they are copied to the high-frequency region. A shaping filter is applied to the source spectral coefficients to pre-compensate for potential echo artifacts, ensuring that when the patches are transposed to the HF region, the temporal envelope continuity is maintained and echoes are reduced or eliminated.

Inventive Principle:
Principle #10Preliminary action

2Measurement precision

If bandwidth extension is performed using transformation to a second domain, then high-frequency reconstruction is enabled, but computational complexity increases

Engineering Contradiction:
Improvespectral reconstruction accuracyVSAvoidcomputational complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent merges the bandwidth extension process with the existing MDCT-based core decoding process by performing spectral patching and temporal noise shaping entirely within the MDCT domain. This eliminates the need for separate transformation stages to other domains, reducing computational complexity while maintaining spectral reconstruction accuracy through parameter-driven post-processing.

Inventive Principle:
Principle #5Merging (Combining)

3Ease of manufacture

If spectral patching is applied without parameter-driven post-processing, then bandwidth extension is simpler, but timbre and color of the original signal are not maintained

Engineering Contradiction:
Improvecoding simplicityVSAvoidtimbre preservation
Core Design Contradiction:
Ease of manufactureVSMeasurement precision

Solution Approach 1:

The patent applies parameter-driven post-processing to the transposed spectral patches, where parameters such as spectral envelope, tilt, and temporal envelope are adjusted to match the characteristics of the target high-frequency region. This ensures that the timbre and color of the original signal are preserved while maintaining coding efficiency through compact parameter representation.

Inventive Principle:
Principle #35Parameter changes

Data Source

PatentUS11049506B2Apparatus and method for encoding and decoding an encoded audio signal using temporal noise/patch shaping
Publication Date: 2021.06.29 FRAUNHOFER GESELLSCHAFT ZUR FORDERUNG DER ANGEWANDTEN FORSCHUNG EV
  • US11049506B2 patent drawing
  • US11049506B2 patent drawing
  • US11049506B2 patent drawing

AI summary

An apparatus for decoding an encoded audio signal, includes: a spectral domain audio decoder for generating a first decoded representation of a first set of first spectral portions being spectral prediction residual values; a frequency regenerator for generating a reconstructed second spectral portion using a first spectral portion of the first set of first spectral portions, wherein the reconstructed second spectral portion additionally includes spectral prediction residual values; and an inverse prediction filter for performing an inverse prediction over frequency using the spectral residual values for the first set of first spectral portions and the reconstructed second spectral portion using prediction filter information included in the encoded audio signal.