Dynamic Time-Frequency Resolution Adaptation in Audio Object Decoders

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current audio object coding schemes, such as MPEG SAOC, are limited by fixed time-frequency resolution, leading to audible artifacts like crosstalk and roughness due to insufficient variability in time-frequency selectivity, which affects the quality of multi-channel audio processing.

Innovation Solution

The solution involves dynamically adapting the time-frequency resolution of the filter bank or transform based on the properties of the audio objects, allowing for high frequency selectivity in quasi-stationary signals and high temporal precision for transient events, while maintaining backward compatibility with standard SAOC data.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If fixed time-frequency resolution is used in audio object coding, then device complexity is reduced and backward compatibility is maintained, but audio quality deteriorates due to audible artifacts like crosstalk and roughness

Engineering Contradiction:
Improveaudio qualityVSAvoidtime-frequency resolution adaptation mechanism
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The patent implements dynamic adaptation of time-frequency resolution by switching between different filter bank configurations (e.g., 64-subband vs. 128-subband) based on signal characteristics. This allows the system to adjust its resolution dynamically rather than being fixed, thereby improving audio quality while managing complexity through conditional adaptation rather than continuous adjustment

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

The system changes key parameters of the filter bank (number of subbands, window lengths) based on detected signal properties such as transient content. By modifying these parameters adaptively, the system resolves the contradiction between fixed complexity and variable performance, achieving high quality where needed without permanently increasing system complexity

Inventive Principle:
Principle #35Parameter changes

2Reliability

If high frequency selectivity is used for quasi-stationary signals, then audio quality improves, but temporal precision deteriorates due to longer analysis windows

Engineering Contradiction:
Improvefrequency selectivityVSAvoidtemporal precision
Core Design Contradiction:
ReliabilityVSLoss of time

Solution Approach 1:

The patent dynamically adjusts the analysis window length and filter bank configuration based on the temporal characteristics of the signal. For quasi-stationary segments, longer windows provide high frequency selectivity, while for transient segments, shorter windows restore temporal precision. This dynamic switching resolves the trade-off between frequency and temporal resolution

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

Different time-frequency resolutions are applied locally to different segments of the audio signal based on their characteristics. Quasi-stationary portions receive high frequency resolution treatment, while transient portions receive high temporal resolution treatment. This localized adaptation allows each segment to be processed with optimal parameters without compromising overall performance

Inventive Principle:
Principle #3Local quality

3Reliability

If high temporal precision is used for transient events, then audio quality improves, but frequency selectivity deteriorates due to shorter analysis windows

Engineering Contradiction:
Improvetemporal precisionVSAvoidfrequency selectivity
Core Design Contradiction:
ReliabilityVSMeasurement precision

Solution Approach 1:

The system dynamically switches to shorter analysis windows and appropriate filter bank configurations when transients are detected, providing high temporal precision for accurate transient representation. After the transient event, the system returns to longer windows for high frequency selectivity, thus resolving the contradiction through time-varying parameter adjustment

Inventive Principle:
Principle #15Dynamics

4Reliability

If dynamic time-frequency resolution adaptation is implemented, then audio quality improves by minimizing artifacts, but bitrate increases due to additional side information

Engineering Contradiction:
Improveperceptual qualityVSAvoidbitrate
Core Design Contradiction:
ReliabilityVSQuantity of substance

Solution Approach 1:

The system efficiently encodes the dynamic parameter changes (window lengths, filter bank configurations) by transmitting only the necessary side information indicating the chosen configuration rather than full parameter sets. This parameter-based adaptation minimizes the additional bitrate required while achieving dynamic resolution adjustment for improved perceptual quality

Inventive Principle:
Principle #35Parameter changes

Data Source

PatentEP2904611B1Encoder, decoder and methods for backward compatible dynamic adaption of time/frequency resolution in spatial-audio-object-coding
Publication Date: 2021.06.23 FRAUNHOFER GESELLSCHAFT ZUR FORDERUNG DER ANGEWANDTEN FORSCHUNG EV
  • EP2904611B1 patent drawingFigure 1A
  • EP2904611B1 patent drawingFigure 1B
  • EP2904611B1 patent drawingFigure 1C

AI summary

A decoder for generating an audio output signal comprising one or more audio output channels from a downmix signal comprising a plurality of time-domain downmix samples is provided. The downmix signal encodes two or more audio object signals. The decoder comprises a window-sequence generator (134) for determining a plurality of analysis windows, wherein each of the analysis windows comprises a plurality of time-domain downmix samples of the downmix signal. Each analysis window of the plurality of analysis windows has a window length indicating the number of the time-domain downmix samples of said analysis window. The window-sequence generator (134) is configured to determine the plurality of analysis windows so that the window length of each of the analysis windows depends on a signal property of at least one of the two or more audio object signals. Moreover, the decoder comprises a t/f-analysis module (135) for transforming the plurality of time-domain downmix samples of each analysis window of the plurality of analysis windows from a time-domain to a time-frequency domain depending on the window length of said analysis window, to obtain a transformed downmix. Furthermore, the decoder comprises an un-mixing unit (136) for un-mixing the transformed downmix based on parametric side information on the two or more audio object signals to obtain the audio output signal. Moreover, an encoder is provided.