Dynamic Time-Frequency Resolution Adaptation in Audio Object Decoders
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current audio object coding schemes, such as MPEG SAOC, are limited by fixed time-frequency resolution, leading to audible artifacts like crosstalk and roughness due to insufficient variability in time-frequency selectivity, which affects the quality of multi-channel audio processing.
Innovation Solution
The solution involves dynamically adapting the time-frequency resolution of the filter bank or transform based on the properties of the audio objects, allowing for high frequency selectivity in quasi-stationary signals and high temporal precision for transient events, while maintaining backward compatibility with standard SAOC data.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If fixed time-frequency resolution is used in audio object coding, then device complexity is reduced and backward compatibility is maintained, but audio quality deteriorates due to audible artifacts like crosstalk and roughness
Solution Approach 1:
The patent implements dynamic adaptation of time-frequency resolution by switching between different filter bank configurations (e.g., 64-subband vs. 128-subband) based on signal characteristics. This allows the system to adjust its resolution dynamically rather than being fixed, thereby improving audio quality while managing complexity through conditional adaptation rather than continuous adjustment
Solution Approach 2:
The system changes key parameters of the filter bank (number of subbands, window lengths) based on detected signal properties such as transient content. By modifying these parameters adaptively, the system resolves the contradiction between fixed complexity and variable performance, achieving high quality where needed without permanently increasing system complexity
2Reliability
If high frequency selectivity is used for quasi-stationary signals, then audio quality improves, but temporal precision deteriorates due to longer analysis windows
Solution Approach 1:
The patent dynamically adjusts the analysis window length and filter bank configuration based on the temporal characteristics of the signal. For quasi-stationary segments, longer windows provide high frequency selectivity, while for transient segments, shorter windows restore temporal precision. This dynamic switching resolves the trade-off between frequency and temporal resolution
Solution Approach 2:
Different time-frequency resolutions are applied locally to different segments of the audio signal based on their characteristics. Quasi-stationary portions receive high frequency resolution treatment, while transient portions receive high temporal resolution treatment. This localized adaptation allows each segment to be processed with optimal parameters without compromising overall performance
3Reliability
If high temporal precision is used for transient events, then audio quality improves, but frequency selectivity deteriorates due to shorter analysis windows
Solution Approach 1:
The system dynamically switches to shorter analysis windows and appropriate filter bank configurations when transients are detected, providing high temporal precision for accurate transient representation. After the transient event, the system returns to longer windows for high frequency selectivity, thus resolving the contradiction through time-varying parameter adjustment
4Reliability
If dynamic time-frequency resolution adaptation is implemented, then audio quality improves by minimizing artifacts, but bitrate increases due to additional side information
Solution Approach 1:
The system efficiently encodes the dynamic parameter changes (window lengths, filter bank configurations) by transmitting only the necessary side information indicating the chosen configuration rather than full parameter sets. This parameter-based adaptation minimizes the additional bitrate required while achieving dynamic resolution adjustment for improved perceptual quality
Data Source
Figure 1A
Figure 1B
Figure 1C
AI summary
A decoder for generating an audio output signal comprising one or more audio output channels from a downmix signal comprising a plurality of time-domain downmix samples is provided. The downmix signal encodes two or more audio object signals. The decoder comprises a window-sequence generator (134) for determining a plurality of analysis windows, wherein each of the analysis windows comprises a plurality of time-domain downmix samples of the downmix signal. Each analysis window of the plurality of analysis windows has a window length indicating the number of the time-domain downmix samples of said analysis window. The window-sequence generator (134) is configured to determine the plurality of analysis windows so that the window length of each of the analysis windows depends on a signal property of at least one of the two or more audio object signals. Moreover, the decoder comprises a t/f-analysis module (135) for transforming the plurality of time-domain downmix samples of each analysis window of the plurality of analysis windows from a time-domain to a time-frequency domain depending on the window length of said analysis window, to obtain a transformed downmix. Furthermore, the decoder comprises an un-mixing unit (136) for un-mixing the transformed downmix based on parametric side information on the two or more audio object signals to obtain the audio output signal. Moreover, an encoder is provided.