Spatial Audio Object Decoding With High-Resolution Un-Mixing

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current audio object coding schemes, such as MPEG SAOC, suffer from limited time-frequency selectivity, leading to audible artifacts like 'halo' effects around tonal sounds due to coarse frequency resolution, which cannot be improved without significantly increasing side information and compromising compatibility with existing standards.

Innovation Solution

An enhanced decoder and encoder system that modifies parametric side information to achieve higher frequency resolution, allowing for dynamic adjustment of time-frequency transforms based on audio object characteristics, enabling high frequency selectivity and temporal precision to minimize inter-object crosstalk and artifacts.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Manufacturing precision

If higher frequency resolution is used to improve spectral separation and reduce artifacts, then manufacturing precision is improved, but device complexity increases

Engineering Contradiction:
Improvespectral separation precisionVSAvoiddecoder complexity
Core Design Contradiction:
Manufacturing precisionVSDevice complexity

Solution Approach 1:

The patent segments the frequency resolution enhancement into separate processing stages: first decoding standard SAOC parameters, then optionally applying high-resolution spectral analysis only to specific frequency regions where artifacts occur. This allows precision improvement without requiring the entire decoder to operate at high complexity.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent applies high frequency resolution processing locally to specific frequency bands and time segments where halo artifacts are detected, rather than uniformly across the entire audio spectrum. This targeted approach improves spectral separation precision where needed while keeping the overall system complexity manageable.

Inventive Principle:
Principle #3Local quality

2Measurement precision

If dynamic time-frequency transform adjustment is implemented to improve frequency selectivity, then measurement precision is improved, but device complexity increases

Engineering Contradiction:
Improvefrequency selectivityVSAvoidtransform adjustment mechanism
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent implements dynamic adjustment of time-frequency transform parameters based on the detected characteristics of the audio signal. The system adapts the transform resolution and type according to the local signal properties, achieving high frequency selectivity when needed while maintaining lower complexity for stationary signals.

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

The system automatically detects signal characteristics and self-adjusts the transform parameters without requiring external control. The decoder analyzes the decoded audio objects and autonomously determines the optimal time-frequency resolution, reducing the burden on external control mechanisms.

Inventive Principle:
Principle #25Self-service

3Loss of information

If modified parametric side information is generated to achieve higher frequency resolution, then loss of information is reduced, but productivity decreases

Engineering Contradiction:
Improvespectral information lossVSAvoidencoding speed
Core Design Contradiction:
Loss of informationVSProductivity

Solution Approach 1:

The patent generates high-resolution parametric side information only for specific frequency bands and time segments where spectral accuracy is critical, rather than for the entire audio signal. This partial application reduces the computational burden on encoding while still recovering the most important spectral information that would otherwise be lost.

Inventive Principle:
Principle #16Partial or excessive action

Data Source

PatentUS11074920B2Encoder, decoder and methods for backward compatible multi-resolution spatial-audio-object-coding
Publication Date: 2021.07.27 FRAUNHOFER GESELLSCHAFT ZUR FORDERUNG DER ANGEWANDTEN FORSCHUNG EV
  • US11074920B2 patent drawing
  • US11074920B2 patent drawing
  • US11074920B2 patent drawing

AI summary

A decoder for generating an un-mixed audio signal including a plurality of un-mixed audio channels is provided. Moreover, an encoder and an encoded audio signal is provided. The decoder includes an un-mixing-information determiner for determining un-mixing information by receiving first parametric side information and second parametric side information on the at least one audio object signal, wherein the frequency resolution of the second parametric side information is higher than that of the first parametric side information. Moreover, the decoder includes an un-mix module for applying the un-mixing information on a downmix signal, to obtain an un-mixed audio signal including the plurality of un-mixed audio channels. The un-mixing-information determiner is configured to determine the un-mixing information by modifying the first parametric information and the second parametric information, such that the modified parametric information has a frequency resolution which is higher than the first frequency resolution.