Audio Object Separation Using Adaptive Time-Frequency Resolution

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current audio object coding techniques face challenges in achieving optimal time-frequency representation for different audio objects, leading to suboptimal separation performance and crosstalk issues, as they often use fixed time-frequency resolutions that do not match the characteristics of the audio objects, resulting in impaired sound quality.

Innovation Solution

The proposed solution involves an audio decoder and method that allow for the selection of the most suitable time-frequency representation for each audio object based on its specific characteristics, using an Enhanced Side Information Estimator and Enhanced Object Separator to compute and apply object-specific side information, enabling improved separation performance and subjective quality.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Device complexity

If fixed time-frequency resolution is used for audio object coding, then device complexity is reduced, but separation performance deteriorates due to mismatch with audio object characteristics

Engineering Contradiction:
Improvetime-frequency representation complexityVSAvoidseparation performance
Core Design Contradiction:
Device complexityVSMeasurement precision

Solution Approach 1:

The patent applies dynamics by transitioning from fixed time-frequency resolution to adaptive time-frequency resolution that changes according to audio object characteristics. The system dynamically selects different time-frequency resolutions based on the temporal and spectral properties of each audio object, allowing the representation to adapt to the specific requirements of different audio content types.

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

The patent implements local quality by applying different time-frequency resolutions to different audio objects or different portions of audio objects based on their local characteristics. This allows high temporal resolution for transient components and high spectral resolution for tonal components, optimizing the representation quality for each specific region of the audio signal.

Inventive Principle:
Principle #3Local quality

2Measurement precision

If adaptive time-frequency representation is used for each audio object, then separation performance is improved, but device complexity increases

Engineering Contradiction:
Improveseparation performanceVSAvoidtime-frequency representation complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent applies parameter changes by modifying the time-frequency resolution parameters according to the characteristics of each audio object. The system changes parameters such as window size, overlap factor, and filter bank configuration to match the temporal and spectral properties of different audio objects, thereby optimizing separation performance without requiring a complete redesign of the system architecture.

Inventive Principle:
Principle #35Parameter changes

3Measurement precision

If high bitrate transmission is used for multi-channel audio content, then audio quality is improved, but resource load increases

Engineering Contradiction:
Improveaudio qualityVSAvoidbitrate
Core Design Contradiction:
Measurement precisionVSQuantity of substance

Solution Approach 1:

The patent applies parameter changes by using parametric techniques to represent audio objects efficiently. Instead of transmitting full-resolution multi-channel audio data, the system transmits compressed parametric information including time-frequency resolution indicators, spatial parameters, and object characteristics, which can then be used to reconstruct the audio objects at the receiver side with high quality but low bitrate.

Inventive Principle:
Principle #35Parameter changes

Data Source

PatentEP2997572B1Audio object separation from mixture signal using object-specific time/frequency resolutions
Publication Date: 2023.01.04 FRAUNHOFER GESELLSCHAFT ZUR FORDERUNG DER ANGEWANDTEN FORSCHUNG EV
  • EP2997572B1 patent drawingFigure 1
  • EP2997572B1 patent drawingFigure 2
  • EP2997572B1 patent drawingFigure 3

AI summary

An audio decoder is proposed for decoding a multi-object audio signal consisting of a downmix signal X and side information PSI. The side information comprises object-specific side information PSIi, for an audio object Si in a time/frequency region R(tR,fR), and object-specific time/frequency resolution information TFRIi indicative of an object-specific time/frequency resolution TFRh of the object-specific side information for the audio object Si in the time/frequency region Κ(tR,fR). The audio decoder comprises an object-specific time/frequency resolution determiner 110 configured to determine the object-specific time/frequency resolution information TFRIi from the side information PSI for the audio object Si . The audio decoder further comprises an object separator 120 configured to separate the audio object si from the downmix signal X using the object-specific side information in accordance with the object-specific time/frequency resolution TFRIi. A corresponding encoder and corresponding methods for decoding or encoding are also described.