Audio Object Decoding With Residual Up-Mixing for Signal Separation

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current audio coding technologies, such as MPEG Surround, fail to effectively separate background and foreground audio signals in applications like Karaoke, where complete separation of audio objects is necessary, and do not allow for flexible loudspeaker configurations at the decoder side.

Innovation Solution

An enhanced audio decoding and encoding method that computes prediction coefficients based on level and inter-correlation information, and uses residual signals to improve up-mixing, allowing for better separation of audio objects and rendering on any loudspeaker configuration, utilizing a generalized TTT encoder and decoder structure.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Quantity of substance

If MPEG Surround downmixing is used to compress multiple audio signals, then bit rate is reduced, but complete separation of background and foreground audio objects cannot be achieved

Engineering Contradiction:
Improvebit rateVSAvoidaudio object separation quality
Core Design Contradiction:
Quantity of substanceVSManufacturing precision

Solution Approach 1:

The patent segments audio objects into foreground and background categories with different downmixing strategies. Foreground objects use individual downmixing to preserve separability, while background objects are downmixed separately. This segmentation enables both bit rate reduction and complete audio object separation by treating different audio components differently in the encoding process.

Inventive Principle:
Principle #1Segmentation

2Device complexity

If MPEG Surround encoder is used with fixed loudspeaker configuration, then decoding is simplified, but decoder cannot adapt to different loudspeaker configurations

Engineering Contradiction:
Improvedecoder complexityVSAvoidloudspeaker configuration flexibility
Core Design Contradiction:
Device complexityVSAdaptability or versatility

Solution Approach 1:

The patent creates a universal downmixing framework that works across multiple loudspeaker configurations. By using object-based downmixing with separate foreground and background processing, the system can adapt to different loudspeaker setups (stereo, 5.1, 7.1, etc.) without requiring configuration-specific encoders. The decoder receives flexible side information that enables adaptation to various output configurations while maintaining manageable complexity through standardized processing blocks.

Inventive Principle:
Principle #6Universality (Multi-functionality)

3Device complexity

If SAOC standard treats all audio objects equally, then encoding is simplified, but complete separation of foreground and background objects is not possible

Engineering Contradiction:
Improveencoding complexityVSAvoidaudio object separation completeness
Core Design Contradiction:
Device complexityVSManufacturing precision

Solution Approach 1:

The patent applies different quality levels and processing strategies to different audio objects based on their importance. Foreground objects receive individual downmixing with preserved separation characteristics, while background objects are downmixed separately with different parameter handling. This local quality differentiation enables complete separation of foreground and background objects while keeping encoding complexity manageable through systematic classification and targeted processing.

Inventive Principle:
Principle #3Local quality

Data Source

PatentUS8407060B2Audio decoder, audio object encoder, method for decoding a multi-audio-object signal, multi-audio-object encoding method, and non-transitory computer-readable medium therefor
Publication Date: 2013.03.26 FRAUNHOFER GESELLSCHAFT ZUR FORDERUNG DER ANGEWANDTEN FORSCHUNG EV
  • US8407060B2 patent drawing
  • US8407060B2 patent drawing
  • US8407060B2 patent drawing

AI summary

An audio decoder for decoding a multi-audio-object signal having an audio signal of a first type and an audio signal of a second type encoded therein is described, the multi-audio-object signal having a downmix signal and side information, the side information having level information of the audio signals of the first and second types in a first predetermined time/frequency resolution, and a residual signal specifying residual level values in a second predetermined time/frequency resolution, the audio decoder having a processor for computing prediction coefficients based on the level information; and an up-mixer for up-mixing the downmix signal based on the prediction coefficients and the residual signal to obtain a first up-mix audio signal approximating the audio signal of the first type and/or a second up-mix audio signal approximating the audio signal of the second type.