Audio Decoder Up-Mixing for Foreground Background Separation

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current audio coding technologies, such as MPEG Surround, fail to effectively separate background and foreground audio signals in applications like Karaoke, where complete separation of audio objects is necessary, and do not allow for flexible loudspeaker configurations at the decoder side.

Innovation Solution

An enhanced audio decoding method that computes prediction coefficients based on level information and residual signals to improve the separation of audio objects, allowing for up-mixing of downmix signals into individual audio channels that can be rendered on any loudspeaker configuration, using a processor and up-mixer to approximate the original audio signals.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Quantity of substance

If MPEG Surround downmixing is used to compress multiple audio signals, then bit rate is reduced, but complete separation of background and foreground audio signals is not achieved

Engineering Contradiction:
Improvebit rateVSAvoidseparation precision of audio objects
Core Design Contradiction:
Quantity of substanceVSManufacturing precision

Solution Approach 1:

The patent segments audio objects into foreground and background categories with different downmixing strategies. Foreground objects (e.g., vocals) are preserved with higher fidelity while background objects (e.g., instruments) are more aggressively downmixed, enabling both bit rate reduction and effective separation for Karaoke applications.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

Different quality levels are applied to different audio objects based on their importance. The encoder selectively maintains high quality for foreground objects while accepting lower quality for background objects, achieving efficient compression without compromising the primary audio content.

Inventive Principle:
Principle #3Local quality

2Quantity of substance

If MPEG Surround encoder is used to downmix audio signals, then bit rate is reduced, but loudspeaker configuration cannot be changed at decoder side

Engineering Contradiction:
Improvebit rateVSAvoidloudspeaker configuration flexibility
Core Design Contradiction:
Quantity of substanceVSAdaptability or versatility

Solution Approach 1:

The patent creates a universal downmix representation that can be decoded for multiple loudspeaker configurations. By encoding spatial relationships and object positions rather than configuration-specific mixing, the system enables flexible rendering on various speaker setups including stereo, surround, and mobile devices.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The decoder dynamically adapts the downmixed signal to different loudspeaker configurations in real-time. The system adjusts spatial rendering and channel assignment based on the target configuration, enabling the same encoded bitstream to be played back optimally on diverse hardware setups.

Inventive Principle:
Principle #15Dynamics

3Adaptability or versatility

If SAOC decoder is used to individually upmix downmix signal, then flexible loudspeaker configuration is enabled, but complete separation of foreground and background objects is not achieved

Engineering Contradiction:
Improveloudspeaker configuration flexibilityVSAvoidseparation precision of audio objects
Core Design Contradiction:
Adaptability or versatilityVSManufacturing precision

Solution Approach 1:

The patent applies different upmixing quality levels to foreground and background objects. Foreground objects receive high-quality individual upmixing with preserved spatial characteristics, while background objects are rendered with lower priority, achieving both configuration flexibility and effective object separation.

Inventive Principle:
Principle #3Local quality

Solution Approach 2:

The upmixing process is segmented into foreground and background processing paths. This segmentation allows the decoder to apply separation-enhancing algorithms specifically to foreground objects while maintaining flexible spatial rendering for both categories across different loudspeaker configurations.

Inventive Principle:
Principle #1Segmentation

Data Source

PatentUS8280744B2Audio decoder, audio object encoder, method for decoding a multi-audio-object signal, multi-audio-object encoding method, and non-transitory computer-readable medium therefor
Publication Date: 2012.10.02 FRAUNHOFER GESELLSCHAFT ZUR FORDERUNG DER ANGEWANDTEN FORSCHUNG EV
  • US8280744B2 patent drawing
  • US8280744B2 patent drawing
  • US8280744B2 patent drawing

AI summary

An audio decoder for decoding a multi-audio-object signal having an audio signal of a first type and an audio signal of a second type encoded therein is described, the multi-audio-object signal having a downmix signal and side information, the side information having level information of the audio signals of the first and second types in a first predetermined time/frequency resolution, and a residual signal specifying residual level values in a second predetermined time/frequency resolution, the audio decoder having a processor for computing prediction coefficients based on the level information; and an up-mixer for up-mixing the downmix signal based on the prediction coefficients and the residual signal to obtain a first up-mix audio signal approximating the audio signal of the first type and/or a second up-mix audio signal approximating the audio signal of the second type.