Audio Object Upmixing for Foreground-Background Separation

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current audio coding technologies, such as MPEG Surround and SAOC, fail to effectively separate background and foreground audio signals in applications like Karaoke, where complete separation of audio objects is necessary, due to their inability to handle uncorrelated audio signals and change loudspeaker configurations.

Innovation Solution

An audio decoder and encoder system that computes prediction coefficient matrices based on level information and inter-correlation parameters to up-mix a downmix signal, allowing for the separation of background and foreground audio signals by using a generalized TTT encoder and decoder structure, enabling flexible rendering on any loudspeaker configuration.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If MPEG Surround decoder is used to upmix downmix signal, then the original input channels can be recovered, but the loudspeaker configuration cannot be changed at the decoder's side

Engineering Contradiction:
Improverecovery of original channelsVSAvoidloudspeaker configuration flexibility
Core Design Contradiction:
ReliabilityVSAdaptability or versatility

Solution Approach 1:

The patent introduces a dynamic upmixing approach where the decoder can adapt to different loudspeaker configurations by computing prediction coefficient matrices based on level information and inter-correlation parameters. The system dynamically adjusts the upmixing process to match the desired output configuration rather than being fixed to the original encoding configuration.

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

The patent changes the parameters used in upmixing by computing prediction coefficient matrices from level information and inter-correlation parameters. This allows the system to transform the downmix signal according to different target configurations by modifying the prediction coefficients rather than using fixed transformation matrices.

Inventive Principle:
Principle #35Parameter changes

2Reliability

If SAOC decoder treats all objects equally, then individual objects can be recovered, but complete separation of background and foreground objects is not possible

Engineering Contradiction:
Improverecovery of individual objectsVSAvoidseparation precision of audio objects
Core Design Contradiction:
ReliabilityVSManufacturing precision

Solution Approach 1:

The patent applies different processing treatments to different types of audio objects. Foreground objects receive enhanced separation processing with computed prediction coefficients that prioritize their isolation from background objects, while background objects are processed differently. This local differentiation enables complete separation for Karaoke and solo applications.

Inventive Principle:
Principle #3Local quality

Solution Approach 2:

The patent segments the audio processing into distinct foreground and background object handling paths. By separating the processing of different object types and applying specific prediction coefficient computations for foreground objects, the system achieves complete separation capability that uniform treatment cannot provide.

Inventive Principle:
Principle #1Segmentation

3Quantity of substance

If downmixing is used to compress multiple audio signals, then bit rate is reduced, but the ability to separate uncorrelated signals is lost

Engineering Contradiction:
Improvebit rate reductionVSAvoidseparation accuracy of audio objects
Core Design Contradiction:
Quantity of substanceVSManufacturing precision

Solution Approach 1:

The patent uses feedback by computing prediction coefficients from level information and inter-correlation parameters that are derived from the downmix signal itself. This feedback mechanism allows the upmixing process to adapt to the actual content of the downmix signal, improving separation accuracy while maintaining bit rate efficiency.

Inventive Principle:
Principle #23Feedback

Solution Approach 2:

The patent replaces traditional mechanical upmixing approaches with a computational approach using prediction coefficient matrices. Instead of using fixed transformation matrices, the system computes adaptive coefficients based on signal characteristics, enabling better separation of uncorrelated signals while maintaining compression efficiency.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Data Source

PatentUS8155971B2Audio decoding of multi-audio-object signal using upmixing
Publication Date: 2012.04.10 FRAUNHOFER GESELLSCHAFT ZUR FORDERUNG DER ANGEWANDTEN FORSCHUNG EV
  • US8155971B2 patent drawing
  • US8155971B2 patent drawing
  • US8155971B2 patent drawing

AI summary

A method for decoding a multi-audio-object signal having audio signals of first and second types encoded therein, the multi-audio-object signal having a downmix signal and side information having level information of the audio signals of the first and second types in a first predetermined time/frequency resolution, the method including computing a prediction coefficient matrix C based on the level information; and up-mixing the downmix signal based on the prediction coefficients to obtain a first and/or a second up-mix audio signal approximating the audio signals of the first and second types, respectively, wherein up-mixing yields the first and/or second up-mix signals S1 and S2 from the downmix signal d according to a computation representable by(S1S2)=D-1⁢{(1C)⁢d+H},with “1” denoting—depending on the number of channels of d—a scalar, or an identity matrix, and D−1 being a matrix uniquely determined by a downmix prescription according to which the audio signals of the first and second types are downmixed into the downmix signal, and which is also included by the side information, and H being a term independent from d.