Audio Object Upmixing for Foreground-Background Separation
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current audio coding technologies, such as MPEG Surround and SAOC, fail to effectively separate background and foreground audio signals in applications like Karaoke, where complete separation of audio objects is necessary, due to their inability to handle uncorrelated audio signals and change loudspeaker configurations.
Innovation Solution
An audio decoder and encoder system that computes prediction coefficient matrices based on level information and inter-correlation parameters to up-mix a downmix signal, allowing for the separation of background and foreground audio signals by using a generalized TTT encoder and decoder structure, enabling flexible rendering on any loudspeaker configuration.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If MPEG Surround decoder is used to upmix downmix signal, then the original input channels can be recovered, but the loudspeaker configuration cannot be changed at the decoder's side
Solution Approach 1:
The patent introduces a dynamic upmixing approach where the decoder can adapt to different loudspeaker configurations by computing prediction coefficient matrices based on level information and inter-correlation parameters. The system dynamically adjusts the upmixing process to match the desired output configuration rather than being fixed to the original encoding configuration.
Solution Approach 2:
The patent changes the parameters used in upmixing by computing prediction coefficient matrices from level information and inter-correlation parameters. This allows the system to transform the downmix signal according to different target configurations by modifying the prediction coefficients rather than using fixed transformation matrices.
2Reliability
If SAOC decoder treats all objects equally, then individual objects can be recovered, but complete separation of background and foreground objects is not possible
Solution Approach 1:
The patent applies different processing treatments to different types of audio objects. Foreground objects receive enhanced separation processing with computed prediction coefficients that prioritize their isolation from background objects, while background objects are processed differently. This local differentiation enables complete separation for Karaoke and solo applications.
Solution Approach 2:
The patent segments the audio processing into distinct foreground and background object handling paths. By separating the processing of different object types and applying specific prediction coefficient computations for foreground objects, the system achieves complete separation capability that uniform treatment cannot provide.
3Quantity of substance
If downmixing is used to compress multiple audio signals, then bit rate is reduced, but the ability to separate uncorrelated signals is lost
Solution Approach 1:
The patent uses feedback by computing prediction coefficients from level information and inter-correlation parameters that are derived from the downmix signal itself. This feedback mechanism allows the upmixing process to adapt to the actual content of the downmix signal, improving separation accuracy while maintaining bit rate efficiency.
Solution Approach 2:
The patent replaces traditional mechanical upmixing approaches with a computational approach using prediction coefficient matrices. Instead of using fixed transformation matrices, the system computes adaptive coefficients based on signal characteristics, enabling better separation of uncorrelated signals while maintaining compression efficiency.
Data Source
AI summary
A method for decoding a multi-audio-object signal having audio signals of first and second types encoded therein, the multi-audio-object signal having a downmix signal and side information having level information of the audio signals of the first and second types in a first predetermined time/frequency resolution, the method including computing a prediction coefficient matrix C based on the level information; and up-mixing the downmix signal based on the prediction coefficients to obtain a first and/or a second up-mix audio signal approximating the audio signals of the first and second types, respectively, wherein up-mixing yields the first and/or second up-mix signals S1 and S2 from the downmix signal d according to a computation representable by(S1S2)=D-1{(1C)d+H},with “1” denoting—depending on the number of channels of d—a scalar, or an identity matrix, and D−1 being a matrix uniquely determined by a downmix prescription according to which the audio signals of the first and second types are downmixed into the downmix signal, and which is also included by the side information, and H being a term independent from d.


