Audio Object Decoding With Residual Up-Mixing for Signal Separation
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current audio coding technologies, such as MPEG Surround, fail to effectively separate background and foreground audio signals in applications like Karaoke, where complete separation of audio objects is necessary, and do not allow for flexible loudspeaker configurations at the decoder side.
Innovation Solution
An enhanced audio decoding and encoding method that computes prediction coefficients based on level and inter-correlation information, and uses residual signals to improve up-mixing, allowing for better separation of audio objects and rendering on any loudspeaker configuration, utilizing a generalized TTT encoder and decoder structure.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Quantity of substance
If MPEG Surround downmixing is used to compress multiple audio signals, then bit rate is reduced, but complete separation of background and foreground audio objects cannot be achieved
Solution Approach 1:
The patent segments audio objects into foreground and background categories with different downmixing strategies. Foreground objects use individual downmixing to preserve separability, while background objects are downmixed separately. This segmentation enables both bit rate reduction and complete audio object separation by treating different audio components differently in the encoding process.
2Device complexity
If MPEG Surround encoder is used with fixed loudspeaker configuration, then decoding is simplified, but decoder cannot adapt to different loudspeaker configurations
Solution Approach 1:
The patent creates a universal downmixing framework that works across multiple loudspeaker configurations. By using object-based downmixing with separate foreground and background processing, the system can adapt to different loudspeaker setups (stereo, 5.1, 7.1, etc.) without requiring configuration-specific encoders. The decoder receives flexible side information that enables adaptation to various output configurations while maintaining manageable complexity through standardized processing blocks.
3Device complexity
If SAOC standard treats all audio objects equally, then encoding is simplified, but complete separation of foreground and background objects is not possible
Solution Approach 1:
The patent applies different quality levels and processing strategies to different audio objects based on their importance. Foreground objects receive individual downmixing with preserved separation characteristics, while background objects are downmixed separately with different parameter handling. This local quality differentiation enables complete separation of foreground and background objects while keeping encoding complexity manageable through systematic classification and targeted processing.
Data Source
AI summary
An audio decoder for decoding a multi-audio-object signal having an audio signal of a first type and an audio signal of a second type encoded therein is described, the multi-audio-object signal having a downmix signal and side information, the side information having level information of the audio signals of the first and second types in a first predetermined time/frequency resolution, and a residual signal specifying residual level values in a second predetermined time/frequency resolution, the audio decoder having a processor for computing prediction coefficients based on the level information; and an up-mixer for up-mixing the downmix signal based on the prediction coefficients and the residual signal to obtain a first up-mix audio signal approximating the audio signal of the first type and/or a second up-mix audio signal approximating the audio signal of the second type.


