Audio Decoder Up-Mixing for Foreground Background Separation
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current audio coding technologies, such as MPEG Surround, fail to effectively separate background and foreground audio signals in applications like Karaoke, where complete separation of audio objects is necessary, and do not allow for flexible loudspeaker configurations at the decoder side.
Innovation Solution
An enhanced audio decoding method that computes prediction coefficients based on level information and residual signals to improve the separation of audio objects, allowing for up-mixing of downmix signals into individual audio channels that can be rendered on any loudspeaker configuration, using a processor and up-mixer to approximate the original audio signals.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Quantity of substance
If MPEG Surround downmixing is used to compress multiple audio signals, then bit rate is reduced, but complete separation of background and foreground audio signals is not achieved
Solution Approach 1:
The patent segments audio objects into foreground and background categories with different downmixing strategies. Foreground objects (e.g., vocals) are preserved with higher fidelity while background objects (e.g., instruments) are more aggressively downmixed, enabling both bit rate reduction and effective separation for Karaoke applications.
Solution Approach 2:
Different quality levels are applied to different audio objects based on their importance. The encoder selectively maintains high quality for foreground objects while accepting lower quality for background objects, achieving efficient compression without compromising the primary audio content.
2Quantity of substance
If MPEG Surround encoder is used to downmix audio signals, then bit rate is reduced, but loudspeaker configuration cannot be changed at decoder side
Solution Approach 1:
The patent creates a universal downmix representation that can be decoded for multiple loudspeaker configurations. By encoding spatial relationships and object positions rather than configuration-specific mixing, the system enables flexible rendering on various speaker setups including stereo, surround, and mobile devices.
Solution Approach 2:
The decoder dynamically adapts the downmixed signal to different loudspeaker configurations in real-time. The system adjusts spatial rendering and channel assignment based on the target configuration, enabling the same encoded bitstream to be played back optimally on diverse hardware setups.
3Adaptability or versatility
If SAOC decoder is used to individually upmix downmix signal, then flexible loudspeaker configuration is enabled, but complete separation of foreground and background objects is not achieved
Solution Approach 1:
The patent applies different upmixing quality levels to foreground and background objects. Foreground objects receive high-quality individual upmixing with preserved spatial characteristics, while background objects are rendered with lower priority, achieving both configuration flexibility and effective object separation.
Solution Approach 2:
The upmixing process is segmented into foreground and background processing paths. This segmentation allows the decoder to apply separation-enhancing algorithms specifically to foreground objects while maintaining flexible spatial rendering for both categories across different loudspeaker configurations.
Data Source
AI summary
An audio decoder for decoding a multi-audio-object signal having an audio signal of a first type and an audio signal of a second type encoded therein is described, the multi-audio-object signal having a downmix signal and side information, the side information having level information of the audio signals of the first and second types in a first predetermined time/frequency resolution, and a residual signal specifying residual level values in a second predetermined time/frequency resolution, the audio decoder having a processor for computing prediction coefficients based on the level information; and an up-mixer for up-mixing the downmix signal based on the prediction coefficients and the residual signal to obtain a first up-mix audio signal approximating the audio signal of the first type and/or a second up-mix audio signal approximating the audio signal of the second type.


