Encoded Audio Scene Bandwidth Extension for Low-Delay Stereo
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing audio processing systems for immersive communication, such as DirAC, face high delays and complexity when converting mono recordings to stereo output due to the use of filter banks, which is suboptimal for low-bitrate scenarios.
Innovation Solution
A method involving a parameter converter that converts DirAC side parameters into stereo parameters using a Short-Time Fourier Transform (STFT) filterbank, allowing parallel processing to achieve low-delay upmixing without additional delay, and combining this with ACELP speech coder processing to maintain the same overall delay as EVS codecs.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If DirAC uses filter-bank domain processing for spatial restoration, then spatial information can be restored from mono channel, but processing delay increases significantly
Solution Approach 1:
The patent replaces the traditional filter-bank domain spatial restoration with a time-domain approach using impulse response convolution. Instead of using complex filter banks to restore spatial information, the system uses a simplified impulse response-based method that operates in the time domain, significantly reducing processing delay while maintaining spatial restoration capability.
Solution Approach 2:
The patent changes the processing domain from frequency domain (filter-bank) to time domain (impulse response). By transforming the spatial restoration mechanism from filter-bank domain to time-domain impulse response convolution, the system achieves the same spatial restoration function with much lower delay, as the impulse response method requires no additional filter-bank analysis/synthesis steps.
2Adaptability or versatility
If FOA upmix with L/R conversion is used for stereo output, then stereo playback is achieved, but overall delay increases compared to direct stereo extraction
Solution Approach 1:
The patent extracts and removes the unnecessary FOA upmix and L/R conversion steps from the traditional DirAC processing chain. By directly extracting stereo information from the encoded audio scene using a simplified time-domain method, the system eliminates the intermediate FOA representation step that causes additional delay, while still achieving high-quality stereo playback.
3Adaptability or versatility
If CLDFB filter-bank analysis/synthesis is used for spatial rendering, then immersive audio is achieved, but processing complexity and delay increase
Solution Approach 1:
The patent substitutes the complex CLDFB filter-bank analysis/synthesis system with a simple time-domain impulse response convolution system. Instead of using sophisticated filter banks to achieve immersive audio rendering, the system uses a computationally efficient impulse response-based approach that achieves the same rendering goal with much lower complexity and delay.
Data Source
AI summary
Apparatus for processing an audio scene representing a sound field, the audio scene comprising information on a transport signal and a set of parameters. The apparatus comprising an output interface for generating a processed audio scene using the set of parameters and the information on the transport signal, wherein the output interface is configured to generate a raw representation of two or more channels using the set of parameters and the transport signal and a multichannel enhancer for generating an enhancement representation of the two or more channels using the transport signal, and a signal combiner for combining the raw representation of the two or more channels and the enhancement representation of the two or more channels to obtain the processed audio scene.


