Stereo-Based Immersive Coding for Low-Bandwidth Spatial Audio
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing audio coding techniques struggle to deliver immersive audio content efficiently over limited bandwidth while maintaining high audio quality, particularly when using multi-channel formats like Dolby Atmos or MPEG-H, due to the need for larger bandwidth and the challenges of spectral distortions and image shifts during spatial rendering.
Innovation Solution
A stereo-based immersive audio coding system that uses a two-channel stereo signal and directional parameters to recreate immersive audio experiences, applying weighting factors in the frequency domain to minimize spectral distortions and decorrelate channel pairs, while reducing the number of channels transmitted.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If multi-channel audio formats (Dolby Atmos, MPEG-H) are used to deliver immersive audio content, then audio quality and immersive experience are improved, but bandwidth requirements increase
Solution Approach 1:
The patent extracts only the essential spatial information (directional parameters) from the full multi-channel audio signal. Instead of transmitting all channels, it separates the audio into a stereo downmix and extracts directional metadata that describes the spatial positioning of sound sources. This extracted metadata is then used at the decoder to reconstruct the immersive spatial experience, significantly reducing bandwidth while maintaining audio quality.
2Productivity
If bandwidth is reduced to lower data rates, then transmission efficiency is improved, but audio quality deteriorates
Solution Approach 1:
The patent changes the representation parameters of the audio signal by transforming it into a different domain. Instead of transmitting raw multi-channel audio data, it uses perceptual audio coding to represent the audio in terms of perceptually relevant parameters (spectral coefficients, directional parameters, spatial metadata). This parameter transformation allows for efficient compression while preserving the perceptual quality of the audio experience.
3Quantity of substance
If spatial rendering is applied to recreate immersive audio from stereo signals, then bandwidth is reduced, but spectral distortions and image shifts occur
Solution Approach 1:
The patent incorporates feedback mechanisms where the encoder analyzes the stereo downmix signal and adjusts the directional parameters based on the actual spatial characteristics present in the downmixed signal. The decoder uses this feedback information along with the transmitted directional metadata to accurately reconstruct the spatial positioning. This feedback loop ensures that spectral distortions and image shifts are minimized by adapting the spatial rendering to the actual content being transmitted.
Data Source
AI summary
Disclosed is an audio codec that represents an immersive signal by a two-channel stereo signal that is a stereo rendering of the immersive signal and directional parameters. The directional parameters may be based on a perceptual model describing the direction of virtual speaker pairs to recreate the perceived location of dominant sounds. Audio processing at the decoder may be performed on the stereo signal in the frequency domain for multiple channel pairs using time-frequency tiles. Spatial localization of the audio signals may use a panning approach by applying weightings to the time-frequency tiles of the stereo signal for each output channel pair. The weightings for the time-frequency tiles may be derived based on the directional parameters, an analysis of the stereo signal, and the output channel layout. The weightings may be used to adaptively process the time-frequency tiles using a de-correlator to reduce or minimize spectral distortions from spatial rendering.


