Multichannel Audio Compression via Spatial Resolution
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current multichannel audio stream compression techniques fail to efficiently reduce bit rate while maintaining sound quality, especially for streams with 5 to 7 channels, as they do not effectively exploit redundancies and are based on unsuitable hypotheses about spatial perception.
Innovation Solution
A method that identifies audio sources in a sound scene, determines their spatial resolution based on psycho-perceptive properties, and generates a compressed stream with information necessary for restoration, exploiting interactions between sources and using spherical harmonics representation to reduce spatial resolution.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Manufacturing precision
If separate encoding of multichannel signals is used, then sound quality is maintained at high bit rates, but compression efficiency deteriorates and bit rate cannot be reduced to 64 kbits/s for 5-7 channels
Solution Approach 1:
The patent combines multiple audio channels into a reduced set of channels (downmix) before encoding, merging spatial information from 5-7 channels into fewer channels while preserving essential spatial perception. This allows efficient compression at 64 kbits/s while maintaining perceived sound quality through selective preservation of spatial characteristics.
Solution Approach 2:
The patent transforms the audio representation by changing parameters from individual channel signals to spatial descriptors (ICTD, ICLD, ICC) that capture essential spatial perception. This parameter transformation enables efficient compression by encoding only the most perceptually relevant spatial characteristics rather than full channel signals.
2Productivity
If downmix to mono or stereo is used, then bit rate is reduced to 64 kbits/s, but spatial perception accuracy deteriorates due to unsuitable monophonic perception hypotheses
Solution Approach 1:
The patent applies different encoding strategies to different spatial regions and sources. By identifying individual sound sources and their spatial characteristics, it preserves high spatial resolution for prominent sources while using coarser encoding for less important sources, optimizing the balance between bit rate and spatial perception accuracy.
Solution Approach 2:
The patent dynamically adapts the encoding precision based on the spatial characteristics and perceptual importance of different sound sources. Spatial resolution and encoding detail are adjusted in real-time according to the acoustic scene, maintaining high spatial accuracy where needed while reducing precision for less critical elements to achieve efficient compression.
3Productivity
If spatial information is added after compression, then compression efficiency is improved, but audible degradation occurs due to binaural unmasking effects
Solution Approach 1:
The patent performs preliminary spatial encoding during the compression phase rather than adding spatial information afterward. By calculating and embedding spatial descriptors (ICTD, ICLD, ICC) during the encoding process itself, it ensures spatial information is properly integrated with the audio signal, avoiding the binaural unmasking artifacts that occur when spatial data is added post-compression.
Data Source
AI summary
A method for compressing an audio stream, including a plurality of signals, describing a sound scene produced by a plurality of sources in a space, by: identifying the sources from an audio stream; determining a frequency band, energy level and spatial position in the space for each of the identified sources; determining, for each identified source, a spatial resolution corresponding to the smallest difference in position of said source in the space which a listener is capable of perceiving, on the basis of: the frequency band, the energy level, and the spatial position of said source; and, on the frequency band, energy level, and spatial position of at least one subset of the other identified sources; generating a compressed stream comprising the information required to restore each identified source with at least the same corresponding spatial resolution.


