Higher Order Ambisonic Audio Data Compression via Vector Decomposition
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Higher order ambisonic (HOA) audio signals, represented by spherical harmonic coefficients, face challenges in efficient compression and transmission due to the large number of channels required, which exceeds the limitations of existing audio equipment and devices, such as mobile devices and broadcast workflows.
Innovation Solution
A vector-based HOA format is introduced, where higher order ambisonic audio data is decomposed into sound and spatial components, with priority information determined based on energy, loudness, and spatial weighting, allowing for the selection of a non-zero subset of components for compression, and specifying these in a unified data object format to reduce channel requirements.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If higher order ambisonic audio data is represented using spherical harmonic coefficients, then the soundfield representation accuracy is improved, but the number of channels increases beyond the limitations of existing audio equipment
Solution Approach 1:
The patent extracts only the most significant spherical harmonic coefficients (those with highest energy contribution) from the complete HOA representation. By selecting a subset of coefficients that capture the dominant spatial information, the system reduces the channel count to match existing equipment capabilities while preserving the essential soundfield characteristics.
Solution Approach 2:
The patent transforms the HOA coefficients from the spherical harmonic domain to the spatial domain using a transformation matrix, and then applies energy-based thresholding to select significant components. This parameter transformation and selective retention changes the representation from a complete high-order model to a truncated version that fits equipment constraints.
2Reliability
If all spherical harmonic coefficients are transmitted, then the audio quality is improved, but the data transmission and storage requirements increase significantly
Solution Approach 1:
The patent extracts and transmits only the significant HOA coefficients that contribute most to audio quality, as determined by energy analysis. This selective extraction reduces the data volume for transmission and storage while maintaining the perceptual quality by preserving the most important spatial audio information.
Solution Approach 2:
Instead of transmitting the complete set of HOA coefficients, the patent transmits a partial set that is sufficient to achieve the desired audio quality. The energy-based selection ensures that the transmitted subset contains enough information for high-quality reproduction without the redundancy of less significant coefficients.
3Adaptability or versatility
If a truncated HOA representation is used to reduce channel count, then compatibility with existing equipment is improved, but the soundfield representation accuracy deteriorates
Solution Approach 1:
The patent applies a transformation from spherical harmonic coefficients to spatial domain coefficients using a predetermined transformation matrix. This parameter change enables the representation to be adapted to different speaker configurations and existing equipment while maintaining soundfield accuracy through energy-based selection of significant components.
Solution Approach 2:
The patent performs energy analysis and coefficient selection in advance, before the actual audio reproduction. By pre-identifying and retaining only the significant coefficients, the system ensures compatibility with existing equipment while minimizing the loss of soundfield representation accuracy.
Data Source
AI summary
In general, techniques are described by which to provide priority information for higher order ambisonic (HOA) audio data. A device comprising a memory and a processor may perform the techniques. The memory stores HOA coefficients of the HOA audio data, the HOA coefficients representative of a soundfield. The processor may decompose the HOA coefficients into a sound component and a corresponding spatial component, the corresponding spatial component defining shape, width, and directions of the sound component, and the corresponding spatial component defined in a spherical harmonic domain. The processor may also determine, based on one or more of the sound component and the corresponding spatial component, priority information indicative of a priority of the sound component relative to other sound components of the soundfield, and specify, in a data object representative of a compressed version of the HOA audio data, the sound component and the priority information.


