Scalable Higher Order Ambisonic Audio Coding Layers
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current technologies for encoding and decoding higher-order ambisonic audio data lack scalability and flexibility, particularly in accommodating varying speaker geometries and acoustic conditions, which limits their ability to provide high-quality soundfield representation across different playback setups.
Innovation Solution
The proposed solution involves scalable coding techniques that divide higher-order ambisonic audio data into multiple layers, including a base layer and enhancement layers, allowing for adaptable reproduction of soundfields. This approach enables improved resolution and accuracy in soundfield representation by combining layers, accommodating various speaker configurations and playback conditions.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If higher-order ambisonic audio data is encoded using a single layer, then the encoding process is simple, but the soundfield reproduction accuracy and resolution are insufficient
Solution Approach 1:
The patent divides the higher-order ambisonic audio data into multiple encoding layers (base layer and enhancement layers), where each layer processes specific spherical harmonic coefficients. This segmentation allows the system to achieve high soundfield reproduction accuracy by combining multiple layers while maintaining manageable encoding complexity through modular processing of different coefficient groups.
2Adaptability or versatility
If a single encoding format is used, then the encoding process is straightforward, but the adaptability to different speaker geometries and playback setups is limited
Solution Approach 1:
The patent creates a universal multi-layer coding structure that can adapt to various speaker geometries and playback configurations. The base layer provides fundamental soundfield representation compatible with simple setups, while enhancement layers add resolution for complex configurations. This multi-functional design allows the same encoding system to serve diverse playback scenarios from stereo to 3D surround sound.
3Measurement precision
If all spherical harmonic coefficients are encoded with high precision, then the soundfield representation is accurate, but the data transmission and processing load increases
Solution Approach 1:
The patent segments spherical harmonic coefficients into different groups that are encoded at different precision levels across multiple layers. The base layer encodes essential coefficients at lower precision for efficient transmission, while enhancement layers encode additional coefficients at higher precision to achieve accurate soundfield representation. This segmented approach balances data volume with representation accuracy.
4Measurement precision
If enhancement layers are added to improve soundfield resolution, then the reproduction quality increases, but the decoding complexity increases
Solution Approach 1:
The patent performs preliminary organization of spherical harmonic coefficients into structured groups before encoding into multiple layers. This preliminary action enables the decoding process to efficiently combine base layer and enhancement layer data by following a predetermined coefficient mapping structure, thereby reducing decoding complexity while maintaining high soundfield reproduction resolution.
Data Source
Figure 1
Figure 2
Figure 3
AI summary
In general, techniques are described for signaling channels for scalable coding of higher order ambisonic audio data. A device comprising a memory and a processor may be configured to perform the techniques. The memory may be configured to store the bitstream. The processor may be configured to obtain, from the bitstream, an indication of a number of channels specified in one or more layers in the bitstream, and obtain the channels specified in the one or more layers in the bitstream based on the indication of the number of channels.