Ambient Higher Order Ambisonic Audio Normalization
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current audio encoding technologies face challenges in efficiently compressing and transmitting higher order ambisonic audio data, particularly due to limitations in legacy audio equipment that struggle with the increased number of audio channels required for 3D soundfield representation, leading to issues with file sizes, latency, and compatibility with older systems.
Innovation Solution
The implementation of a method for normalization and inverse normalization of ambient higher order ambisonic audio coefficients, using a device with processors and memory to store and process these coefficients, which involves a spatial audio encoding device that performs decomposition, energy compensation, and quantization to reduce the number of channels and improve coding efficiency, and a decoding device that reverses these processes to maintain audio quality.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If higher order ambisonic audio data is used to represent 3D soundfield, then the soundfield representation quality is improved, but the number of audio channels increases leading to larger file sizes and compatibility issues with legacy equipment
Solution Approach 1:
The patent segments the higher order ambisonic audio data into different spherical harmonic coefficient groups that can be selectively processed. By dividing the complete HOA data into manageable segments corresponding to different spatial frequencies and directions, the system can compress the data while preserving essential soundfield characteristics, thus reducing the effective number of channels needed for representation.
Solution Approach 2:
The patent extracts and processes specific spherical harmonic coefficients that contribute most significantly to the ambient soundfield representation. By identifying and separating the most important coefficients from the complete HOA dataset, the system can maintain soundfield quality while reducing the total channel count required for transmission and storage.
2Measurement precision
If higher order ambisonic audio data is used, then the 3D soundfield representation is improved, but the processing complexity and latency increase for legacy audio equipment
Solution Approach 1:
The patent performs preliminary normalization of spherical harmonic coefficients before encoding and transmission. By pre-processing the HOA data to establish proper energy scaling and normalization factors, the system reduces the computational burden on legacy equipment during playback, as the heavy normalization calculations are already completed during encoding.
Solution Approach 2:
The patent transforms the HOA data by changing parameters such as spherical harmonic coefficient scaling and energy normalization. These parameter transformations are designed to make the data more compatible with legacy audio equipment while preserving the 3D soundfield information, effectively adapting the high-dimensional data to lower-dimensional processing requirements.
3Measurement precision
If higher order ambisonic audio data is used, then the ambient soundfield representation is improved, but the energy compensation and normalization requirements increase
Solution Approach 1:
The patent implements energy compensation mechanisms that use feedback from the normalized spherical harmonic coefficients. By calculating the energy distribution across different HOA channels and using this information to adjust normalization factors, the system maintains accurate ambient soundfield representation while optimizing energy usage and reducing excessive compensation requirements.
Data Source
AI summary
In general, techniques are directed to performing normalization with respect to ambient higher order ambisonic audio data. A device configured to decode higher order ambisonic audio data may perform the techniques. The device may include a memory and one or more processors. The memory may be configured to store an audio channel that provides a normalized ambient higher order ambisonic coefficient representative of at least a portion of an ambient component of a soundfield. The one or more processors may be configured to perform inverse normalization with respect to the audio channel.


