Spatial Audio Metadata Encoding Using DCT Codebooks
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current audio encoding technologies face challenges in effectively encoding spatial audio metadata, particularly for inputs other than microphone-array captured signals, such as loudspeaker signals and Ambisonic signals, with existing methods focusing on direction and energy ratio parameters but lacking efficient compression of coherence data.
Innovation Solution
The proposed solution involves an apparatus and method that determine a codebook for encoding spread and surround coherence values based on energy ratio and azimuth values, using discrete cosine transformation and entropy encoding to efficiently compress and transmit spatial audio metadata, including coherence parameters, across various audio signal formats.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If coherence parameters are encoded using traditional methods, then encoding simplicity is maintained, but encoding efficiency and compression ratio deteriorate
Solution Approach 1:
The patent applies parameter changes by transforming coherence parameters from the time domain to the frequency domain using Discrete Cosine Transform (DCT). This transformation changes the representation of the data, concentrating energy in fewer coefficients, which enables more efficient entropy coding and achieves better compression ratios without significantly increasing processing complexity.
Solution Approach 2:
The patent segments the coherence parameters into different frequency sub-bands using DCT, allowing independent processing and coding of each sub-band. This segmentation enables selective coding strategies where important low-frequency components are preserved while less important high-frequency components are compressed more aggressively, improving overall encoding efficiency.
2Quantity of substance
If spatial audio metadata is compressed to reduce bitrate, then transmission efficiency is improved, but reconstruction quality deteriorates
Solution Approach 1:
The patent uses DCT to transform coherence parameters into a frequency domain representation where energy is concentrated in fewer coefficients. This parameter change enables selective retention of important coefficients during compression, allowing high-quality reconstruction even at reduced bitrates by preserving the most significant spectral components.
Solution Approach 2:
The patent replaces traditional time-domain compression methods with frequency-domain transformation and entropy coding. This substitution allows for more efficient information packing by exploiting the statistical properties of DCT coefficients, achieving better compression-to-quality ratio through mathematical transformation rather than simple quantization.
3Measurement precision
If coherence parameters are encoded with high precision, then spatial audio fidelity is improved, but data size increases
Solution Approach 1:
The patent transforms coherence parameters using DCT, which concentrates the information into fewer, more significant coefficients. This parameter transformation enables high-fidelity representation with reduced data size by retaining only the most important spectral components rather than preserving all time-domain samples at full precision.
Solution Approach 2:
The patent moves the representation of coherence parameters from the time domain to the frequency domain through DCT. This dimensional transformation reveals the underlying spectral structure of the data, allowing for more efficient compression while maintaining fidelity by preserving the essential spectral characteristics rather than all temporal details.
Data Source
AI summary
An apparatus comprising means for: receiving values for sub-bands of a frame of an audio signal, the values comprising at least one azimuth value, at least one elevation value at least one energy ratio value and at least one spread and/or surround coherence value for each sub-band; determining a codebook for encoding at least one spread and/or surround coherence value for each sub-band based on the at least one energy ratio value and at least one azimuth value for each sub-band for a frame; discrete cosine transforming at least one vector, the at least one vector comprising the at least one spread and/or surround coherence value for a sub-band for the frame; and encoding a first number of components of the discrete cosine transformed vector based on the determined codebook.


