Spatial Audio Quantization Scheme Selection for Direction Precision
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing spatial audio encoding technologies face challenges in efficiently encoding directional components like elevation and azimuth with minimal bits, especially for multi-channel audio inputs, and require high bitrate and high quality lossless encoding.
Innovation Solution
A system combining fixed and variable bitrate coding with a quantization scheme that distributes bits based on variance, using a determining procedure to select quantization schemes for azimuth and elevation values, and employs a spherical grid quantization for direction parameters.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Device complexity
If fixed bitrate coding is used for spatial metadata encoding, then device complexity is reduced, but measurement precision of directional parameters deteriorates
Solution Approach 1:
The patent implements dynamic bit allocation that adapts to the variance characteristics of directional parameters. The system calculates variance for each time-frequency block and dynamically adjusts the number of bits allocated to encode azimuth and elevation values, transitioning from static fixed-bitrate to dynamic variable-bitrate encoding based on actual signal characteristics
Solution Approach 2:
The patent changes the encoding parameters by using different quantization schemes (uniform vs. non-uniform) based on the variance of directional parameters. When variance is high, non-uniform quantization with more bits is applied; when variance is low, uniform quantization with fewer bits is used, thereby optimizing the trade-off between precision and bitrate
2Measurement precision
If variable bitrate coding is used for spatial metadata encoding, then measurement precision of directional parameters is improved, but device complexity increases
Solution Approach 1:
The patent segments the encoding process into distinct stages: variance calculation, threshold comparison, and conditional quantization selection. By dividing the directional parameters into different time-frequency blocks and applying different quantization strategies to each segment based on its characteristics, the system manages complexity through structured segmentation rather than monolithic processing
3Manufacturing precision
If high bitrate encoding is used for spatial metadata, then manufacturing precision of audio quality is improved, but loss of energy increases
Solution Approach 1:
The patent changes the encoding parameters dynamically by adjusting the number of bits allocated to directional parameters based on their variance. This allows the system to maintain high audio quality when necessary (high variance regions) while reducing bitrate when quality requirements are naturally met (low variance regions), optimizing the energy-quality trade-off
4Device complexity
If uniform quantization is used for directional parameters, then device complexity is reduced, but measurement precision deteriorates
Solution Approach 1:
The patent applies local quality by using different quantization schemes (uniform or non-uniform) for different time-frequency blocks based on their local variance characteristics. Instead of applying a single uniform quantization strategy globally, the system adapts the quantization method to the local properties of each parameter segment, thereby optimizing precision where needed while maintaining simplicity where sufficient
Data Source
Figure 1
Figure 2
Figure 3a
AI summary
There is disclosed inter alia an apparatus for spatial audio signal encoding comprising means for receiving for each time frequency block of a sub band of an audio frame a spatial audio parameter comprising an azimuth and an elevation; determining a first distortion measure for the audio frame by determining a first distance measure for each time frequency block and summing the first distance measure for each time frequency block; determining a second distortion measure for the audio frame by determining a second distance measure for each time frequency block and summing the second distance measure for each time frequency block, and selecting either the first quantization scheme or the second quantization scheme for quantising the elevation and the azimuth for all time frequency blocks of the sub band of the audio frame, wherein the selecting is dependent on the first and second distortion measures.