Directional Audio Parameter Encoding With Mixed Time-Frequency Resolution
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current methods for encoding directional audio coding parameters, such as DirAC metadata, face challenges in achieving low bit-rate transmission while maintaining high quality, particularly due to the large amount of data required for 3D audio scenes, and previous solutions have compromised on spatial resolution or limited applications to specific scenarios like teleconferences.
Innovation Solution
The proposed solution involves encoding directional audio coding parameters with different time or frequency resolutions for diffuseness and direction parameters, using a parameter calculator to calculate these parameters at varying resolutions and a quantizer to reduce data by quantizing and encoding them efficiently, while also considering the power of the audio signal for weighted averaging and grouping.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If directional audio coding parameters are encoded with high spatial resolution, then audio quality is improved, but data rate increases
Solution Approach 1:
The audio spectrum is segmented into multiple frequency bands, and within each band, diffuseness and direction parameters are calculated separately. This allows selective encoding where high resolution is applied only where needed (low diffuseness regions) while high diffuseness regions use coarser resolution, resolving the contradiction between overall spatial resolution and total data rate
Solution Approach 2:
Different time/frequency resolutions are applied locally based on the diffuseness characteristic of each region. Low diffuseness regions (directional sound) receive high resolution encoding, while high diffuseness regions (diffuse sound) receive low resolution encoding. This local adaptation maintains audio quality where needed while reducing data rate in regions where high resolution provides minimal perceptual benefit
2Measurement precision
If uniform high resolution is used for all parameters, then audio quality is maintained, but bit-rate increases
Solution Approach 1:
The resolution parameter is dynamically changed based on the diffuseness value of each time/frequency bin. The encoder selects from multiple predefined resolution levels (e.g., 1/4, 1/2, full resolution) according to the local diffuseness characteristic, optimizing the trade-off between quality and bit-rate for each region independently
Solution Approach 2:
Instead of applying full high resolution encoding to all parameters uniformly, the system applies high resolution only partially to regions where it provides perceptual benefit (low diffuseness areas). This partial application of high resolution significantly reduces total bit-rate while maintaining quality where it matters most
3Quantity of substance
If different resolutions are used for diffuseness and direction parameters, then data rate is reduced, but encoding complexity increases
Solution Approach 1:
The diffuseness parameter is calculated first for each time/frequency bin, and its value is used to pre-determine the resolution level to be applied to the direction parameters. This preliminary calculation of diffuseness guides the subsequent encoding process, simplifying the overall complexity by providing a clear decision rule before direction parameter encoding
Data Source
AI summary
An apparatus for encoding directional audio coding parameters including diffuseness parameters and direction parameters includes: a parameter calculator for calculating the diffuseness parameters with a first time or frequency resolution and for calculating the direction parameters with a second time or frequency resolution; and a quantizer and encoder processor for generating a quantized and encoded representation of the diffuseness parameters and the direction parameters.


