DirAC Spatial Audio Coding With Selective Diffuse Compensation
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing spatial audio coding methods face challenges in efficiently synthesizing higher-order Ambisonics signals while maintaining audio quality, particularly in low-bitrate transmission scenarios, due to quantization errors and energy fluctuations in directional components.
Innovation Solution
The method involves synthesizing diffuse sound components only up to a certain order and performing energy compensation on low-order and mid-order sound field components, using diffuseness and direction data to correct for energy loss and artifacts.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If diffuse sound components are synthesized up to higher orders, then spatial audio quality is improved, but computational complexity increases
Solution Approach 1:
The patent segments the sound field components into different groups based on their diffuseness characteristics. Low-order components are processed with full diffuse sound synthesis, while high-order components use simplified processing. This segmentation allows the system to achieve good spatial audio quality by focusing computational resources where they are most needed, rather than uniformly processing all components at high cost.
Solution Approach 2:
The patent applies different processing qualities to different parts of the sound field representation. Specifically, low-order Ambisonics components (which have greater impact on perceived spatial quality) receive full diffuse sound synthesis processing, while high-order components use energy compensation only. This local quality approach optimizes the balance between overall audio quality and computational complexity.
2Measurement precision
If energy compensation is applied to all sound field components, then audio quality is improved, but processing time increases
Solution Approach 1:
The patent applies energy compensation selectively rather than universally. Instead of applying energy compensation to all sound field components equally, the method applies full energy compensation only to low-order components where it has the greatest impact on audio quality. For high-order components, simplified or no energy compensation is applied, as these components contribute less to the overall perceived quality. This partial action approach maintains audio quality while significantly reducing processing time.
3Productivity
If quantization is applied during low-bitrate transmission, then transmission efficiency is improved, but quantization errors increase
Solution Approach 1:
The patent changes the representation parameters of the sound field to make them more robust to quantization. By using Ambisonics coefficients as the transmission format instead of raw multi-channel audio, the system achieves better compression efficiency. Additionally, the selective energy compensation approach applied after decoding helps correct quantization errors in the most perceptually important low-order components, maintaining signal accuracy despite low-bitrate transmission.
Data Source
AI summary
An apparatus generating a sound field description includes an input signal analyzer acquiring diffuseness data from the input signal; a sound component generator for generating, from the input signal, one or more sound field components of a first group of sound field components comprising for each sound field component a direct component and a diffuse component, and for generating, from the input signal, a second group of sound field components comprising only a direct component. The sound component generator is configured to perform an energy compensation when generating the first group of sound field components, the energy compensation depending on the diffuseness data and at least one of a number of sound field components in the second group, a number of diffuse components in the first group, a maximum order of sound field components of the first group and a maximum order of sound field components of the second group.


