Ambisonics Audio Coding with Hybrid SPAR-DirAC Reconstruction
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing Ambisonics audio coding technologies face inefficiencies in encoding and decoding, particularly at high bit rates, with Spatial Audio Reconstruction (SPAR) lacking the ability to recover higher Ambisonics orders from lower orders and Directional Audio Coding (DirAC) having inferior coding efficiency and complexity issues.
Innovation Solution
A combined coding scheme that integrates SPAR and DirAC, where SPAR downmixes and parametrically encodes Ambisonics signals, and DirAC enhances the spatial resolution by analyzing and synthesizing intermediate signals, allowing flexible and efficient encoding and decoding across various bit rates.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If SPAR encoder is used for Ambisonics audio coding, then coding efficiency is improved, but ability to recover higher Ambisonics orders from lower orders is lost
Solution Approach 1:
The patent merges SPAR and DirAC coding schemes into a hybrid system where SPAR processes the Ambisonics signal and DirAC enhances spatial resolution. The combination allows the system to achieve both efficient coding and the ability to recover higher Ambisonics orders, as DirAC's directional parameters supplement SPAR's reconstruction capabilities.
Solution Approach 2:
The patent creates a composite coding structure that integrates parameters from both SPAR (spatial reconstruction parameters) and DirAC (directional audio coding parameters). This composite approach combines the strengths of both methods, enabling efficient coding while preserving the ability to reconstruct higher-order spatial information.
2Measurement precision
If DirAC is used for Ambisonics audio coding, then spatial resolution is enhanced, but coding efficiency deteriorates
Solution Approach 1:
The patent applies partial DirAC processing by using it only for specific frequency ranges or time segments where spatial resolution enhancement is most beneficial, while relying on SPAR for other portions. This selective application maintains spatial resolution advantages while reducing the overall computational burden and improving coding efficiency.
3Measurement precision
If DirAC is used for Ambisonics audio coding, then spatial analysis capability is improved, but system complexity increases
Solution Approach 1:
The patent segments the audio signal processing into distinct frequency bands or time segments, applying DirAC analysis only where it provides the most value. This segmentation reduces the overall computational complexity while maintaining enhanced spatial analysis capabilities in critical regions of the audio spectrum.
Data Source
Figure 1~2
Figure 3a
Figure 3b
AI summary
The application provides a method (500) for encoding an Ambisonics input audio signal. The method (500) comprises providing (501) the input audio signal to a SPAR encoder and to a DirAC analyzer and parameter encoder. Furthermore, the method (500) comprises generating (502) an encoder bit stream based on output of the SPAR encoder and based on output of the DirAC analyzer and parameter encoder. The application also provides a method for decoding the encoder bitstream by generating an intermediate Ambisonics signal using a SPAR decoder based on the encoder bitstream and processing the intermediate Ambisonics signal using a DirAC synthesizer to provide an output audio signal for rendering. The DirAC synthesizer may use DirAC parameters of the bitstreams or DirAC parameters obtained by analysing the intermediate Ambisonics signal.