DirAC Spatial Audio Coding Format Conversion and Mixing
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current audio signal processing technologies lack a universal scheme to efficiently encode, transmit, and reproduce complex audio scenes composed of different 3D audio representations, such as channel-based, object-based, and scene-based formats, particularly in supporting audio objects and Ambisonics formats.
Innovation Solution
A DirAC-based spatial audio coding system that converts and combines different audio formats into a common format, enabling efficient parametric coding and manipulation of audio objects, allowing for dialogue enhancement and flexible handling of audio scenes across various loudspeaker layouts and headphones.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If multiple dedicated coding schemes are used for different audio formats (channel-based, object-based, scene-based), then each format can be efficiently coded, but the system complexity increases and a universal scheme is lacking
Solution Approach 1:
The patent applies universality by designing a single parametric coding scheme that can handle multiple audio representations (channel-based, object-based, and scene-based formats) through a unified framework. The DirAC technique serves as a universal representation that can encode different audio formats using the same coding principles, eliminating the need for multiple dedicated coding schemes while maintaining efficiency across all formats.
Solution Approach 2:
The patent utilizes parameter changes by representing different audio formats through variations of parametric descriptors within the DirAC framework. By adjusting parameters such as direction of arrival, diffuseness, and spatial distribution within the unified coding scheme, the system can adapt to encode channel-based, object-based, and scene-based audio representations without requiring separate coding mechanisms for each format.
2Adaptability or versatility
If DirAC is used as a common format for mixing different audio formats, then format compatibility improves, but the ability to directly process audio objects with metadata is limited
Solution Approach 1:
The patent applies segmentation by separating the audio scene into distinct audio objects, each with its own DirAC parameters and metadata. This allows individual audio objects to be processed, manipulated, and encoded independently while maintaining their identity within the mixed audio scene. The segmentation enables direct processing of audio objects with their associated metadata while still using DirAC as the common representation format.
3Productivity
If a universal DirAC-based scheme is implemented, then encoding efficiency across formats improves, but the complexity of converting and combining different formats increases
Solution Approach 1:
The patent applies the intermediary principle by using DirAC parameters as a mediator between different audio formats. The conversion process involves transforming various audio representations into DirAC parametric form, which serves as an intermediate representation that simplifies the combination process. This intermediary step enables efficient encoding by providing a common language for mixing different audio formats while managing conversion complexity through standardized parametric transformations.
Data Source
AI summary
An apparatus for generating a description of a combined audio scene, includes: an input interface for receiving a first description of a first scene in a first format and a second description of a second scene in a second format, wherein the second format is different from the first format; a format converter for converting the first description into a common format and for converting the second description into the common format, when the second format is different from the common format; and a format combiner for combining the first description in the common format and the second description in the common format to obtain the combined audio scene.


