DirAC Spatial Audio Coding Format Conversion and Mixing

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current audio signal processing technologies lack a universal scheme to efficiently encode, transmit, and reproduce complex audio scenes composed of different 3D audio representations, such as channel-based, object-based, and scene-based formats, particularly in supporting audio objects and Ambisonics formats.

Innovation Solution

A DirAC-based spatial audio coding system that converts and combines different audio formats into a common format, enabling efficient parametric coding and manipulation of audio objects, allowing for dialogue enhancement and flexible handling of audio scenes across various loudspeaker layouts and headphones.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If multiple dedicated coding schemes are used for different audio formats (channel-based, object-based, scene-based), then each format can be efficiently coded, but the system complexity increases and a universal scheme is lacking

Engineering Contradiction:
Improvesupport for multiple audio formatsVSAvoidnumber of coding schemes
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The patent applies universality by designing a single parametric coding scheme that can handle multiple audio representations (channel-based, object-based, and scene-based formats) through a unified framework. The DirAC technique serves as a universal representation that can encode different audio formats using the same coding principles, eliminating the need for multiple dedicated coding schemes while maintaining efficiency across all formats.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The patent utilizes parameter changes by representing different audio formats through variations of parametric descriptors within the DirAC framework. By adjusting parameters such as direction of arrival, diffuseness, and spatial distribution within the unified coding scheme, the system can adapt to encode channel-based, object-based, and scene-based audio representations without requiring separate coding mechanisms for each format.

Inventive Principle:
Principle #35Parameter changes

2Adaptability or versatility

If DirAC is used as a common format for mixing different audio formats, then format compatibility improves, but the ability to directly process audio objects with metadata is limited

Engineering Contradiction:
Improveformat compatibilityVSAvoiddirect processing of audio objects
Core Design Contradiction:
Adaptability or versatilityVSEase of operation

Solution Approach 1:

The patent applies segmentation by separating the audio scene into distinct audio objects, each with its own DirAC parameters and metadata. This allows individual audio objects to be processed, manipulated, and encoded independently while maintaining their identity within the mixed audio scene. The segmentation enables direct processing of audio objects with their associated metadata while still using DirAC as the common representation format.

Inventive Principle:
Principle #1Segmentation

3Productivity

If a universal DirAC-based scheme is implemented, then encoding efficiency across formats improves, but the complexity of converting and combining different formats increases

Engineering Contradiction:
Improveencoding efficiencyVSAvoidformat conversion and combination process
Core Design Contradiction:
ProductivityVSDevice complexity

Solution Approach 1:

The patent applies the intermediary principle by using DirAC parameters as a mediator between different audio formats. The conversion process involves transforming various audio representations into DirAC parametric form, which serves as an intermediate representation that simplifies the combination process. This intermediary step enables efficient encoding by providing a common language for mixing different audio formats while managing conversion complexity through standardized parametric transformations.

Inventive Principle:
Principle #24Intermediary (Mediator)

Data Source

PatentUS12058501B2Apparatus, method and computer program for encoding, decoding, scene processing and other procedures related to DirAC based spatial audio coding
Publication Date: 2024.08.06 FRAUNHOFER GESELLSCHAFT ZUR FORDERUNG DER ANGEWANDTEN FORSCHUNG EV
  • US12058501B2 patent drawing
  • US12058501B2 patent drawing
  • US12058501B2 patent drawing

AI summary

An apparatus for generating a description of a combined audio scene, includes: an input interface for receiving a first description of a first scene in a first format and a second description of a second scene in a second format, wherein the second format is different from the first format; a format converter for converting the first description into a common format and for converting the second description into the common format, when the second format is different from the common format; and a format combiner for combining the first description in the common format and the second description in the common format to obtain the combined audio scene.