Spatial Audio Stem Encoding for Adaptive Diffusion Playback

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current digital audio reproduction systems fail to provide an optimal acoustic experience for home theater listeners due to the lack of artistic control over reverberation and diffusion, as they are typically mixed for large cinema environments and lack flexibility to adapt to smaller spaces, leading to a less than ideal sound experience even with high-quality surround sound systems.

Innovation Solution

The system encodes 'dry' audio tracks with time-variable metadata representing desired diffusion and mix parameters, allowing for customization of playback based on the local environment, using a configurable audio diffusion processor that dynamically adjusts reverberation and delay to create a perceptually diffuse sound effect.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If audio is mixed for large cinema environments, then the soundtrack is optimized for theater presentation with proper dialog and spatial imaging, but the same mix performs poorly in small home listening environments due to excessive reverberation and inappropriate spatial characteristics

Engineering Contradiction:
Improvesoundtrack performanceVSAvoidadaptability to different playback spaces
Core Design Contradiction:
ReliabilityVSAdaptability or versatility

Solution Approach 1:

The audio signal is segmented into direct sound components and reverberant sound components separately. The direct sound is preserved for spatial localization while the reverberant sound is processed independently to match the acoustic characteristics of different playback environments (cinema vs. home theater).

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The system dynamically adjusts the amount and characteristics of reverberation added to the audio signal based on the detected or selected playback environment. Metadata parameters control the reverberation time, decay rate, and spatial diffusion to adapt the mix for either large cinema spaces or small home environments.

Inventive Principle:
Principle #15Dynamics

2Reliability

If different mixes are released for home and cinema listening, then each environment receives optimized audio, but this approach is economically unviable and impossible for legacy content where original multi-track stems are unavailable

Engineering Contradiction:
Improveacoustic experience qualityVSAvoidproduction complexity and cost
Core Design Contradiction:
ReliabilityVSEase of manufacture

Solution Approach 1:

Reverberation and spatial effects are pre-calculated and encoded as metadata parameters during the mastering stage. This preliminary action allows the same audio stem to be adaptively rendered for different environments at playback without requiring separate mixes or access to original multi-track sessions.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

Metadata acts as an intermediary between the audio signal and the playback environment characteristics. The metadata contains pre-computed reverberation parameters that mediate the interaction between the fixed audio mix and variable playback spaces, enabling adaptive rendering without re-mixing.

Inventive Principle:
Principle #24Intermediary (Mediator)

3Ease of operation

If reverberation is introduced in the soundtrack, then spatial ambiance is enhanced, but the reverberation characteristics become mismatched with the actual playback space, especially in small rooms with short reverberation times

Engineering Contradiction:
Improvespatial ambianceVSAvoidreverberation characteristic accuracy
Core Design Contradiction:
Ease of operationVSManufacturing precision

Solution Approach 1:

The system changes key reverberation parameters (reverberation time, decay rate, spatial diffusion coefficient) based on the target playback environment. For small home theaters with short natural reverberation, the metadata specifies shorter decay times and lower diffusion levels, while cinema mixes use longer reverberation parameters.

Inventive Principle:
Principle #35Parameter changes

4Reliability

If diffuse sound is synthesized using time-to-frequency transform techniques and reverberation processing, then simulated reverberant signals are generated, but the computational demand is excessively high for practical implementations

Engineering Contradiction:
Improvediffuse sound qualityVSAvoidprocessing efficiency
Core Design Contradiction:
ReliabilityVSProductivity

Solution Approach 1:

The computationally intensive reverberation processing is extracted from the real-time playback path and performed in advance during the encoding stage. The results are stored as compact metadata parameters that can be applied efficiently at playback without requiring heavy computational resources.

Inventive Principle:
Principle #2Taking out (Extraction)

Applied Scientific Principles

This section explains which scientific principles are used to turn an abstract innovation direction into a practical engineering solution.

Function Achieved in This Case

This approach enables a single mix to play well in both cinema and home environments, providing improved acoustic fidelity by allowing for user-selectable diffusion effects, enhancing the psychoacoustic experience by simulating accurate reverberation and diffusion, thus addressing the limitations of existing technologies.

Implementation Method 1

The metadata includes at least one parameter capable of being decoded to configure a perceptually diffuse audio effect in at least one audio channel. The system processes the digital audio signal with the perceptually diffuse audio effect to simulate accurate reverberation and diffusion.

Methodology Applied
Scientific EffectReverberation: Reverberation

Data Source

PatentUS8908874B2Spatial audio encoding and reproduction
Publication Date: 2014.12.09 DTS INC(US)
  • US8908874B2 patent drawing
  • US8908874B2 patent drawing
  • US8908874B2 patent drawing

AI summary

A method and apparatus processes multi-channel audio by encoding, transmitting or recording “dry” audio tracks or “stems” in synchronous relationship with time-variable metadata controlled by a content producer and representing a desired degree and quality of diffusion. Audio tracks are compressed and transmitted in connection with synchronized metadata representing diffusion and preferably also mix and delay parameters. The separation of audio stems from diffusion metadata facilitates the customization of playback at the receiver, taking into account the characteristics of local playback environment.