Audio Object Extraction Using Projection-Space Channel Correlation

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Channel-based audio content formats are inefficient in adapting to various playback configurations, leading to degraded listening experiences due to mismatched playback settings and limitations in binaural rendering, particularly in separating and positioning audio objects.

Innovation Solution

A method and system for extracting audio objects from channel-based audio content by identifying projection spaces and determining correlations between channels to separate and re-render objects, using source separation techniques to produce clean multi-channel or mono representations adaptable to different playback settings.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Ease of manufacture

If channel-based audio content formats are used, then the audio content can be created and stored with predefined physical locations optimized for specific playback settings, but the listening experience degrades when played back with different playback configurations due to mismatch between settings

Engineering Contradiction:
Improveaudio content creationVSAvoidplayback configuration adaptability
Core Design Contradiction:
Ease of manufactureVSAdaptability or versatility

Solution Approach 1:

The patent segments the audio content into individual audio objects with independent spatial position metadata, separating the audio signal from its spatial representation. This allows each object to be independently positioned and rendered, enabling adaptation to different playback configurations without degrading the listening experience.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent implements dynamic rendering where audio objects are re-positioned and re-rendered in real-time based on the detected playback configuration. The system dynamically adjusts the spatial parameters of audio objects to match the actual speaker arrangement, transforming static channel-based content into dynamically adaptable object-based content.

Inventive Principle:
Principle #15Dynamics

2Device complexity

If channel-based format is used for binaural rendering, then a limited number of head-related transfer functions (HRTFs) specific to speaker positions can be used, but for other positions interpolation of HRTFs is required which degrades the binaural listening experience

Engineering Contradiction:
ImproveHRTF implementationVSAvoidbinaural listening experience
Core Design Contradiction:
Device complexityVSReliability

Solution Approach 1:

The patent changes the spatial representation parameter from fixed channel assignments to continuous 3D position coordinates. By representing audio objects with precise spatial parameters (azimuth, elevation, distance), the system can select or interpolate HRTFs based on actual object positions rather than being constrained to predefined speaker positions, significantly improving binaural rendering quality.

Inventive Principle:
Principle #35Parameter changes

3Device complexity

If source separation techniques are not used, then the processing is simpler, but objects mixed in the same channel cannot be separated leading to incorrect position estimations

Engineering Contradiction:
Improveprocessing complexityVSAvoidobject position estimation
Core Design Contradiction:
Device complexityVSMeasurement precision

Solution Approach 1:

The patent extracts individual audio objects from mixed channel signals using source separation techniques. By separating the mixed signals into distinct object components, each with its own spatial position metadata, the system achieves accurate position estimation for each object even when multiple objects were originally mixed in the same channel.

Inventive Principle:
Principle #2Taking out (Extraction)

Data Source

PatentUS10275685B2Projection-based audio object extraction from audio content
Publication Date: 2019.04.30 DOLBY LABORATORIES LICENSING CORP
  • US10275685B2 patent drawing
  • US10275685B2 patent drawing

AI summary

A method is disclosed for audio object extraction from an audio content which includes identifying a first set of projection spaces including a first subset for a first channel and a second subset for a second channel of the plurality of channels. The method may further include determining a first set of correlations between the first and second channels, each of the first set of correlations corresponding to one of the first subset of projection spaces and one of the second subset of projection spaces. Still further, the method may include extracting an audio object from an audio signal of the first channel at least in part based on a first correlation among the first set of correlations and the projection space from the first subset corresponding to the first correlation, the first correlation being greater than a first predefined threshold. Corresponding system and computer program products are also disclosed.