Binaural Audio Spatial Enhancement Using Object-Based Masks

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Rendering user-generated content with immersive audio is difficult due to challenges in enhancing spatial characteristics of audio signals.

Innovation Solution

A method involving the extraction of audio objects from multi-channel audio signals, generation of a spatial enhancement mask based on spatial information, and application of this mask to binaural audio signals to enhance spatial characteristics, along with processing residues to emphasize specific directions and adjusting levels and timbre.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If spatial enhancement processing is applied to binaural audio signals, then spatial perception and immersive experience are improved, but processing complexity and computational requirements increase

Engineering Contradiction:
Improvespatial perception accuracyVSAvoidprocessing complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The audio signal is segmented into multiple audio objects with distinct spatial characteristics. Each audio object is processed independently through the spatial enhancement pipeline, allowing complex processing to be distributed and managed separately rather than applied to the entire mixed signal at once.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

Spatial enhancement masks are pre-computed based on spatial information from multi-channel audio signals before being applied to binaural audio signals. This preliminary processing of spatial characteristics separates the complex computational tasks from the final rendering stage.

Inventive Principle:
Principle #10Preliminary action

2Measurement precision

If multiple audio capture devices are used to capture spatial audio, then spatial characteristics are improved, but device synchronization and integration become more difficult

Engineering Contradiction:
Improvespatial characteristics accuracyVSAvoiddevice integration complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

A processing system acts as an intermediary between multiple audio capture devices and the final output. This intermediary receives signals from various devices, performs synchronization and spatial analysis, and integrates the data into a unified spatial representation, isolating the complexity of multi-device coordination from the rendering pipeline.

Inventive Principle:
Principle #24Intermediary (Mediator)

3Measurement precision

If spatial enhancement mask is applied to binaural audio signal, then spatial perception is enhanced, but processing time and computational resources increase

Engineering Contradiction:
Improvespatial perceptionVSAvoidprocessing time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

Spatial enhancement masks are generated in advance based on spatial information from multi-channel audio signals. By pre-computing these masks before applying them to binaural signals, the system separates the computationally intensive spatial analysis from the time-critical audio rendering operation.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The spatial enhancement mask is applied selectively to specific audio objects and their corresponding spatial regions rather than uniformly across the entire audio spectrum. This localized processing reduces the overall computational burden while maintaining enhancement quality where it matters most.

Inventive Principle:
Principle #3Local quality

Data Source

PatentUS20260046587A1Spatial enhancement for user-generated content
Publication Date: 2026.02.12 DOLBY LABORATORIES LICENSING CORP
  • US20260046587A1 patent drawing
  • US20260046587A1 patent drawing
  • US20260046587A1 patent drawing

AI summary

Methods, systems, and media for enhancing audio content are provided. In some embodiments, a method for enhancing audio content involves receiving a multi-channel audio signal from a first audio capture device and a binaural audio signal from a second audio capture device. The method may further involve extracting one or more objects from the multi-channel audio signal. The method may further involve generating a spatial enhancement mask based on spatial information associated with the one or more objects. The method may further involve applying the spatial enhancement mask to the binaural audio signal to enhance spatial characteristics of the binaural audio signal to generate an enhanced binaural audio signal. The method may further involve generating output binaural audio signal based on the enhanced binaural audio signal.