HRTF-Based Binaural Signal Post-Processing for Object Separation

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing audio post-processing systems for binaural audio struggle with frequency-dependent level and time differences, leading to suboptimal manipulation and separation of objects in binaural audio signals.

Innovation Solution

A method involving signal transformation, spatial analysis, and object extraction using estimated head-related transfer functions (HRTFs) to separate and process binaural audio signals into main and residual components, allowing for frequency-dependent and time-dependent processing.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If linear gains are used for audio source separation, then channel-based signals such as stereo audio can be processed effectively, but frequency-dependent level and time differences in binaural audio cannot be handled properly

Engineering Contradiction:
Improveadaptability to binaural audioVSAvoidseparation accuracy
Core Design Contradiction:
Adaptability or versatilityVSManufacturing precision

Solution Approach 1:

The patent changes the processing parameters from fixed linear gains to frequency-dependent parameters including HRTF-based level differences and phase differences. This allows the system to adapt to the frequency-dependent characteristics of binaural audio while maintaining accurate source separation, resolving the contradiction between versatility for binaural audio and precision for separation.

Inventive Principle:
Principle #35Parameter changes

Solution Approach 2:

The system dynamically adjusts processing parameters based on frequency and spatial information extracted from the binaural audio signal. By making the gains frequency-dependent and adapting them to the specific characteristics of each frequency band, the system achieves both broad adaptability to binaural formats and precise separation accuracy.

Inventive Principle:
Principle #15Dynamics

2Ease of operation

If relative changes due to rotation are performed on the full mix or coherent element only, then re-orientation can be achieved, but object-specific processing cannot be applied

Engineering Contradiction:
Improvere-orientation capabilityVSAvoidobject-specific processing capability
Core Design Contradiction:
Ease of operationVSAdaptability or versatility

Solution Approach 1:

The patent segments the binaural audio signal into multiple frequency bands and spatial components, allowing independent processing of each segment. This segmentation enables the system to perform re-orientation on the full mix while simultaneously applying object-specific processing to individual frequency components and spatial elements.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The system applies different processing quality and parameters to different local regions of the audio signal. By making the processing local to specific frequency bands and spatial locations, the system can maintain ease of operation for overall re-orientation while enabling sophisticated object-specific processing where needed.

Inventive Principle:
Principle #3Local quality

Data Source

PatentUS12413929B2Binaural signal post-processing
Publication Date: 2025.09.09 DOLBY LABORATORIES LICENSING CORP
  • US12413929B2 patent drawing
  • US12413929B2 patent drawing
  • US12413929B2 patent drawing

AI summary

A method of audio processing includes performing spatial analysis on a binaural signal to estimate level differences and phase differences characteristic of a binaural filter of the binaural signal, performing object extraction on the binaural audio signal using the estimated level and phase differences to generate a left/right main component signal and a left/right residual component signal. The system may process the left/right main and left/right residual components differently using different object processing parameters for e.g. repositioning, equalization, compression, upmixing, channel remapping or storage to generate a processed binaural signal that provides an improved listening experience. Repositioning may be based on head tracking sensor data.