Sparse HRIR Decomposition for Real-Time Spatial Audio

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current methods for obtaining listener-specific Head-Related Transfer Functions (HRTFs) are time-consuming and burdensome, requiring extensive measurements to replicate sound scattering in virtual reality environments, and lack efficient processing techniques for real-time spatial audio synthesis.

Innovation Solution

A method involving semi-non-negative matrix factorization to decompose Head-Related Impulse Responses (HRIRs) into direction-independent and direction-dependent filters, allowing for efficient convolution in the time domain and real-time processing of auditory signals for spatial audio synthesis in virtual reality.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If traditional HRTF measurement methods are used to obtain listener-specific head-related transfer functions, then accurate spatial audio reproduction is achieved, but the measurement process becomes time-consuming and burdensome requiring thousands of measurements

Engineering Contradiction:
ImproveHRTF measurement accuracyVSAvoidmeasurement time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The patent pre-computes a small set of basis HRIRs (head-related impulse responses) covering representative head positions before actual use. These basis HRIRs are stored and later combined through linear interpolation to generate HRIRs for arbitrary head positions, eliminating the need for time-consuming measurements at every possible position.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

Instead of measuring actual HRIRs for each listener at each position, the patent creates synthetic copies by interpolating between pre-computed basis HRIRs. This copying approach through linear interpolation provides sufficient accuracy for spatial audio applications while dramatically reducing measurement requirements.

Inventive Principle:
Principle #26Copying

2Measurement precision

If full HRIR filters are used for convolution operations to maintain audio quality, then accurate spatial rendering is achieved, but computational cost increases

Engineering Contradiction:
Improvespatial audio accuracyVSAvoidcomputational cost
Core Design Contradiction:
Measurement precisionVSPower

Solution Approach 1:

The patent decomposes the full HRIR filter into a linear combination of a small number of basis HRIRs (typically 6-12 basis functions). This segmentation allows the system to process only the small set of basis filters once, then efficiently combine them using simple weight coefficients during runtime, dramatically reducing computational cost while maintaining accuracy.

Inventive Principle:
Principle #1Segmentation

3Adaptability or versatility

If extensive HRTF measurements are performed to capture sound scattering for all source locations, then complete spatial coverage is achieved, but the process becomes tedious and burdensome to the listener

Engineering Contradiction:
Improvespatial coverageVSAvoidlistener burden
Core Design Contradiction:
Adaptability or versatilityVSEase of operation

Solution Approach 1:

The patent creates a universal set of basis HRIRs that can represent sound scattering for any source location and listener head position through linear interpolation. This single universal basis set serves all possible spatial configurations, eliminating the need for separate measurements for each position and making the system universally applicable without burdening the listener.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Data Source

PatentUS10237676B2Sparse decomposition of head related impulse responses with applications to spatial audio rendering
Publication Date: 2019.03.19 UNIV OF MARYLAND
  • US10237676B2 patent drawing
  • US10237676B2 patent drawing
  • US10237676B2 patent drawing

AI summary

This application describes methods of signal processing and spatial audio synthesis. One such method includes accepting an auditory signal and generating an impression of auditory virtual reality by processing the auditory signal to impute a spatial characteristic on it via convolution with a plurality of head-related impulse responses. The processing is performed in a series of steps, the steps including: performing a first convolution of an auditory signal with a characteristic-independent, mixed-sign filter and performing a second convolution of the result of first convolution with a characteristic-dependent, sparse, non-negative filter. In some described methods, the first convolution can be pre-computed and the second convolution can be performed in real-time, thereby resulting in a reduction of computational complexity in said methods of signal processing and spatial audio synthesis.