Hybrid BRIR/HRTF Data for Sparse Binaural Spatial Rendering

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing binaural rendering technologies face challenges with sparse, noisy, or corrupted binaural room impulse response (BRIR) and head-related transfer function (HRTF) data sets, leading to suboptimal spatial audio rendering with artifacts like coloration in timbre and imprecise localization, especially when room effects are not accounted for.

Innovation Solution

A binaural renderer that combines user-specific, potentially sparse or corrupted BRIR/HRTF data sets with predefined, high-quality data sets, applying perceptual matching to ensure accurate directional perception and uncolored timbre, even with low directional resolution.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If user-specific BRIR/HRTF data sets are used for binaural rendering, then personalization and accuracy are improved, but data quality and completeness deteriorate due to sparsity, noise, or corruption

Engineering Contradiction:
Improvedirectional accuracyVSAvoiddata quality
Core Design Contradiction:
Measurement precisionVSReliability

Solution Approach 1:

The patent combines user-specific BRIR/HRTF data sets with predefined high-quality data sets to create a hybrid data set. The rendering device merges the personalized measurements with reference data, using the predefined data to supplement and correct the user-specific data, thereby maintaining both personalization and data quality reliability.

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The predefined BRIR/HRTF data sets act as an intermediary reference standard. When user-specific data is sparse, noisy, or corrupted, the system uses the predefined data as a mediator to fill gaps and correct errors, ensuring consistent and reliable rendering without requiring perfect user measurements.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Manufacturing precision

If high-resolution BRIR/HRTF data sets are used, then rendering quality is improved, but data sparsity and directional coverage deteriorate

Engineering Contradiction:
Improverendering qualityVSAvoiddata density
Core Design Contradiction:
Manufacturing precisionVSQuantity of substance

Solution Approach 1:

The system merges dense predefined data with user-specific measurements to achieve both high rendering quality and comprehensive directional coverage. The predefined data provides dense sampling across all directions, while user-specific data adds personalized characteristics, resulting in a combined data set that is both dense and personalized.

Inventive Principle:
Principle #5Merging (Combining)

3Device complexity

If room effects are excluded from rendering, then processing complexity is reduced, but spatial accuracy and naturalness deteriorate

Engineering Contradiction:
Improveprocessing complexityVSAvoidspatial accuracy
Core Design Contradiction:
Device complexityVSMeasurement precision

Solution Approach 1:

The predefined BRIR data sets include pre-computed room effects and acoustic characteristics. By using these pre-processed data, the system incorporates accurate spatial information and room interactions without performing complex real-time room modeling, thus maintaining high spatial accuracy while reducing processing complexity.

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS12425800B2Spatial audio representation and rendering
Publication Date: 2025.09.23 NOKIA TECHNOLOGIES OY
  • US12425800B2 patent drawing
  • US12425800B2 patent drawing
  • US12425800B2 patent drawing

AI summary

An apparatus including circuitry configured to: obtain a spatial audio signal including at least one audio signal and spatial metadata associated with the at least one audio signal; obtain at least one data set related to binaural rendering; obtain at least one pre-defined data set related to binaural rendering; and generate a binaural audio signal based on a combination of at least part of the at least one data set and the at least one pre-defined data set, and the spatial audio signal.