Binaural Renderer Adjusting HRTF for Large Sound Source Size

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current audio signal processing methods for binaural rendering struggle to accurately reproduce the three-dimensionality of sound sources, particularly when the size of the sound source is large relative to the distance from the listener, as they fail to differentiate between audio signals based on object sizes, leading to inadequate spatial audio representation.

Innovation Solution

An audio signal processing device and method that utilizes a binaural renderer to generate 2-channel audio by adjusting the head-related transfer function (HRTF) based on the distance and size of the sound source, incorporating pseudo HRTFs and interaural cross-correlation (IACC) adjustments to simulate the size and distance of sound sources, allowing for more accurate spatial audio representation.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If binaural rendering is performed using conventional HRTF methods, then the processing is simple and fast, but the three-dimensionality and spatial accuracy of sound sources are inadequate, especially when object size is large relative to distance

Engineering Contradiction:
Improvespatial accuracyVSAvoidprocessing complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent segments the sound source into multiple virtual point sources distributed across the object's spatial extent. Instead of treating the sound source as a single point, it divides the object into multiple locations and generates HRTF-based audio signals for each point source, then combines them. This segmentation enables accurate representation of large objects while maintaining manageable processing through systematic decomposition.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent transitions from conventional 2D stereo audio to 3D spatial audio by incorporating the height dimension and spatial distribution of sound sources. It uses three-dimensional position information and generates binaural signals that exploit vertical localization cues through head-related transfer functions, enabling accurate spatial representation in three-dimensional space rather than flat stereo.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

2Measurement precision

If multiple HRTFs are used to accurately represent large sound sources, then spatial accuracy improves, but computational complexity increases significantly

Engineering Contradiction:
Improvespatial accuracyVSAvoidcomputational power
Core Design Contradiction:
Measurement precisionVSPower

Solution Approach 1:

The patent divides the sound source into a finite number of discrete virtual point sources rather than attempting continuous representation. This segmentation into manageable points reduces computational complexity while maintaining spatial accuracy, as each point source can be processed independently with standard HRTF convolution.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent uses a sufficient number of virtual point sources to achieve accurate spatial representation without over-processing. By selecting an appropriate number of points based on object size and distance, it achieves the necessary spatial accuracy without the excessive computational burden of using too many points, balancing precision and efficiency.

Inventive Principle:
Principle #16Partial or excessive action

3Ease of operation

If conventional binaural rendering is used, then the implementation is straightforward, but the ability to reproduce size-dependent spatial cues is lost

Engineering Contradiction:
Improveimplementation easeVSAvoidspatial information
Core Design Contradiction:
Ease of operationVSLoss of information

Solution Approach 1:

The patent pre-calculates and stores HRTF-based impulse responses for multiple virtual point sources at different spatial locations before actual audio playback. This preliminary preparation of spatial filtering data allows the system to maintain implementation simplicity during runtime while having accurate spatial information ready, reducing real-time computational burden while preserving spatial cues.

Inventive Principle:
Principle #10Preliminary action

Applied Scientific Principles

This section explains which scientific principles are used to turn an abstract innovation direction into a practical engineering solution.

Function Achieved in This Case

The solution effectively enhances the three-dimensional audio experience by accurately simulating sound source sizes and distances, improving the spatial representation and reducing tone distortion, while reducing computational complexity through the use of pseudo HRTFs and IACC adjustments.

Implementation Method 1

The binaural renderer may determine a characteristic of a head related transfer function (HRTF) based on the distance from the listener to the sound source and the size of the object simulated by the sound source, and may perform binaural rendering on the input audio signal using the HRTF.

Methodology Applied
Scientific EffectHead-related transfer function (HRTF):

Implementation Method 2

The HRTF may be a pseudo HRTF generated by adjusting an initial time delay of an HRTF corresponding to a path from the listener to the sound source based on the distance from the listener to the sound source and the size of the object simulated by the sound source.

Methodology Applied
Scientific EffectTime delay adjustment:

Data Source

PatentUS10349201B2Apparatus and method for processing audio signal to perform binaural rendering
Publication Date: 2019.07.09 GAUDI AUDIO LAB
  • US10349201B2 patent drawing
  • US10349201B2 patent drawing
  • US10349201B2 patent drawing

AI summary

Disclosed is an audio signal processing device for performing binaural rendering on an input audio signal. The audio signal processing device includes a reception unit configured to receive the input audio signal, a binaural renderer configured to generate a 2-channel audio by performing binaural rendering on the input audio signal, and an output unit configured to output the 2-channel audio. The binaural renderer performs binaural rendering on the input audio signal based on a distance from a listener to a sound source corresponding to the input audio signal and a size of an object simulated by the sound source.