Stereo Audio Synthesis Using HRTF Spatial Filtering

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current methods for converting multi-channel surround sound to stereo fail to preserve the immersive audio experience, as they lack the ability to accurately simulate the spatial cues and frequency responses associated with three-dimensional audio environments, resulting in a loss of realism when played back through stereo headphones.

Innovation Solution

A method and system that utilize filter response functions based on azimuth and elevation coordinates to synthesize left and right stereo output signals, allowing for the simulation of 3D audio environments by applying filter responses to audio inputs, which can include moving audio sources along calculated trajectories, and upsampling to create additional virtual sources, enhancing the listener's experience with spatial audio cues.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If multi-channel surround sound is down-mixed to stereo for headphone playback, then compatibility with stereo devices is improved, but the immersive audio experience and spatial realism are lost

Engineering Contradiction:
Improvecompatibility with stereo devicesVSAvoidspatial audio information
Core Design Contradiction:
Adaptability or versatilityVSLoss of information

Solution Approach 1:

The patent applies Head-Related Transfer Functions (HRTFs) to transform spatial audio information from multiple dimensions (multi-channel surround) into a format suitable for two-dimensional stereo headphone playback. The HRTFs encode azimuth and elevation coordinates into filter responses that preserve 3D spatial cues in the stereo output, effectively adding a dimensional transformation layer that maintains spatial realism while adapting to stereo device constraints.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

2Ease of operation

If traditional down-mixing methods are used to convert surround sound to stereo, then device compatibility is improved, but spatial cues and frequency responses associated with 3D audio environments are degraded

Engineering Contradiction:
Improveplayback on stereo headphonesVSAvoidspatial audio accuracy
Core Design Contradiction:
Ease of operationVSMeasurement precision

Solution Approach 1:

The patent transforms spatial parameters (azimuth and elevation coordinates) into filter response parameters through HRTF-based processing. Each audio channel is convolved with impulse responses corresponding to the spatial location of virtual loudspeakers, changing the frequency and temporal parameters of the audio signal to encode spatial information. This parameter transformation preserves spatial accuracy while enabling stereo headphone playback.

Inventive Principle:
Principle #35Parameter changes

3Adaptability or versatility

If phantom loudspeaker techniques are used to create spatial perception over headphones, then the sensation of multiple loudspeakers is achieved, but the accuracy of spatial positioning and frequency response varies with elevation and azimuth

Engineering Contradiction:
Improvevirtual loudspeaker positioningVSAvoidspatial positioning accuracy
Core Design Contradiction:
Adaptability or versatilityVSManufacturing precision

Solution Approach 1:

The patent introduces HRTF impulse responses as intermediary elements between the audio signal and the listener's ears. These impulse responses act as filters that mediate the spatial positioning information, encoding azimuth and elevation coordinates into frequency and temporal characteristics that the human auditory system can interpret for accurate spatial localization. The HRTFs serve as a translation layer that bridges virtual loudspeaker positions and perceived spatial location.

Inventive Principle:
Principle #24Intermediary (Mediator)

Data Source

PatentEP3406088B1Synthesis of signals for immersive audio playback
Publication Date: 2022.03.02 SPHEREO SOUND LTD
  • EP3406088B1 patent drawingFigure 1
  • EP3406088B1 patent drawingFigure 2
  • EP3406088B1 patent drawingFigure 3

AI summary

A method for synthesizing sound includes receiving one or more first inputs (80), each including a respective monaural audio track (82). One or more second inputs are received, indicating respective three-dimensional (3D) source locations having azimuth and elevation coordinates to be associated with the first inputs. Each of the first inputs is assigned respective left and right filter responses based on filter response functions that depend upon the azimuth and elevation coordinates of the respective 3D source locations. Left and right stereo output signals (94) are synthesized by applying the respective left and right filter responses to the first inputs.