3D Audio Virtual Sound Sources Using AI-Based HRTF Mapping
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional methods for generating 3D audio virtualization are limited by non-individualized, incorrect, or missing Head-Related Transfer Functions (HRTFs), leading to a compromised 3D spatial perception and listener experience, particularly in scenarios where sound sources are not accurately reproduced due to missing or incorrectly positioned acoustic transducers.
Innovation Solution
The use of spatial Mapping Transfer Functions (MTFs) generated through machine learning, which transform HRTFs to accurately position virtual sound sources in 3D space by convolving and mixing audio data, utilizing individualized HRTF data sets and AI algorithms to simulate sound sources at intended locations.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If conventional HRTF methods are used for 3D audio virtualization, then the system can simulate sound sources in 3D space, but the spatial perception is compromised due to non-individualized or incorrect HRTFs
Solution Approach 1:
The patent changes the parameters of HRTF from fixed, non-individualized values to dynamic, individualized parameters by using machine learning models that adapt to each user's unique acoustic characteristics. This allows the system to generate accurate HRTFs tailored to individual listeners, resolving the contradiction between spatial perception accuracy and HRTF reliability.
Solution Approach 2:
The patent creates a virtual copy of the user's acoustic characteristics through machine learning models that replicate individual HRTF patterns. Instead of using generic HRTF data, the system generates personalized HRTF copies for each user, enabling accurate spatial perception without requiring physical measurements of each user's anatomy.
2Device complexity
If headset type devices use only two transducers positioned close to the ear canal, then the device structure is simplified, but suitable HRTFs for far-field sound sources cannot be generated
Solution Approach 1:
The patent replaces the mechanical approach of using multiple physical transducers positioned at specific locations with a computational approach using machine learning models. Instead of physically replicating far-field sound sources with multiple transducers, the system uses AI algorithms to generate accurate HRTFs that simulate the acoustic effects of distant sound sources, thereby maintaining simple device structure while achieving high HRTF generation accuracy.
3Adaptability or versatility
If sound sources are positioned in far-field spherical space with no corresponding physical transducers, then the 3D audio experience is enhanced, but the required HRTFs are missing or distorted
Solution Approach 1:
The patent performs preliminary action by pre-training machine learning models with comprehensive HRTF data from multiple sources and scenarios. These pre-trained models contain the necessary HRTF information for far-field sound sources before the actual audio playback occurs, allowing the system to generate complete and accurate HRTFs on-demand without requiring physical measurements or pre-existing individualized data.
Solution Approach 2:
The patent introduces machine learning models as an intermediary between the desired 3D audio experience and the available HRTF data. The AI models act as mediators that synthesize complete HRTF information from partial or generic data, filling in the missing information for far-field sound sources and enabling accurate spatial audio reproduction without direct physical measurement.
Data Source
AI summary
Embodiments of the present disclosure provide systems and/or methods for reproducing three-dimensional sound through the generation of acoustic virtual sound sources located within a three-dimensional space surrounding a designated user (listener) or origin point. The present disclosure provides 3D audio virtualization by generating a spatial Mapping Transfer Function (MTF) from data sets of measured HRTFs, modeled HRTFs or a combination thereof that transforms a known HRTF for a real acoustic transducer or loudspeaker in a system to a new HRTF for a virtual acoustic transducer, loudspeaker or sound object in a system. The MTF may be generated using data analysis algorithms and may be produced with high accuracy using supervised machine learning (AI). Convolving MTFs with audio data, existent HRTFs in the system and subsequently mixing the result into present audio or sound reproduction channels may enable existing acoustic transducers or loudspeakers to reproduce a virtual sound source.


