Personalized HRTF via Video Capture for Spatial Audio
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Personalized audio delivery devices, such as headphones and earbuds, fail to accurately convey sound direction due to the absence of interaction with human anatomy, as the sound does not reflect or scatter with the outer ear, head, and torso, leading to a lack of perceived spatialization.
Innovation Solution
A system utilizing a video capture device to capture images of a user's anatomy, which analyzes features like demographics, accessories, and anatomy to determine a personalized Head-Related Transfer Function (HRTF) through image processing and machine learning techniques, allowing for the spatialization of sound by applying the predicted HRTF to audio output.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If personalized audio delivery devices output sound directly into the ear canal, then sound reproduction accuracy is improved, but spatial perception capability deteriorates because the sound does not interact with human anatomy
Solution Approach 1:
The patent captures images of the user's pinna and head anatomy, then creates a digital 3D model that copies the unique anatomical features. This digital model is used to generate a personalized HRTF that replicates how sound would naturally interact with the user's anatomy, enabling spatial perception without actual sound-anatomy interaction through headphones
Solution Approach 2:
The patent introduces HRTF (Head-Related Transfer Function) as an intermediary that mediates between the audio output and the user's perception. The HRTF processes the audio signal to simulate the acoustic effects of sound interacting with the user's unique anatomy, allowing the brain to interpret spatial information even though the sound bypasses the actual anatomical structures
2Measurement precision
If HRTF personalization is implemented to enable spatial perception, then spatialization accuracy is improved, but system complexity increases due to anatomy capture and processing requirements
Solution Approach 1:
The patent enables the user to capture their own anatomical images using their mobile device's camera, eliminating the need for specialized measurement equipment or professional operators. The system then automatically processes these images through machine learning models to generate the personalized HRTF, making the complex personalization process accessible and automated
Solution Approach 2:
The patent replaces traditional mechanical measurement systems (physical scanners, laser range finders, or manual anthropometry) with an optical system using standard camera images. Computer vision and machine learning algorithms substitute for complex mechanical measurement and processing equipment, simplifying the overall system while maintaining accuracy
Data Source
AI summary
A video is received from a video capture device. The video capture device has a front facing camera and a display screen which displays the video captured by the video capture device in real time to a user. One or more images of a pinna and head of the user in the video are used to automatically determine one or more features associated with the user. The one or more features include an anatomy of the user, a demographic of the user, a latent feature of the user, and an indication of an accessory worn by the user. Based on the one or more features and one or more HRTF models, a head related transfer function (HRTF) is determined which is personalized to the user.


