HRTF Audio Processing for Head-Tracked 3D Sound Realism

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current audio processing methods struggle to simulate a realistic sound field effectively, particularly in adjusting interaural level and time differences for personalized audio experiences.

Innovation Solution

A method and device that acquire a user's head image and audio signal, determine head attitude angles and sound source distance, and input these into a head-related transfer function to process left and right channel audio signals, adjusting loudness and time differences to mimic real-world audio experiences.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Manufacturing precision

If conventional audio processing methods are used to simulate sound field, then the processing method is simple, but the audio playing effect is not close to reality

Engineering Contradiction:
Improveaudio playing effect realismVSAvoidaudio processing complexity
Core Design Contradiction:
Manufacturing precisionVSDevice complexity

Solution Approach 1:

The system pre-calculates and stores head-related transfer functions (HRTF) for various head attitudes and sound source distances before actual audio playback. When processing audio, the system only needs to retrieve pre-computed HRTF data based on detected head attitude and distance, rather than performing complex real-time calculations, thus achieving realistic 3D audio effects without excessive computational complexity

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The system dynamically adjusts audio processing parameters based on real-time head attitude detection and sound source distance estimation. By continuously tracking head movements and updating the HRTF application accordingly, the system creates a dynamic 3D audio experience that adapts to user behavior, improving realism while maintaining manageable processing complexity through event-driven updates

Inventive Principle:
Principle #15Dynamics

2Adaptability or versatility

If generic audio processing is used, then the processing is fast, but it does not consider user-specific head attitude angles and sound source distances

Engineering Contradiction:
Improveuser-specific audio processingVSAvoidaudio processing efficiency
Core Design Contradiction:
Adaptability or versatilityVSProductivity

Solution Approach 1:

The system pre-computes and stores head-related transfer functions for multiple discrete head attitudes and sound source distances. Instead of calculating custom HRTF for every user situation in real-time, the system selects from pre-prepared HRTF data sets that match detected head attitude ranges, achieving user-specific adaptation without sacrificing processing speed

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The system creates simplified representations of complex acoustic environments by using pre-measured or pre-simulated HRTF data that captures the essential characteristics of sound propagation for different head attitudes. These copied HRTF profiles are then applied to audio signals, providing personalized 3D audio effects while maintaining computational efficiency

Inventive Principle:
Principle #26Copying

Data Source

PatentUS11425524B2Method and device for processing audio signal
Publication Date: 2022.08.23 DOUYIN VISION CO LTD
  • US11425524B2 patent drawing
  • US11425524B2 patent drawing
  • US11425524B2 patent drawing

AI summary

The embodiments of the present disclosure disclose a method and device for processing an audio signal. A specific embodiment of the method includes acquiring a head image of a target user and an audio signal to be processed; determining head attitude angles of the target user based on the head image, and determining a distance between a target sound source and the head of the target user; and inputting the head attitude angles, the distance and the audio signal to be processed into a preset head related transfer function to obtain a processed left channel audio signal and a processed right channel audio signal, wherein the head related transfer function is used to characterize a correspondence between the head attitude angles, the distance and the audio signal to be processed, and the processed left channel audio signal and the processed right channel audio signal.