ITD Crossfader for Binaural Audio Without Clicking Artifacts
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
In virtual reality, augmented reality, and mixed reality environments, accurately presenting audio signals to users while minimizing sonic artifacts such as 'clicking' sounds is challenging due to the need for rapid changes in audio signals to reflect object positions and orientations, which can compromise the immersiveness of the experience.
Innovation Solution
A wearable head device processes input audio signals to generate left and right output audio signals by applying a delay process, adjusting gains, and applying head-related transfer functions (HRTFs) to accurately simulate interaural time differences, using a combination of delay modules and cross-faders to smooth transitions and reduce sonic artifacts.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If rapid changes are applied to audio signals to reflect object positions and orientations, then spatial accuracy is improved, but sonic artifacts such as clicking sounds increase
Solution Approach 1:
The system pre-calculates and stores head-related transfer functions (HRTFs) for multiple discrete spatial positions before they are needed. When the sound source moves between positions, the system prepares the next HRTF in advance and transitions smoothly between pre-computed values, avoiding real-time calculation artifacts and clicking sounds while maintaining spatial accuracy.
Solution Approach 2:
The system dynamically selects and transitions between different HRTF sets based on the current and target spatial positions of sound sources. By implementing smooth cross-fading transitions between HRTF applications as the sound source moves, the system adapts to changing spatial requirements without generating audible artifacts, resolving the contradiction between spatial accuracy and artifact reduction.
2Measurement precision
If multiple HRTF sets are applied to accurately represent different spatial positions, then spatial resolution is improved, but computational complexity increases
Solution Approach 1:
The continuous spatial environment is segmented into multiple discrete positional sectors, each associated with a specific HRTF set. Instead of computing HRTFs continuously for every possible position, the system divides space into manageable segments and assigns pre-computed HRTFs to each segment, reducing computational complexity while maintaining sufficient spatial resolution for binaural rendering.
Solution Approach 2:
HRTF sets for all discrete spatial positions are pre-computed and stored in memory before runtime. This preliminary action eliminates the need for real-time HRTF calculation during audio rendering, significantly reducing computational complexity while preserving spatial resolution by selecting from the pre-computed set based on current sound source position.
3Object-generated harmful factors
If smooth transitions are applied between audio signals, then sonic artifacts are reduced, but temporal responsiveness decreases
Solution Approach 1:
The system implements periodic cross-fading transitions at predetermined temporal intervals or position thresholds rather than continuously smoothing all movements. This periodic approach allows for abrupt position changes without artifacts at specific critical moments while maintaining smooth transitions during normal movement, balancing artifact reduction with temporal responsiveness.
Solution Approach 2:
The transition smoothness is dynamically adjusted based on the speed and distance of sound source movement. For small positional changes or slow movements, smooth cross-fading is applied to prevent artifacts. For large or rapid movements, the system reduces transition time to maintain temporal responsiveness, adaptively balancing artifact reduction with responsiveness to different movement scenarios.
Data Source
AI summary
Examples of the disclosure describe systems and methods for presenting an audio signal to a user of a wearable head device. In an example, a received first input audio signal is processed to generate a left output audio signal and a right output audio signal presented to ears of the user. Processing the first input audio signal comprises applying a delay process to the first input audio signal to generate a left audio signal and a right audio signal; adjusting gains of the left audio signal and the right audio signal; applying head-related transfer functions (HRTFs) to the left and right audio signals to generate the left and right output audio signals. Applying the delay process to the first input audio signal comprises applying an interaural time delay (ITD) to the first input audio signal, the ITD determined based on the source location.


