Binaural Downmixing Separation for Low-Latency Head Tracking
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing binaural processing techniques for wireless audio systems face challenges in maintaining low latency and efficient head-tracking due to channel limitations and high latency in wireless transmission protocols, preventing seamless three-dimensional sound experiences.
Innovation Solution
Separate channels into head-tracked and fixed channels, performing binaural downmixing on fixed channels at the source device and head-tracking on head-tracked channels at the output device, using encoding techniques to transmit within channel limitations, such as mid-side encoding and Stereo Quadraphony, to reduce computational load.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Quantity of substance
If binaural processing is performed on the source device prior to wireless transmission, then channel limitations are satisfied, but head-tracking latency exceeds perceptual thresholds
Solution Approach 1:
The patent segments the audio processing into two distinct parts: binaural downmixing is performed on the source device for non-head-tracked channels, while head-tracking processing is performed separately on the output device for head-tracked channels. This segmentation allows each processing stage to operate independently, satisfying channel limitations while maintaining low latency for head-tracking by keeping the sensor data processing local to the output device.
Solution Approach 2:
The patent introduces an intermediary approach where the source device transmits both the original multichannel audio signal and pre-downmixed stereo audio to the output device. The output device then acts as an intermediary that combines these signals, applying head-tracking only to the appropriate channels. This intermediary transmission strategy resolves the contradiction by allowing the output device to perform low-latency head-tracking on received channels without requiring high-bandwidth wireless transmission of processed head-tracked audio.
2Loss of time
If head-tracking processing is performed on the output device, then latency is reduced below perceptual thresholds, but computational load increases
Solution Approach 1:
The patent applies partial action by performing head-tracking processing only on the subset of channels that require head-tracking, rather than processing all audio channels. The source device performs binaural downmixing only on non-head-tracked channels, while the output device performs head-tracking only on head-tracked channels. This partial processing approach reduces the overall computational load compared to processing all channels at full resolution, while still achieving low latency for the critical head-tracked portions.
3Loss of information
If multichannel audio is transmitted wirelessly, then audio quality is maintained, but channel limitations of wireless protocols are exceeded
Solution Approach 1:
The patent applies preliminary action by performing binaural downmixing on the source device for non-head-tracked channels before wireless transmission. This pre-processing converts multichannel audio into a compact stereo format that fits within wireless protocol channel limitations. The source device transmits both the original multichannel signal and the pre-downmixed stereo signal, allowing the output device to reconstruct high-quality audio without exceeding transmission channel constraints.
Data Source
Figure 1
Figure 2
Figure 3
AI summary
Techniques for separating binaural downmixing and head tracking processing between a source device and an output device are described. Embodiments include receiving, by a source device, an audio signal comprising a plurality of channels and selecting a first subset of the plurality of channels as head-tracked channels. Embodiments include performing, by the source device, binaural downmixing on a second subset of the plurality of channels that is different than the first subset of the plurality of channels to produce a downmixed second subset of the plurality of channels. Embodiments include transmitting, by the source device, the first subset of the plurality of channels and the downmixed second subset of the plurality of channels to an output device. Embodiments include performing, by the output device, binaural downmixing on the first subset of the plurality of channels based on positional data captured via one or more sensors associated with the output device.