Multi-Modal 6DoF Tracking for Low-Latency Spatial Audio
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional systems for dynamic spatial audio rendering lack accurate six degrees of freedom (6DoF) tracking and require specialized hardware or complex sensor setups, leading to latency and inefficiencies in mobile and VR applications.
Innovation Solution
A source device employs multi-modal sensor fusion using computer vision algorithms and inertial measurement units to track user orientation and position, eliminating the need for specialized hardware and reducing latency by performing all calculations at the source device.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If conventional systems use specialized hardware or complex sensor setups for 6DoF tracking, then tracking accuracy is improved, but device complexity and cost increase
Solution Approach 1:
The patent combines computer vision algorithms with IMU sensor data to achieve accurate 6DoF tracking. The system merges visual information from the camera with inertial measurement data to determine user orientation and position, eliminating the need for specialized tracking hardware while maintaining high measurement precision.
Solution Approach 2:
The source device performs all tracking calculations itself using its own camera and IMU sensors. The system serves its own tracking needs by processing sensor data locally to generate spatial audio rendering, without requiring external base stations or specialized user device sensors.
2Speed
If conventional systems use two-way roundtrip motion-to-sound latency for head-tracking, then tracking responsiveness is improved, but latency increases
Solution Approach 1:
Instead of the conventional approach where the user device tracks motion and sends data to the source device for audio rendering, this system inverts the workflow: the source device captures sensor data directly and performs all spatial audio rendering calculations itself, eliminating the two-way communication loop and reducing latency.
3Measurement precision
If conventional VR systems use lighthouse tracking with external base stations, then six DoF tracking accuracy is improved, but device complexity and cost increase
Solution Approach 1:
The patent extracts the tracking functionality from external infrastructure (lighthouse base stations) and implements it using the source device's own sensors and computer vision capabilities. The system takes out the dependency on external tracking infrastructure and achieves 6DoF tracking using only the source device's camera and IMU.
4Measurement precision
If conventional systems use inside-out tracking with multiple sensors, then six DoF tracking is achieved, but reliability for external source device audio rendering decreases
Solution Approach 1:
The patent uses computer vision algorithms as an intermediary to bridge the gap between sensor data and spatial audio rendering. The CV algorithms process visual information to complement IMU data, providing more reliable and accurate orientation and position information for external source device audio rendering compared to inside-out tracking alone.
Data Source
AI summary
A device includes a memory configured to store multi-channel audio content. The device also includes one or more processors coupled to the memory and configured to obtain first information based on first sensor data from a first sensor and to obtain second information based on second sensor data from a second sensor. The one or more processors are further configured to select, based on the first information, the second information, or a combination thereof, a determination scheme. The one or more processors are configured to generate, based on the determination scheme, determination information associated with an audio output device. The determination information indicates an orientation, a position, or a combination thereof. The one or more processors are configured to generate, based on the determination information and the multi-channel audio content, a spatial audio output associated with the audio output device.


