Multi-Modal 6DoF Tracking for Low-Latency Spatial Audio

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional systems for dynamic spatial audio rendering lack accurate six degrees of freedom (6DoF) tracking and require specialized hardware or complex sensor setups, leading to latency and inefficiencies in mobile and VR applications.

Innovation Solution

A source device employs multi-modal sensor fusion using computer vision algorithms and inertial measurement units to track user orientation and position, eliminating the need for specialized hardware and reducing latency by performing all calculations at the source device.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If conventional systems use specialized hardware or complex sensor setups for 6DoF tracking, then tracking accuracy is improved, but device complexity and cost increase

Engineering Contradiction:
Improvetracking accuracyVSAvoidhardware complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent combines computer vision algorithms with IMU sensor data to achieve accurate 6DoF tracking. The system merges visual information from the camera with inertial measurement data to determine user orientation and position, eliminating the need for specialized tracking hardware while maintaining high measurement precision.

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The source device performs all tracking calculations itself using its own camera and IMU sensors. The system serves its own tracking needs by processing sensor data locally to generate spatial audio rendering, without requiring external base stations or specialized user device sensors.

Inventive Principle:
Principle #25Self-service

2Speed

If conventional systems use two-way roundtrip motion-to-sound latency for head-tracking, then tracking responsiveness is improved, but latency increases

Engineering Contradiction:
Improvetracking responsivenessVSAvoidlatency
Core Design Contradiction:
SpeedVSLoss of time

Solution Approach 1:

Instead of the conventional approach where the user device tracks motion and sends data to the source device for audio rendering, this system inverts the workflow: the source device captures sensor data directly and performs all spatial audio rendering calculations itself, eliminating the two-way communication loop and reducing latency.

Inventive Principle:
Principle #13The other way round (Inversion)

3Measurement precision

If conventional VR systems use lighthouse tracking with external base stations, then six DoF tracking accuracy is improved, but device complexity and cost increase

Engineering Contradiction:
Improvesix DoF tracking accuracyVSAvoidsystem complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent extracts the tracking functionality from external infrastructure (lighthouse base stations) and implements it using the source device's own sensors and computer vision capabilities. The system takes out the dependency on external tracking infrastructure and achieves 6DoF tracking using only the source device's camera and IMU.

Inventive Principle:
Principle #2Taking out (Extraction)

4Measurement precision

If conventional systems use inside-out tracking with multiple sensors, then six DoF tracking is achieved, but reliability for external source device audio rendering decreases

Engineering Contradiction:
Improvesix DoF trackingVSAvoidaudio rendering reliability
Core Design Contradiction:
Measurement precisionVSReliability

Solution Approach 1:

The patent uses computer vision algorithms as an intermediary to bridge the gap between sensor data and spatial audio rendering. The CV algorithms process visual information to complement IMU data, providing more reliable and accurate orientation and position information for external source device audio rendering compared to inside-out tracking alone.

Inventive Principle:
Principle #24Intermediary (Mediator)

Data Source

PatentUS20260052354A1Method and system of multi-modal tracking for dynamic spatial audio rendering
Publication Date: 2026.02.19 QUALCOMM INC
  • US20260052354A1 patent drawing
  • US20260052354A1 patent drawing
  • US20260052354A1 patent drawing

AI summary

A device includes a memory configured to store multi-channel audio content. The device also includes one or more processors coupled to the memory and configured to obtain first information based on first sensor data from a first sensor and to obtain second information based on second sensor data from a second sensor. The one or more processors are further configured to select, based on the first information, the second information, or a combination thereof, a determination scheme. The one or more processors are configured to generate, based on the determination scheme, determination information associated with an audio output device. The determination information indicates an orientation, a position, or a combination thereof. The one or more processors are configured to generate, based on the determination information and the multi-channel audio content, a spatial audio output associated with the audio output device.