Deep Inertial Pose Prediction for Occluded XR Tracking
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional extended reality systems face tracking loss when controllers or devices become obstructed or exit the field of view, leading to interruptions and loss of seamless interaction with virtual environments.
Innovation Solution
A deep inertial prediction system utilizing a machine learned model, such as a convolutional neural network (CNN), in conjunction with an Extended Kalman Filter (EKF), to maintain tracking by integrating IMU data and correcting for biases, ensuring seamless user experience by predicting device poses even when visual tracking is lost.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If conventional visual tracking systems are used to track device pose, then tracking accuracy is maintained when devices are visible, but tracking is lost when devices become obstructed or exit field of view
Solution Approach 1:
The system performs preliminary actions by training deep learning models offline to predict device pose from IMU data. When visual tracking is lost, these pre-trained models immediately activate to predict and maintain pose information, preventing tracking loss without requiring real-time visual data.
Solution Approach 2:
The patent introduces IMU sensors as an intermediary between the device and the tracking system. When visual tracking fails, the IMU data serves as a mediator that feeds into deep learning models to infer device pose, maintaining the tracking chain without direct visual contact.
2Reliability
If IMU data is used for pose prediction, then tracking can be maintained during visual occlusion, but drift and error accumulate over time
Solution Approach 1:
The system implements feedback by continuously comparing deep learning pose predictions with available visual tracking data. When visual tracking is available, it corrects the predicted pose, creating a feedback loop that prevents drift accumulation while maintaining continuity during occlusion.
Solution Approach 2:
The patent merges multiple data sources - IMU measurements, deep learning predictions, and visual tracking data - into a unified pose estimation system. This combination allows the system to leverage the strengths of each method while compensating for their individual weaknesses, particularly reducing drift through sensor fusion.
3Productivity
If deep learning models are trained on synthetic data, then training efficiency is improved, but realism and generalization to real-world scenarios may be reduced
Solution Approach 1:
The system creates synthetic copies of real-world tracking scenarios through rendered virtual environments. These synthetic datasets replicate real device motions and visual appearances, providing realistic training data without requiring extensive real-world data collection, thus maintaining both efficiency and generalization.
Solution Approach 2:
The patent employs parameter changes by systematically varying synthesis parameters such as device colors, textures, lighting conditions, and camera positions during data generation. This creates diverse training samples that improve model robustness and generalization to different real-world conditions while maintaining training efficiency.
Data Source
AI summary
A system configured to determine poses of a tracked device in a physical environment and to utilize the poses as an input to control or manipulate a virtual environment or mixed reality environment. In some cases, the system may include a fall back tracking system for when the main tracking system loses visual tracking of the tracked device.


