Metaverse User Localization Using Visual-IMU Fusion and Stable Keypoints
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing user localization methods in the metaverse suffer from poor accuracy due to reliance on visual tracking that fails in low-light or dynamic environments, leading to unrealistic user displacements and loss of tracking, and lack of integration with inertial movements.
Innovation Solution
A method using inertial sensors and cameras to generate data, processed by ML models to extract stable key points, filter out dynamic points, and integrate visual and sensor data with weighted importance to enhance localization accuracy.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If visual tracking is used for user localization, then localization can be implemented, but tracking reliability deteriorates in low-light or dynamic environments
Solution Approach 1:
The patent combines visual tracking data from cameras with inertial measurement unit (IMU) sensor data to create a hybrid localization system. The sensor fusion module integrates both data sources, using visual information when available and reliable, and switching to or supplementing with inertial data when visual tracking fails or is unreliable, thereby maintaining continuous and accurate user localization across diverse environmental conditions.
Solution Approach 2:
The patent introduces an intermediary sensor fusion module that mediates between visual tracking and inertial measurement systems. This intermediary component processes and reconciles data from both sources, filtering out inconsistencies and combining complementary information to produce robust localization results that are more reliable than either system alone.
2Measurement precision
If visual tracking is used for user localization, then localization can be implemented, but localization precision deteriorates due to errors in tracking
Solution Approach 1:
The patent implements feedback mechanisms where the system continuously monitors the quality and reliability of visual tracking data. When tracking errors are detected or confidence levels drop, the system adjusts its reliance on visual data versus inertial data accordingly, using feedback to dynamically optimize the fusion weights and maintain high localization precision despite imperfect visual tracking.
Solution Approach 2:
The patent supplements the visual tracking mechanism with an inertial measurement system that operates on different physical principles. By replacing reliance on purely visual mechanisms with a hybrid approach that includes inertial sensors, the system compensates for visual tracking limitations and maintains precise localization even when visual data becomes inaccurate or unavailable.
3Measurement precision
If key points are extracted from visual data for localization, then user position can be determined, but unrealistic user movements occur due to dynamic objects
Solution Approach 1:
The patent extracts and separates reliable tracking information from unreliable visual data by identifying and filtering out key points that correspond to dynamic objects. The system selectively extracts only those visual features that provide accurate user position information, removing harmful contributions from moving objects that would otherwise cause unrealistic user movement artifacts.
Solution Approach 2:
The patent segments the visual field into different regions and classifies extracted key points based on their reliability and stability. By segmenting the tracking data into reliable and unreliable components, the system can process only the trustworthy information for localization while discarding or weighting down contributions from dynamic objects that would introduce distortion.
Data Source
AI summary
The present disclosure provides a method for intelligent user localization in a metaverse, including: detecting movements of a wearable head gear configured to present virtual content to a user, and generating sensor data and visual data using an inertial sensor and a camera, respectively, mapping the visual data to a virtual world using an image associated with the visual data to localize the user in the virtual world; providing the visual data and the sensor data to a first Machine Learning (ML) model and a second ML model, respectively; extracting a plurality of key points from the visual data and distinguishing stable key points and dynamic key points; and removing visual impacts corresponding to the visual data having a relatively low weightage, and providing a relatively high weightage to the sensor data processed through the second ML model.


