XR Pose Estimation Using Predicted 3D Landmarks
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing AR/VR headset systems face challenges in accurately estimating pose due to fast head motions, dynamic objects, and minimal distinguishing visual features, leading to drift accumulation and tracking loss, especially in indoor environments with motion blur and dynamic occlusions, which degrade user experience.
Innovation Solution
The system predicts 3D landmarks for partially visible objects using historical context and household object information, generating new landmarks to aid in robust visual matches, thereby enhancing pose estimation and reducing drift.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If visual SLAM techniques are used for pose estimation in indoor environments, then localization accuracy can be achieved, but fast head motions and dynamic occlusions cause loss of visual matches leading to drift accumulation and tracking loss
Solution Approach 1:
The system predicts 3D landmarks for future frames before they are actually observed, using the current camera pose and predicted future pose to generate expected landmark positions. This preliminary action ensures that landmarks are ready for matching even before the visual data is captured, preventing tracking loss during fast motions.
Solution Approach 2:
The patent introduces predicted 3D landmarks as an intermediary element that bridges the gap between visual SLAM and IMU data. These predicted landmarks serve as a mediator that can be matched with actual visual features when available, or used with IMU predictions when visual matches are lost, ensuring continuous and reliable pose estimation.
2Speed
If IMU-based pose estimation is used as a stop-gap measure when visual matches are unavailable, then short-term motion prediction is possible, but drift accumulation occurs in the estimated pose
Solution Approach 1:
The system uses a feedback mechanism where predicted 3D landmarks are continuously refined by comparing with actual visual features when they become available. The prediction error is fed back to adjust future predictions, and when visual matches are recovered, the system corrects accumulated drift by aligning predicted landmarks with actual observed landmarks.
Solution Approach 2:
The patent dynamically changes the weighting parameters between visual SLAM and IMU-based estimation based on the availability and quality of visual matches. When visual matches are abundant, visual weighting is increased; when matches are lost, IMU weighting is increased. This adaptive parameter adjustment optimizes the balance between speed and accuracy in different conditions.
3Adaptability or versatility
If the user suspends and resumes AR/VR device at different viewpoints, then device portability and flexibility are improved, but re-localization failures and delayed re-localization occur
Solution Approach 1:
The system performs preliminary mapping and stores comprehensive 3D landmark information during the initial mapping phase, including landmarks that may be visible from future viewpoints. When the user resumes at a different viewpoint, these pre-stored landmarks are quickly matched with the current view, enabling fast re-localization without requiring extensive new data collection.
4Device complexity
If conventional SLAM techniques are used in environments with minimal distinguishing visual features, then system simplicity is maintained, but visual matches are insufficient leading to drift accumulation
Solution Approach 1:
The patent introduces predicted 3D landmarks as an intermediary that enhances the effectiveness of limited visual features. By predicting where landmarks should be based on the current map and camera pose, the system can infer the presence and position of landmarks even when they are not clearly visible or distinguishable in the current frame, effectively amplifying the information from minimal visual features.
Data Source
AI summary
There is provided a method and device for estimating a pose in an XR environment by detecting a transition of an XR device from a first position to a second position, extracting at least one first set of objects from a real-world scene from a list of the plurality of 3D objects at the first position of the XR device, predicting at least one second set of 3D objects, from the list of the plurality of 3D objects, at the second position of the XR device, and estimating the pose at the second position of the XR device, using the at least one first extracted object at the first position and the at least one second predicted object at the second position of the XR device.


