Head-Mounted Display Pose Tracking with Salient-Point Maps
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing augmented reality (AR) and virtual reality (VR) technologies face challenges in accurately determining the pose of a user's head and display system within a real-world environment, requiring complex hardware setups and leading to discomfort due to mismatches between accommodative and vergence states.
Innovation Solution
A display system that utilizes patch-based frame-to-frame tracking and descriptor-based map-to-frame tracking to determine pose, leveraging imaging devices to identify salient points and adjust virtual content placement based on head pose without requiring fixed emitters, while maintaining accurate alignment with real-world coordinates.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If complex hardware setups with fixed emitters are used to determine pose, then measurement precision is improved, but device complexity increases
Solution Approach 1:
The patent uses image processing to create a virtual copy of the real-world environment with salient points and descriptors. Instead of using physical fixed emitters, the system captures images, extracts salient points, and maintains a map of these points to determine pose. This virtual copying approach eliminates the need for complex physical emitter hardware while maintaining measurement precision.
Solution Approach 2:
The patent replaces mechanical/optical emitter systems with computational image processing. Instead of using physical emitters that require complex alignment and hardware, the system uses imaging devices to capture visual information, extracts salient points, and computes pose through algorithmic processing of image data and descriptor matching.
2Measurement precision
If fixed emitters are used for pose determination, then measurement precision is improved, but ease of operation deteriorates
Solution Approach 1:
The system automatically extracts salient points from images and builds a map of the environment without requiring manual configuration of emitters. The pose determination is performed self-service through automated image processing, salient point extraction, and descriptor matching, eliminating the need for user setup or calibration of physical emitter systems.
Solution Approach 2:
By creating a virtual map copied from captured images, the system eliminates the need for physical emitter setup. The virtual environment model is automatically generated from visual data, making the system easier to operate while maintaining precision, as no physical hardware alignment is required.
3Productivity
If patch-based frame-to-frame tracking is used, then processing requirements are reduced, but measurement precision may worsen
Solution Approach 1:
The patent segments the image processing into distinct functional components: extracting salient points from current images, matching them with the virtual map, and computing pose. This segmentation allows efficient processing by handling only the essential features (salient points) rather than processing entire images, while maintaining precision through focused feature matching.
Solution Approach 2:
The system focuses computational resources on local salient points rather than processing the entire image uniformly. By identifying and processing only the most distinctive features (salient points with unique descriptors), the system achieves efficient processing while maintaining high measurement precision through localized feature analysis.
4Adaptability or versatility
If descriptor-based map-to-frame tracking is used, then adaptability is improved, but processing requirements increase
Solution Approach 1:
The system performs preliminary action by pre-processing images to extract salient points and their descriptors, then storing this information in a virtual map. This preparation allows rapid pose determination in subsequent frames by simply matching current image features against the pre-built map, improving adaptability to different environments while managing processing requirements through efficient data structures.
Data Source
Figure 1
Figure 2
Figure 3A~3C
AI summary
To determine the head pose of a user, a head-mounted display system having an imaging device can obtain a current image of a real-world environment, with points corresponding to salient points which will be used to determine the head pose. The salient points are patch-based and include: a first salient point being projected onto the current image from a previous image, and with a second salient point included in the current image being extracted from the current image. Each salient point is subsequently matched with real-world points based on descriptor-based map information indicating locations of salient points in the real‑world environment. The orientation of the imaging devices is determined based on the matching and based on the relative positions of the salient points in the view captured in the current image. The orientation may be used to extrapolate the head pose of the wearer of the head-mounted display system.