Keyframe Selection for AR Map Initialization
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing augmented reality systems face challenges in accurately and efficiently selecting keyframes for map generation and camera position determination due to requirements for human interaction, computational expense, and assumptions about feature distances, leading to suboptimal map accuracy and complexity.
Innovation Solution
A method for selecting a first image from a plurality of images based on determining image features, camera poses, and reconstruction errors, where the image is chosen if the reconstruction error meets a predetermined criterion for scene reconstruction, allowing for automated keyframe selection and improved map generation.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Ease of operation
If five-point-pose algorithm is used for map initialisation, then relative camera pose can be estimated, but human interaction is required and stereo baseline requirement must be understood by users
Solution Approach 1:
The system automatically selects keyframes and determines camera poses without requiring user intervention. The automated keyframe selection algorithm evaluates multiple images and selects appropriate keyframes based on pre-determined criteria, eliminating the need for users to manually choose images or understand stereo baseline requirements.
Solution Approach 2:
The patent changes the selection criterion from manual user input to automated evaluation based on reconstruction error and pose accuracy. The system evaluates camera poses and selects keyframes that satisfy pre-determined criteria, transforming the initialisation process from human-dependent to system-dependent.
2Reliability
If five-point-pose algorithm is used, then map initialisation can be performed, but long uninterrupted tracked features are required which may fail with camera rotation
Solution Approach 1:
The patent segments the feature tracking process into discrete keyframe selections rather than requiring continuous tracking. By selecting multiple keyframes at different positions and using them individually for pose estimation, the system avoids the need for long uninterrupted tracked features and can handle camera rotation more effectively.
3Measurement precision
If model-based method with GRIC score is used, then 3D camera poses can be estimated, but computation of re-projection errors for both homography and epi-polar models is required
Solution Approach 1:
The patent extracts and uses only the essential information needed for keyframe selection - specifically reconstruction error and pose accuracy - rather than computing comprehensive re-projection errors for multiple models. This selective approach maintains sufficient accuracy while significantly reducing computational requirements.
Solution Approach 2:
The system performs partial computation by evaluating only the necessary parameters for keyframe selection rather than exhaustively computing all possible re-projection errors for both homography and epi-polar models. This partial action approach is sufficient for the task while being computationally efficient.
4Productivity
If fixed threshold for temporal distance or track length is used, then keyframe selection can be simplified, but accuracy of 3D map generation is compromised
Solution Approach 1:
The patent changes the selection criteria from fixed temporal or length thresholds to dynamic evaluation based on reconstruction error and pose accuracy. This allows the system to adaptively select keyframes that ensure 3D map generation accuracy while maintaining efficient processing speed.
Data Source
AI summary
A method of selecting a first image from a plurality of images for constructing a coordinate system of an augmented reality system. A first image feature in the first image corresponding to the feature of the marker is determined. A second image feature in a second image is determined based on a second pose of a camera, said second image feature having a visual match to the first image feature. A reconstructed position of the feature of the marker in a three-dimensional (3D) space is determined based on positions of the first and second image features, the first and the second camera pose. A reconstruction error is determined based on the reconstructed position of the feature of the marker and a pre-determined position of the marker.


