AR Pose Determination via Silhouette Mutual Information
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing augmented reality systems face challenges in accurately determining the pose of a 3D model relative to a video flux of a real object, especially when the 3D model has no texture or when the real object is difficult to segment, leading to unreliable pose computation.
Innovation Solution
A computer-implemented method that captures video flux, extracts 2D images of the real object, and determines the pose of the 3D model by rewarding mutual information between virtual and actual images, allowing for accurate augmentation without relying on texture correlation or pre-trained neural networks.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If key point matching methods (FERN or SIFT) are used for pose computation, then pose determination can be performed when texture is available, but the method fails when the 3D model has no texture or when the real object is hard to segment
Solution Approach 1:
The patent changes the fundamental parameter used for pose determination from texture-based key point matching to geometry-based silhouette matching. By using the outline/silhouette of the object in the image and comparing it with the projected silhouette of the 3D model, the method becomes applicable to untextured objects and those that are difficult to segment, while maintaining reliability through geometric constraints
Solution Approach 2:
The patent segments the problem into two independent components: (1) extracting the silhouette/outline of the object from the image, and (2) projecting the 3D model silhouette based on candidate poses. This segmentation allows each component to be optimized independently and combined through silhouette matching, resolving the limitation of requiring both texture and segmentability
2Measurement precision
If learning-based approaches with pre-trained neural networks are used, then pose prediction can be achieved for trained objects, but the method does not work well on real objects that were unseen during training
Solution Approach 1:
Instead of using pre-trained neural networks that require copying training data patterns, the patent uses a geometric copying approach where the 3D model's silhouette is projected and directly compared with the observed object silhouette. This eliminates the need for training on specific objects and enables immediate generalization to any object with a known 3D model
Solution Approach 2:
The patent inverts the traditional approach by instead of training the system to recognize objects, it projects the known 3D model in multiple candidate poses and determines which projection best matches the observed silhouette. This inversion eliminates the generalization problem of neural networks on unseen objects
3Productivity
If EPnP algorithm is used for pose computation from key points, then computation is efficient O(n), but the method requires accurate key point correspondence which is unavailable when texture is missing
Solution Approach 1:
The patent extracts only the essential geometric feature (silhouette/outline) from the image, removing the dependency on texture information and key point detection. This extraction approach maintains computational efficiency while improving reliability by focusing on the most robust geometric invariant - the object's silhouette
Data Source
AI summary
A computer-implemented method of augmented reality includes capturing the video flux with a video camera, extracting, from the video flux, one or more 2D images each representing the real object, and obtaining a 3D model representing the real object. The method also includes determining a pose of the 3D model relative to the video flux, among candidate poses. The determining rewards a mutual information, for at least one 2D image and for each given candidate pose, which represents a mutual dependence between a virtual 2D rendering and the at least one 2D image. The method also includes augmenting the video flux based on the pose. This forms an improved solution of augmented reality for augmenting a video flux of a real scene including a real object.


