Mixed Reality Tracking Using Dynamic Feature Segmentation
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current mixed reality (MR) technologies face challenges in accurately aligning real and virtual spaces, particularly in obtaining the position and orientation of cameras and tracking target objects, especially when markers are outside the camera's visual field or when advanced preparations like CAD systems or 3D scanners are not used.
Innovation Solution
An information processing apparatus that extracts feature information from images, detects indices on tracking target objects, estimates their position and orientation, and classifies features to construct a tracking target model, allowing for accurate camera and object positioning without requiring markers within the visual field or pre-created models.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If Visual SLAM is used to obtain camera position and orientation, then the position and orientation can be estimated based on feature points, but feature points from moving objects like tools held by users are eliminated, reducing tracking accuracy for dynamic objects
Solution Approach 1:
The patent segments feature points into two categories: those belonging to the background scene and those belonging to moving objects. By processing these segments separately and allowing moving object features to be retained despite motion, the system achieves both accurate camera localization and effective tracking of dynamic objects like tools held by users.
Solution Approach 2:
The patent introduces dynamic adaptability by allowing the set of tracked feature points to change over time. Features from moving objects are dynamically added to the tracking list when detected, enabling the system to adapt to dynamic environments without requiring pre-defined models or markers.
2Measurement precision
If model-based tracking with pre-created 3D models is used, then position and orientation can be obtained using edge or optical flow information, but advanced preparations like CAD systems or 3D scanners are required, increasing system complexity
Solution Approach 1:
The patent performs preliminary extraction of edge information and optical flow data directly from captured images, eliminating the need for separate 3D modeling processes. By preparing feature data in advance from the actual captured images rather than from pre-created models, the system achieves accurate tracking without requiring CAD systems or 3D scanners.
Solution Approach 2:
The patent creates a simplified representation of the target object using extracted edge information and optical flow patterns from images, rather than using complex pre-created 3D models. This copying approach captures the essential geometric and motion information needed for tracking while avoiding the complexity of traditional 3D modeling pipelines.
3Measurement precision
If markers are used for alignment, then position and orientation can be obtained, but markers must be within the camera's visual field, limiting tracking capability when markers are outside the view
Solution Approach 1:
The patent uses edge information and optical flow patterns as intermediary representations of the target object's position and motion. These intermediaries can be extracted from image data even when the actual marker or object is at the boundary or outside the optimal viewing range, extending the effective tracking distance beyond what traditional marker-based systems allow.
Solution Approach 2:
The patent transitions from two-dimensional marker detection to three-dimensional spatial reasoning by using optical flow information that captures motion across the image plane. This dimensional expansion allows the system to infer the position and orientation of objects even when they are partially outside the camera's direct line of sight or at the edges of the visual field.
Data Source
AI summary
An apparatus includes an extraction unit configured to extract a plurality of feature information from an image obtained by capturing a real space including a tracking target object, an index estimation unit configured to detect an index arranged on the tracking target object from the image, and estimate a position and orientation of the index, a target object estimation unit configured to estimate a position and orientation of the tracking target object based on the position and orientation of the index and a tracking target model, a classification unit configured to classify the plurality of feature information based on a position and orientation of a camera capturing the real space and the position and orientation of the tracking target object, and a construction unit configured to add feature information determined as belonging to the tracking target object by the classification unit, to the tracking target model.


