Trinocular Camera Pose Tracking via Multi-Frame Feature Matching
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Camera tracking methods using monocular cameras face challenges in accurately recovering 3D information and camera motion due to accumulation errors, especially in large-scale scenes, and struggle with occlusions when using stereo cameras.
Innovation Solution
A method and apparatus for tracking camera pose using at least three cameras, which involves extracting and tracking features across multiple frames, removing dynamic trajectories, and estimating camera pose based on scale-invariant feature transform (SIFT) descriptors and geometry constraints, employing a two-phase processing approach to balance accuracy and speed.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Device complexity
If monocular cameras are used for camera tracking, then device complexity is reduced and cost is lowered, but measurement precision of 3D information and camera motion deteriorates due to accumulation errors
Solution Approach 1:
The patent combines multiple monocular cameras (at least three) to form a multi-camera system that merges their individual tracking results. This combination allows the system to recover 3D information and camera motion more accurately by integrating data from multiple viewpoints, thereby resolving the precision limitation of single monocular cameras while maintaining relative simplicity compared to stereo systems.
Solution Approach 2:
The patent transitions from 2D image plane tracking to 3D space tracking by utilizing depth information from multiple camera viewpoints. By incorporating the temporal dimension (multiple frames) and spatial dimension (multiple camera positions), the system reconstructs 3D camera trajectories and scene geometry, overcoming the fundamental 2D-to-3D ambiguity of monocular tracking.
2Measurement precision
If stereo cameras are used to recover camera motion and depth maps, then measurement precision of 3D information is improved, but difficulty of handling occlusions increases
Solution Approach 1:
The patent merges observations from at least three camera viewpoints to handle occlusions. When a point is occluded in one camera's view, the system can still track it using corresponding points from other cameras' views. This multi-view integration provides redundant information that compensates for occlusions, making the tracking more robust compared to stereo camera systems.
Solution Approach 2:
The patent applies different processing strategies to different regions of the image based on occlusion status. For occluded regions, the system relies on temporal consistency and multi-view correspondence; for non-occluded regions, it uses direct depth estimation. This localized adaptation optimizes performance across different scene conditions.
3Reliability
If features are tracked across multiple frames to improve tracking stability, then reliability of camera pose estimation is improved, but loss of time increases due to processing multiple frames
Solution Approach 1:
The patent segments the feature tracking process into distinct phases: feature extraction from multiple frames, feature matching across frames, and pose estimation. By organizing the processing in a structured pipeline, the system efficiently handles multiple frames without excessive computational overhead, balancing reliability with processing speed.
Solution Approach 2:
The patent performs preliminary feature extraction and matching across multiple frames before final pose estimation. By pre-processing and identifying corresponding features in advance, the system reduces the computational burden during the actual tracking phase, thereby maintaining reliability while minimizing time loss.
Data Source
AI summary
A camera pose tracking apparatus may track a camera pose based on frames photographed using at least three cameras, may extract and track at least one first feature in multiple-frames, and may track a pose of each camera in each of the multiple-frames based on first features. When the first features are tracked in the multiple-frames, the camera pose tracking apparatus may track each camera pose in each of at least one single-frame based on at least one second feature of each of the at least one single-frame. Each of the at least one second feature may correspond to one of the at least one first feature, and each of the at least one single-frame may be a previous frame of an initial frame of which the number of tracked second features is less than a threshold, among frames consecutive to multiple-frames.


