Monocular 3D Localization via Ground Plane Scale Correction
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Monocular Structure from Motion (SFM) systems face challenges in large-scale autonomous navigation due to scale drift, particularly in real-world autonomous driving scenarios where the ground plane is difficult to estimate accurately from low-textured road surfaces, leading to high translational errors and limited accuracy compared to stereo systems.
Innovation Solution
A real-time computer vision system using a single camera for 3D moving object localization that combines object detection and monocular SFM through ground plane estimation, tracks feature points for 3D orientation, and corrects scale drift using cues from sparse features and dense stereo visual data, employing a data-driven mechanism for cue combination and adaptive observation covariances.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Ease of manufacture
If monocular SFM is used for autonomous navigation, then cost and calibration requirements are reduced, but scale drift occurs leading to low accuracy
Solution Approach 1:
The patent introduces an intermediary ground plane model as a mediator between camera observations and 3D reconstruction. By estimating the ground plane parameters (orientation and position) from image data, the system resolves the scale ambiguity inherent in monocular SFM. The ground plane acts as a reference frame that anchors the 3D reconstruction, enabling accurate measurement of distances and positions without requiring stereo cameras or complex calibration procedures.
2Measurement precision
If ground plane estimation is performed on low-textured road surfaces, then monocular SFM accuracy is improved, but estimation becomes challenging and unreliable
Solution Approach 1:
The patent merges multiple complementary approaches to ground plane estimation: sparse feature-based estimation (using detected road surface features) and dense pixel-based estimation (using pixel correspondence across frames). By combining these methods, the system leverages the strengths of both approaches - sparse features provide geometric constraints while dense matching provides continuous coverage. This fusion enables reliable ground plane estimation even on low-textured surfaces where either method alone would fail.
Solution Approach 2:
The patent employs parameter changes by adapting the estimation method based on scene conditions. When sparse features are available, the system uses feature-based plane fitting; when dense pixel correspondence is available, it uses pixel-based homography estimation. The system dynamically switches between these parameter estimation approaches based on the quality and availability of input data, ensuring robust performance across diverse road surface conditions.
3Device complexity
If two-view estimation is used for relative pose, then computation is simplified, but translational errors increase for narrow baseline forward motion
Solution Approach 1:
The patent transitions from two-view geometry to multi-view geometry by incorporating the ground plane as an additional dimensional constraint. Instead of estimating pose from only two image planes, the system uses three-dimensional information from the ground plane to resolve ambiguities in the translation estimation. This dimensional addition provides extra geometric constraints that disambiguate the scale and direction of motion, significantly improving translation accuracy while maintaining computational efficiency through the structured nature of the ground plane model.
Data Source
AI summary
Systems and methods are disclosed for autonomous driving with only a single camera by moving object localization in 3D with a real-time framework that harnesses object detection and monocular structure from motion (SFM) through the ground plane estimation; tracking feature points on moving cars a real-time framework to and use the feature points for 3D orientation estimation; and correcting scale drift with ground plane estimation that combines cues from sparse features and dense stereo visual data.


