Monocular Ground-Plane Extraction Using Self-Supervised Depth Maps
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing machine vision systems require multiple cameras or expensive LiDAR systems to extract three-dimensional ground plane information from images, which is costly and computationally intensive.
Innovation Solution
The use of self-supervised depth networks to generate three-dimensional reconstructions from monocular images, allowing for the calculation of surface normals and extraction of ground plane information without the need for additional hardware.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If multiple cameras or LiDAR systems are used to extract ground plane information, then measurement precision is improved, but device complexity and cost increase
Solution Approach 1:
The system uses the monocular camera's own image data to generate depth information through self-supervised learning, eliminating the need for external depth sensors. The depth network processes the single camera input to produce depth maps and surface normals that enable ground plane extraction, making the system self-sufficient without additional hardware
Solution Approach 2:
The patent replaces complex mechanical/optical depth sensing systems (multiple cameras, LiDAR) with a computational approach using a trained depth network. The neural network substitutes physical depth measurement mechanisms with algorithmic depth estimation, achieving comparable accuracy with simpler hardware
2Measurement precision
If multiple cameras or LiDAR systems are used to extract ground plane information, then measurement precision is improved, but cost increases
Solution Approach 1:
The system replaces expensive, durable sensors (LiDAR, stereo cameras) with a single inexpensive monocular camera. The cost savings from using a low-cost camera are offset by the computational resources required for the depth network, but the overall system cost is significantly reduced compared to hardware-based depth sensing solutions
3Measurement precision
If traditional depth systems are used to generate depth maps, then measurement precision is improved, but use of energy increases
Solution Approach 1:
The depth network is trained offline in advance on large datasets, performing the computationally intensive learning process beforehand. During actual operation, the pre-trained network efficiently processes incoming images to generate depth maps, reducing real-time computational energy requirements while maintaining high accuracy
Data Source
AI summary
A method for controlling an agent to navigate through an environment includes generating a depth map associated with a monocular image of the environment. The method also includes generating a group of surface normal. Each surface normal of the group of surface normals is associated with a respective polygon of a group of polygons associated with the depth map. The method further includes identifying one or more ground planes in the depth map based on the group of surface normal. The method further includes controlling the agent to navigate through the environment based on identifying the one or more ground planes.


