Monocular Ground-Plane Extraction Using Self-Supervised Depth Maps

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing machine vision systems require multiple cameras or expensive LiDAR systems to extract three-dimensional ground plane information from images, which is costly and computationally intensive.

Innovation Solution

The use of self-supervised depth networks to generate three-dimensional reconstructions from monocular images, allowing for the calculation of surface normals and extraction of ground plane information without the need for additional hardware.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If multiple cameras or LiDAR systems are used to extract ground plane information, then measurement precision is improved, but device complexity and cost increase

Engineering Contradiction:
Improveground plane extraction accuracyVSAvoidhardware complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The system uses the monocular camera's own image data to generate depth information through self-supervised learning, eliminating the need for external depth sensors. The depth network processes the single camera input to produce depth maps and surface normals that enable ground plane extraction, making the system self-sufficient without additional hardware

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The patent replaces complex mechanical/optical depth sensing systems (multiple cameras, LiDAR) with a computational approach using a trained depth network. The neural network substitutes physical depth measurement mechanisms with algorithmic depth estimation, achieving comparable accuracy with simpler hardware

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

2Measurement precision

If multiple cameras or LiDAR systems are used to extract ground plane information, then measurement precision is improved, but cost increases

Engineering Contradiction:
Improveground plane extraction accuracyVSAvoidsystem cost
Core Design Contradiction:
Measurement precisionVSEase of manufacture

Solution Approach 1:

The system replaces expensive, durable sensors (LiDAR, stereo cameras) with a single inexpensive monocular camera. The cost savings from using a low-cost camera are offset by the computational resources required for the depth network, but the overall system cost is significantly reduced compared to hardware-based depth sensing solutions

Inventive Principle:
Principle #27Cheap short-living objects (Disposable)

3Measurement precision

If traditional depth systems are used to generate depth maps, then measurement precision is improved, but use of energy increases

Engineering Contradiction:
Improvedepth map accuracyVSAvoidcomputational energy consumption
Core Design Contradiction:
Measurement precisionVSUse of energy by moving object

Solution Approach 1:

The depth network is trained offline in advance on large datasets, performing the computationally intensive learning process beforehand. During actual operation, the pre-trained network efficiently processes incoming images to generate depth maps, reducing real-time computational energy requirements while maintaining high accuracy

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS12307695B2System and method for self-supervised monocular ground-plane extraction
Publication Date: 2025.05.20 TOYOTA JIDOSHA KK
  • US12307695B2 patent drawing
  • US12307695B2 patent drawing
  • US12307695B2 patent drawing

AI summary

A method for controlling an agent to navigate through an environment includes generating a depth map associated with a monocular image of the environment. The method also includes generating a group of surface normal. Each surface normal of the group of surface normals is associated with a respective polygon of a group of polygons associated with the depth map. The method further includes identifying one or more ground planes in the depth map based on the group of surface normal. The method further includes controlling the agent to navigate through the environment based on identifying the one or more ground planes.