Monocular Depth Maps Cleaned with LiDAR Point Cloud Fusion

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Monocular camera images used for depth estimation contain noise, including pixels without depth information, multiple depth information for a single pixel, and occlusion, which degrades the performance of monocular depth estimation models.

Innovation Solution

A method and apparatus that utilize a lidar sensor to generate a point cloud, convert it to a camera coordinate system, remove noise using a convex hull algorithm and deep learning-based segmentation, and delete specific regions like the sky or lidar blind spots to create a clean depth map for training monocular depth estimation models.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Device complexity

If a monocular camera image is used for depth estimation, then the system is simple and cost-effective, but the depth map contains noise and inconsistent depth information

Engineering Contradiction:
Improvesystem complexityVSAvoiddepth map quality
Core Design Contradiction:
Device complexityVSManufacturing precision

Solution Approach 1:

The patent combines monocular camera imaging with lidar point cloud data to create a hybrid depth estimation system. The camera captures visual information while the lidar provides accurate depth measurements, and their results are fused to produce a depth map that maintains the simplicity of monocular systems while achieving the precision of multi-sensor systems.

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The patent introduces an intermediary processing system that uses the camera image as a reference to filter and validate lidar point cloud data. This intermediary process removes noise and inconsistent depth information by comparing lidar measurements with corresponding camera image regions, producing a cleaned depth map.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Ease of manufacture

If raw monocular camera images are used for training, then data acquisition is simple, but the model performance degrades due to noise

Engineering Contradiction:
Improvedata acquisition easeVSAvoidmodel performance
Core Design Contradiction:
Ease of manufactureVSReliability

Solution Approach 1:

The patent applies preliminary noise removal and data cleaning processing to the training data before model training begins. By pre-processing the monocular camera images and corresponding depth maps to remove noise and inconsistent information, the training data quality is improved, which directly enhances model performance and prevents learning from corrupted data.

Inventive Principle:
Principle #10Preliminary action

3Device complexity

If depth information is estimated from monocular images alone, then the system is straightforward, but overfitting occurs and verification accuracy is low

Engineering Contradiction:
Improvesystem simplicityVSAvoidverification accuracy
Core Design Contradiction:
Device complexityVSMeasurement precision

Solution Approach 1:

The patent merges monocular camera data with lidar point cloud data to create a more robust training dataset. This combination provides the model with both visual appearance information and accurate depth measurements, enabling the model to generalize better to out-of-sample data and reducing overfitting while maintaining system simplicity.

Inventive Principle:
Principle #5Merging (Combining)

Applied Scientific Principles

This section explains which scientific principles are used to turn an abstract innovation direction into a practical engineering solution.

Function Achieved in This Case

The method improves the performance of monocular depth estimation models by providing a noise-free depth map, ensuring consistent depth information and removing irrelevant data, thereby enhancing model accuracy and preventing overfitting.

Implementation Method 1

a light detection and ranging (lidar) sensor that generates a point cloud corresponding to the monocular camera image

Methodology Applied
Scientific EffectLight detection and ranging (lidar): LIDAR

Data Source

PatentUS12354289B2Apparatus for generating a depth map of a monocular camera image and a method thereof
Publication Date: 2025.07.08 HYUNDAI MOTOR CO LTD
  • US12354289B2 patent drawing
  • US12354289B2 patent drawing
  • US12354289B2 patent drawing

AI summary

An apparatus and a method remove noise from a point cloud and generate a depth map of a monocular camera image. The apparatus includes a camera sensor that photographs a monocular camera image. The apparatus also includes a lidar sensor that generates the point cloud corresponding to the monocular camera image. The apparatus also includes a controller that removes noise from the point cloud and generates the depth map of the monocular camera image based on the point cloud from which the noise is removed.