Small Baseline Stereo Camera Depth Estimation via LiDAR Transfer Learning

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Stereo cameras with small baselines face challenges in achieving high depth resolution due to structural constraints, such as in smartphones, wearable AR/VR devices, and drones, where existing methods fail to produce accurate depth maps.

Innovation Solution

A deep learning-based method that uses transfer learning from a wide baseline-stereo network to estimate high-resolution depth maps, leveraging a LiDAR sensor for training and generating pseudo-LiDAR data by down-sampling the depth map.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If a wider baseline is used in stereo camera, then depth resolution is improved, but device size and structural complexity increase

Engineering Contradiction:
Improvedepth resolutionVSAvoidbaseline distance
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent replaces the mechanical approach of increasing physical baseline distance with a computational approach using deep learning networks. Instead of physically separating cameras wider apart, the system uses neural networks to enhance depth estimation accuracy from the limited baseline available in mobile devices, thus resolving the contradiction between measurement precision and device complexity.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Solution Approach 2:

The patent changes the approach from physical parameter (baseline distance) to computational parameter (deep learning model parameters). By training networks on synthetic data with varying baselines and using transfer learning, the system adapts to different baseline configurations without requiring physical changes to the camera hardware, thereby improving depth resolution without increasing device complexity.

Inventive Principle:
Principle #35Parameter changes

2Measurement precision

If deep learning network is trained with LiDAR data, then depth estimation accuracy is improved, but data generation complexity increases

Engineering Contradiction:
Improvedepth estimation accuracyVSAvoiddata generation process
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent creates synthetic training data by copying and rendering 3D scenes using computer graphics. Instead of requiring actual LiDAR scans and complex data collection pipelines, the system generates realistic synthetic image pairs and corresponding depth maps from 3D models, thereby improving training accuracy while simplifying the data generation process.

Inventive Principle:
Principle #26Copying

Solution Approach 2:

The patent performs preliminary training of deep learning networks on synthetic data before deploying to real-world applications. By pre-training on easily generated synthetic datasets with ground truth depth information, the system avoids the need for complex real-world data collection and annotation processes, thus improving accuracy without increasing operational complexity.

Inventive Principle:
Principle #10Preliminary action

3Measurement precision

If transfer learning from wide baseline network is applied, then depth resolution in small baseline is improved, but computational resources and training time increase

Engineering Contradiction:
Improvedepth resolutionVSAvoidcomputational resources
Core Design Contradiction:
Measurement precisionVSUse of energy by moving object

Solution Approach 1:

The patent performs preliminary training of a wide baseline deep learning network on synthetic data, then transfers the learned features to a small baseline network. This pre-training phase completes the heavy computational work once, allowing the final small baseline model to run more efficiently on mobile devices, thus improving depth resolution while managing computational resource consumption.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent segments the deep learning network into shared feature extraction layers and task-specific layers. By segmenting the architecture and using transfer learning, the system reuses computational resources from the wide baseline network in the small baseline network, reducing redundant computations and lowering overall energy consumption while maintaining high depth resolution.

Inventive Principle:
Principle #1Segmentation

Data Source

PatentUS20240202951A1Depth estimation method for small baseline-stereo camera through lidar sensor fusion
Publication Date: 2024.06.20 KOREA ELECTRONICS TECH INST
  • US20240202951A1 patent drawing
  • US20240202951A1 patent drawing
  • US20240202951A1 patent drawing

AI summary

There is provided a depth estimation method for a small baseline-stereo camera through LiDAR sensor fusion. A depth map estimation method according to an embodiment may estimate a high-resolution depth map from a small baseline-stereo image based on deep learning, by using transfer learning from a deep learning network that is trained to estimate a depth map from a wide baseline-stereo image. Accordingly, in a device which has a small baseline-stereo camera installed therein due to structural constraints, such as a smartphone, a wearable AR/VR device, a drone, 3D image quality can be enhanced. In addition, according to embodiments, pseudo-LiDAR data may be generated by using a depth map estimated from a small baseline-stereo image, and may be used for replacing or reinforcing LiDAR data.