Vehicle Depth Estimation Using Multi-Camera Panoramic Loss

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing methods for obtaining accurate depth data in vehicles, such as LiDAR and stereo cameras, incur significant hardware and maintenance costs and require additional processing to convert sparse measurements into dense data, while also necessitating additional sensors.

Innovation Solution

A Deep Neural Network (DNN) with an encoder-decoder architecture is trained using a multi-task learning loss function that combines disparity, smoothness, semantic segmentation, and panoramic loss terms to estimate dense depth maps from surround view images captured by multiple cameras, generating dense depth data without the need for additional hardware beyond existing cameras.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If LiDAR is added to obtain accurate depth data, then measurement precision is improved, but device complexity and cost increase

Engineering Contradiction:
Improvedepth data accuracyVSAvoidhardware complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent creates a virtual copy of the LiDAR depth measurement capability through software processing. By training a neural network on paired LiDAR and image data, the system learns to generate LiDAR-equivalent depth maps from standard camera images, effectively copying the depth sensing function without physical LiDAR hardware.

Inventive Principle:
Principle #26Copying

Solution Approach 2:

The patent replaces the mechanical LiDAR system (which uses rotating mirrors or fixed arrays to emit and detect light) with a computational approach using deep neural networks. The mechanical depth sensing system is substituted with an algorithmic system that processes image data to infer depth information.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

2Measurement precision

If LiDAR is used to obtain depth data, then measurement precision is improved, but use of energy increases

Engineering Contradiction:
Improvedepth data accuracyVSAvoidpower consumption
Core Design Contradiction:
Measurement precisionVSUse of energy by moving object

Solution Approach 1:

The patent creates a virtual copy of the LiDAR depth measurement capability through software processing. By training a neural network on paired LiDAR and image data, the system learns to generate LiDAR-equivalent depth maps from standard camera images, effectively copying the depth sensing function without physical LiDAR hardware.

Inventive Principle:
Principle #26Copying

Solution Approach 2:

The patent replaces the mechanical LiDAR system (which uses rotating mirrors or fixed arrays to emit and detect light) with a computational approach using deep neural networks. The mechanical depth sensing system is substituted with an algorithmic system that processes image data to infer depth information.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

3Measurement precision

If LiDAR is used to obtain depth data, then measurement precision is improved, but loss of substance increases due to sparse measurements requiring additional processing

Engineering Contradiction:
Improvedepth data accuracyVSAvoiddata sparsity
Core Design Contradiction:
Measurement precisionVSLoss of substance

Solution Approach 1:

The patent creates a virtual copy of the LiDAR depth measurement capability through software processing. By training a neural network on paired LiDAR and image data, the system learns to generate LiDAR-equivalent depth maps from standard camera images, effectively copying the depth sensing function without physical LiDAR hardware.

Inventive Principle:
Principle #26Copying

Solution Approach 2:

The patent merges the strengths of LiDAR (accurate depth measurement) with the advantages of camera systems (dense pixel data and lower cost). By combining LiDAR ground truth data during training with image data, the neural network learns to produce dense depth maps that capture both accurate depth information and complete spatial coverage.

Inventive Principle:
Principle #5Merging (Combining)

4Measurement precision

If stereo cameras are added to obtain depth data, then measurement precision is improved, but device complexity increases

Engineering Contradiction:
Improvedepth data accuracyVSAvoidsensor quantity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent makes the existing camera system multi-functional by enabling it to perform both standard image capture and depth estimation. Through neural network processing, the same camera hardware that captures visual scenes also generates depth maps, eliminating the need for dedicated depth-sensing sensors.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The patent creates a virtual copy of the LiDAR depth measurement capability through software processing. By training a neural network on paired LiDAR and image data, the system learns to generate LiDAR-equivalent depth maps from standard camera images, effectively copying the depth sensing function without physical LiDAR hardware.

Inventive Principle:
Principle #26Copying

Data Source

PatentUS12008817B2Systems and methods for depth estimation in a vehicle
Publication Date: 2024.06.11 GM GLOBAL TECHNOLOGY OPERATIONS LLC
  • US12008817B2 patent drawing
  • US12008817B2 patent drawing
  • US12008817B2 patent drawing

AI summary

Methods and system for training a neural network for depth estimation in a vehicle. The methods and systems receive respective training image data from at least two cameras. Fields of view of adjacent cameras of the at least two cameras partially overlap. The respective training image data is processed through a neural network providing depth data and semantic segmentation data as outputs. The neural network is trained based on a loss function. The loss function combines a plurality of loss terms including at least a semantic segmentation loss term and a panoramic loss term. The panoramic loss term includes a similarity measure regarding overlapping image patches of the respective image data that each correspond to a region of overlapping fields of view of the adjacent cameras. The semantic segmentation loss term quantifies a difference between ground truth semantic segmentation data and the semantic segmentation data output from the neural network.