Neural Camera Map Building from RGB Video Without LiDAR

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current autonomous vehicle and robotics systems face challenges in accurately reconstructing 3D maps of environments using camera data due to limited perspective and reliance on expensive LiDAR sensors, which are bulky and costly.

Innovation Solution

A neural camera model is employed to learn pixel-wise ray surfaces for self-supervised depth and pose estimation across various camera geometries, enabling consistent map construction from RGB video sequences without the need for calibrated camera models or expensive sensors.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If LiDAR sensors are used for accurate 3D map construction, then measurement precision is improved, but device complexity and cost increase

Engineering Contradiction:
Improvedepth estimation accuracyVSAvoidsensor system complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent replaces complex mechanical LiDAR sensing systems with a neural network-based computational approach that processes standard camera images. The neural camera model learns to estimate depth and ray surfaces directly from 2D image data, substituting physical light measurement devices with algorithmic processing of optical data.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Solution Approach 2:

The patent creates a virtual 3D representation (digital copy) of the environment through neural network processing of 2D images. Instead of directly measuring 3D space with LiDAR, the system generates a computational copy of the scene geometry that can be used for navigation and mapping purposes.

Inventive Principle:
Principle #26Copying

2Measurement precision

If multiple sensors are used to improve environmental perception, then measurement precision is improved, but device complexity increases

Engineering Contradiction:
Improveenvironmental perception accuracyVSAvoidsensor configuration complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent makes the standard camera serve multiple functions: it captures both 2D visual information for object recognition and 3D geometric information for depth estimation and mapping. The neural camera model enables the same sensor to provide both semantic and metric information, eliminating the need for separate LiDAR hardware.

Inventive Principle:
Principle #6Universality (Multi-functionality)

3Manufacturing precision

If calibrated camera models are used for 3D reconstruction, then manufacturing precision is improved, but ease of manufacture deteriorates

Engineering Contradiction:
Improvemap construction accuracyVSAvoidsystem setup complexity
Core Design Contradiction:
Manufacturing precisionVSEase of manufacture

Solution Approach 1:

The patent implements self-calibration where the neural network automatically learns and adapts to the specific camera's optical characteristics during operation. The system performs its own calibration by learning the relationship between 2D image coordinates and 3D ray surfaces directly from data, eliminating the need for manual calibration procedures.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The patent transforms the rigid, pre-determined parameters of traditional camera calibration into learnable, adaptive parameters. Instead of fixed calibration matrices obtained through complex procedures, the system uses neural network weights that automatically adjust to capture the camera's projection characteristics, making the system more tolerant of manufacturing variations.

Inventive Principle:
Principle #35Parameter changes

Data Source

PatentUS11615544B2Systems and methods for end-to-end map building from a video sequence using neural camera models
Publication Date: 2023.03.28 TOYOTA JIDOSHA KK
  • US11615544B2 patent drawing
  • US11615544B2 patent drawing
  • US11615544B2 patent drawing

AI summary

Systems and methods for map construction using a video sequence captured on a camera of a vehicle in an environment, comprising: receiving a video sequence from the camera, the video sequence including a plurality of image frames capturing a scene of the environment of the vehicle; using a neural camera model to predict a depth map and a ray surface for the plurality of image frames in the received video sequence; and constructing a map of the scene of the environment based on image data captured in the plurality of frames and depth information in the predicted depth maps.