EnforceNet Monocular Camera Localization in Sparse LiDAR Point Clouds

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current localization methods for autonomous robots, such as autonomous cars, face challenges with high-definition GPS costs and availability, and camera-based localization stability under varying lighting conditions and scale drift, especially in environments like parking garages where LiDAR scans are sparse and expensive.

Innovation Solution

A novel neural network structure, EnforceNet, that combines camera images with LiDAR point clouds to estimate camera pose using depth projections, incorporating a resistor module for state-value prediction and pose regression, enabling efficient localization even in sparse LiDAR environments and varying lighting conditions.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Ease of manufacture

If camera-based localization is used, then cost is reduced and ubiquity is improved, but stability under varying lighting conditions deteriorates

Engineering Contradiction:
ImprovecostVSAvoidstability under varying lighting conditions
Core Design Contradiction:
Ease of manufactureVSReliability

Solution Approach 1:

The patent merges camera-based visual odometry with LiDAR point cloud data processing. The system combines camera images for feature extraction with LiDAR depth information for geometric constraints, creating a hybrid localization approach that leverages the strengths of both modalities while mitigating their individual weaknesses in terms of lighting sensitivity and computational efficiency

Inventive Principle:
Principle #5Merging (Combining)

2Adaptability or versatility

If visual odometry is used, then localization can be achieved without GPS, but accuracy deteriorates due to scale drift

Engineering Contradiction:
Improvelocalization without GPSVSAvoidlocalization accuracy
Core Design Contradiction:
Adaptability or versatilityVSMeasurement precision

Solution Approach 1:

The patent introduces LiDAR point cloud depth projections as an intermediary element that bridges camera-based visual odometry and GPS-like precision. The depth projections from LiDAR serve as geometric anchors that constrain the scale drift accumulation in visual odometry, providing metric accuracy without requiring GPS signals

Inventive Principle:
Principle #24Intermediary (Mediator)

3Measurement precision

If LiDAR localization is used, then localization accuracy is improved, but computational resources required increase significantly

Engineering Contradiction:
Improvelocalization accuracyVSAvoidcomputational resources
Core Design Contradiction:
Measurement precisionVSUse of energy by moving object

Solution Approach 1:

The patent segments the LiDAR point cloud into depth projections at multiple levels or regions, processing only the relevant portions rather than the entire point cloud. This segmentation approach reduces the computational burden while maintaining localization accuracy by focusing calculations on the most informative depth layers

Inventive Principle:
Principle #1Segmentation

4Reliability

If depth projections are sampled from LiDAR point cloud, then localization stability is improved, but processing time increases

Engineering Contradiction:
Improvelocalization stabilityVSAvoidprocessing time
Core Design Contradiction:
ReliabilityVSLoss of time

Solution Approach 1:

The patent performs preliminary sampling of depth projections from the LiDAR point cloud before the main localization computation. By pre-processing and selecting representative depth projections in advance, the system reduces the computational time required during actual localization while maintaining the stability benefits of depth-based constraints

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS11380003B2Monocular camera localization in large scale indoor sparse LiDAR point cloud
Publication Date: 2022.07.05 BLACK SESAME TECH INC
  • US11380003B2 patent drawing
  • US11380003B2 patent drawing
  • US11380003B2 patent drawing

AI summary

A method of camera localization, comprising, receiving a camera image, receiving a LiDAR point cloud, estimating an initial camera pose for the camera image, sampling an initial set of depth projections within the LiDAR point cloud, measuring a similarity of the initial camera pose to the initial set of depth projections and deriving a subsequent set of depth projections based on the measured similarity.