BEV Mask Generation Using LIDAR Memory Across Map Tiles

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing methods for generating top-down tile representations and BEV masks in autonomous systems rely heavily on perspective image data, leading to inaccuracies due to camera calibration errors, depth estimation inaccuracies, and occlusions, which propagate uncertainties and lack self-correction mechanisms, especially in dynamic environments.

Innovation Solution

Generating BEV masks using top-down LIDAR data and other sensor data, combined with a robust memory mechanism to refine and merge predictions across multiple tile images, enhancing accuracy and precision.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Ease of manufacture

If perspective image data from cameras is used to generate top-down tile representations and BEV masks, then the system can utilize available and affordable camera systems, but the accuracy and precision of environmental feature location predictions deteriorate due to camera calibration errors, depth estimation inaccuracies, and occlusions

Engineering Contradiction:
Improveavailability and affordability of camera systemsVSAvoidaccuracy of environmental feature location predictions
Core Design Contradiction:
Ease of manufactureVSMeasurement precision

Solution Approach 1:

The patent combines multiple data sources including top-down camera images, perspective camera images, and LIDAR data to generate BEV masks. By merging these complementary data sources, the system achieves both cost-effectiveness (using affordable cameras) and high accuracy (through multi-sensor fusion that compensates for individual sensor limitations).

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The patent introduces an intermediary processing framework that transforms perspective images and LIDAR data into BEV space representations. This intermediary BEV mask generation process acts as a mediator that reconciles the different data formats and coordinate systems, enabling accurate feature location prediction while utilizing diverse sensor inputs.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Productivity

If perspective images are projected into top-down views for tile generation, then the system can create map representations using standard camera systems, but uncertainties propagate and accumulate across tiles leading to inconsistent BEV masks

Engineering Contradiction:
Improvecomputational efficiency of map generationVSAvoidconsistency of BEV masks across tiles
Core Design Contradiction:
ProductivityVSReliability

Solution Approach 1:

The patent segments the environment into multiple overlapping tiles, each processed independently to generate local BEV masks. The overlap between adjacent tiles ensures consistent feature detection across boundaries, while the segmentation enables parallel processing that maintains computational efficiency.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent implements a feedback mechanism where detected features from one tile inform and refine the processing of adjacent tiles. This feedback loop ensures that uncertainties do not accumulate uncontrollably, as each tile's processing is guided by validated features from neighboring regions, maintaining overall consistency.

Inventive Principle:
Principle #23Feedback

3Device complexity

If traditional perspective image processing methods are used, then the system can operate with existing camera infrastructure, but the method lacks self-correction mechanisms to refine predictions and reduce errors

Engineering Contradiction:
Improvesimplicity of processing pipelineVSAvoidaccuracy of feature location predictions
Core Design Contradiction:
Device complexityVSMeasurement precision

Solution Approach 1:

The patent implements self-service through iterative refinement processes where the system uses its own output (initial feature detections) to improve its input (refined BEV masks). The memory mechanism stores and reuses detected features across multiple processing passes, allowing the system to self-correct errors without external intervention.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The patent replaces traditional mechanical perspective projection methods with a computational approach using neural networks and probabilistic models. This substitution enables sophisticated error correction and uncertainty quantification that go beyond simple geometric transformations, improving precision while maintaining reasonable complexity.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Data Source

PatentUS20260080692A1Mask generation for feature detection in autonomous and semi-autonomous systems and applications
Publication Date: 2026.03.19 NVIDIA CORP
  • US20260080692A1 patent drawing
  • US20260080692A1 patent drawing
  • US20260080692A1 patent drawing

AI summary

In various examples, systems and methods are described that may be used to generate a mask of a geographic area and corresponding vector representations of one or more environmental features included in the area. In some embodiments, the method and system may generate or obtain a first tile image representing a portion of a path surface corresponding to a geographic area. One or more features may be extracted from the data where the extracted features may indicate environmental characteristics associated with the path surface—e.g., lane boundaries, medians, traffic signs, signals, etc. Additionally, the method or system may generate a mask corresponding to the tile image and vector representations of the portion of the geographic area. In some embodiments, the locations of the environmental features associated with the mask may be reconciled with one or more previously generated masks that include some of the same environmental features.