3D Vision-Language Driving Plans From Multi-View Spatial Alignment

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional vision-language models (VLMs) are typically trained to operate using 2D video data and, therefore, cannot perceive the full three-dimensional (3D) volume of space surrounding a given AV, leading to inaccurate assessment of distances, sizes, and relative positions of objects within the environment, and therefore cannot effectively make safe driving decisions.

Innovation Solution

A computer-implemented method for controlling a vehicle using a vision-language model (VLM) trained to interpret three-dimensional (3D) data, including a projector that is configured to process multi-view image features and a 3D position encoding to generate aligned image features, and a projector that is configured to generate aligned image features, and a projector that is configured to generate aligned image features, and a projector that is configured to process multi-view image features and a 3D position encoding to generate aligned image features.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If conventional VLMs are trained to operate using 2D video data, then the model structure and training process are simpler, but the vehicle cannot perceive the full 3D volume of space surrounding it, leading to inaccurate assessment of distances, sizes, and relative positions of objects

Engineering Contradiction:
Improveassessment accuracy of distances, sizes, and relative positionsVSAvoidmodel complexity for processing 3D data
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent transitions from 2D video data processing to 3D spatial understanding by introducing a projector that maps 3D point cloud data into a 2D image space. This dimensionality transformation allows the VLM to leverage its existing 2D processing capabilities while incorporating 3D spatial information from LiDAR sensors, thereby improving measurement precision without requiring complete redesign of the model architecture.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

2Reliability

If conventional VLMs process only 2D video data, then the training data requirements are lower, but the system cannot effectively make safe driving decisions due to limited environmental perception

Engineering Contradiction:
Improvesafety of driving decisionsVSAvoidtraining data volume and annotation requirements
Core Design Contradiction:
ReliabilityVSQuantity of substance

Solution Approach 1:

The patent merges 2D video data from cameras with 3D point cloud data from LiDAR sensors into a unified representation. The projector aligns the 3D point cloud with the 2D image space, creating combined training samples that contain both visual and spatial information. This merging approach enables the VLM to learn from diverse data sources simultaneously, improving driving decision reliability while efficiently utilizing available training data.

Inventive Principle:
Principle #5Merging (Combining)

3Loss of information

If the VLM is enhanced to interpret 3D data, then the environmental assessment depth and accuracy improve, but the computational complexity and processing requirements increase

Engineering Contradiction:
Improvecompleteness of environmental informationVSAvoidcomputational energy consumption
Core Design Contradiction:
Loss of informationVSUse of energy by moving object

Solution Approach 1:

The patent introduces a projector as an intermediary component that bridges 3D point cloud data and 2D image processing. Instead of directly processing 3D data through the entire VLM pipeline, the projector transforms and aligns the 3D point cloud into the 2D image space, where it can be processed by the existing VLM architecture. This intermediary approach minimizes computational overhead while maximizing information completeness.

Inventive Principle:
Principle #24Intermediary (Mediator)

Data Source

PatentUS20250381981A1Techniques for autonomous driving with language
Publication Date: 2025.12.18 NVIDIA CORP
  • US20250381981A1 patent drawing
  • US20250381981A1 patent drawing
  • US20250381981A1 patent drawing

AI summary

In various embodiments, a computer-implemented method for controlling a vehicle includes performing a visual-language alignment operation based on a set of multi-view image features and a three-dimensional position encoding to generate a set of aligned image features, causing a language model to generate a driving plan for operating the vehicle based on the set of aligned image features, wherein the driving plan includes a description of a three-dimensional trajectory for the vehicle; and controlling the vehicle to move based on the driving plan.