3D Vision-Language Driving Plans From Multi-View Spatial Alignment
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional vision-language models (VLMs) are typically trained to operate using 2D video data and, therefore, cannot perceive the full three-dimensional (3D) volume of space surrounding a given AV, leading to inaccurate assessment of distances, sizes, and relative positions of objects within the environment, and therefore cannot effectively make safe driving decisions.
Innovation Solution
A computer-implemented method for controlling a vehicle using a vision-language model (VLM) trained to interpret three-dimensional (3D) data, including a projector that is configured to process multi-view image features and a 3D position encoding to generate aligned image features, and a projector that is configured to generate aligned image features, and a projector that is configured to generate aligned image features, and a projector that is configured to process multi-view image features and a 3D position encoding to generate aligned image features.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If conventional VLMs are trained to operate using 2D video data, then the model structure and training process are simpler, but the vehicle cannot perceive the full 3D volume of space surrounding it, leading to inaccurate assessment of distances, sizes, and relative positions of objects
Solution Approach 1:
The patent transitions from 2D video data processing to 3D spatial understanding by introducing a projector that maps 3D point cloud data into a 2D image space. This dimensionality transformation allows the VLM to leverage its existing 2D processing capabilities while incorporating 3D spatial information from LiDAR sensors, thereby improving measurement precision without requiring complete redesign of the model architecture.
2Reliability
If conventional VLMs process only 2D video data, then the training data requirements are lower, but the system cannot effectively make safe driving decisions due to limited environmental perception
Solution Approach 1:
The patent merges 2D video data from cameras with 3D point cloud data from LiDAR sensors into a unified representation. The projector aligns the 3D point cloud with the 2D image space, creating combined training samples that contain both visual and spatial information. This merging approach enables the VLM to learn from diverse data sources simultaneously, improving driving decision reliability while efficiently utilizing available training data.
3Loss of information
If the VLM is enhanced to interpret 3D data, then the environmental assessment depth and accuracy improve, but the computational complexity and processing requirements increase
Solution Approach 1:
The patent introduces a projector as an intermediary component that bridges 3D point cloud data and 2D image processing. Instead of directly processing 3D data through the entire VLM pipeline, the projector transforms and aligns the 3D point cloud into the 2D image space, where it can be processed by the existing VLM architecture. This intermediary approach minimizes computational overhead while maximizing information completeness.
Data Source
AI summary
In various embodiments, a computer-implemented method for controlling a vehicle includes performing a visual-language alignment operation based on a set of multi-view image features and a three-dimensional position encoding to generate a set of aligned image features, causing a language model to generate a driving plan for operating the vehicle based on the set of aligned image features, wherein the driving plan includes a description of a three-dimensional trajectory for the vehicle; and controlling the vehicle to move based on the driving plan.


