Vision-Language Planning Models for Context-Aware Driving Trajectories

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing autonomous driving systems face challenges in integrating language comprehension with vision-based planning systems, leading to suboptimal performance and generalization across diverse driving scenarios.

Innovation Solution

A Vision-Language-Planning (VLP) Foundation model is introduced, utilizing contrastive learning to fuse language comprehension with vision models, enhancing planning capabilities by extracting high-level semantic information from textual cues and improving trajectory prediction.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If vision-based planning systems are used in autonomous driving, then the system can process visual information from cameras and sensors, but the system lacks language comprehension capabilities leading to suboptimal performance in diverse driving scenarios

Engineering Contradiction:
Improveperformance across diverse driving scenariosVSAvoidlanguage comprehension capability
Core Design Contradiction:
Adaptability or versatilityVSLoss of information

Solution Approach 1:

The patent merges vision-based planning systems with language comprehension models by integrating a language model into the autonomous driving architecture. The language model processes textual descriptions of driving scenarios and generates natural language plans that are then executed by the vision-based planning system, creating a unified system that leverages both visual and linguistic capabilities for improved adaptability across diverse driving scenarios

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The patent introduces a language model as an intermediary component between the sensory input systems and the planning execution system. This language model acts as a mediator that translates visual scene understanding into natural language representations, which then guide the planning system, thereby bridging the gap between vision processing and decision-making without direct coupling

Inventive Principle:
Principle #24Intermediary (Mediator)

2Measurement precision

If traditional vision models are used for planning, then the system can process visual data, but the system lacks high-level semantic understanding from textual cues

Engineering Contradiction:
Improvetrajectory prediction accuracyVSAvoidsemantic information from text
Core Design Contradiction:
Measurement precisionVSLoss of information

Solution Approach 1:

The patent combines traditional vision models with language models in a unified architecture where both models process their respective inputs (visual data and textual cues) and their outputs are integrated to inform planning decisions. The language model extracts semantic information from textual descriptions of the environment and desired outcomes, while the vision model processes sensor data, and their combined insights improve trajectory prediction accuracy

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The patent adds a linguistic dimension to the traditional vision-only processing pipeline by introducing natural language as an additional information modality. The language model processes textual descriptions of driving scenarios, objectives, and constraints, adding a semantic layer of understanding that complements the spatial-temporal information from vision models, thereby enhancing overall planning precision

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

3Reliability

If the system integrates language comprehension with vision models, then semantic understanding improves, but the device complexity increases

Engineering Contradiction:
Improvecontextually aware planning decisionsVSAvoidsystem architecture complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The patent segments the autonomous driving system into distinct functional modules: a language comprehension module that processes textual input, a vision processing module that handles sensor data, and a planning module that integrates both. This segmentation allows each component to specialize in its function while reducing the complexity of any single component, as the language model only needs to understand text and generate plans without directly processing sensor data

Inventive Principle:
Principle #1Segmentation

4Productivity

If contrastive learning is used to fuse language and vision features, then planning capabilities are enhanced, but the training complexity and computational resources increase

Engineering Contradiction:
Improveplanning capability efficiencyVSAvoidtraining process complexity
Core Design Contradiction:
ProductivityVSDevice complexity

Solution Approach 1:

The patent applies contrastive learning during the training phase to pre-align the feature representations from language and vision models before deployment. By performing this alignment work during training using contrastive objectives that maximize agreement between linguistic and visual feature spaces, the system achieves improved planning capabilities without requiring complex real-time computations during actual autonomous driving operations

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS12528507B2Systems and methods for vision-language planning (VLP) foundation models for autonomous driving
Publication Date: 2026.01.20 ROBERT BOSCH GMBH
  • US12528507B2 patent drawing
  • US12528507B2 patent drawing
  • US12528507B2 patent drawing

AI summary

Methods and systems for training an autonomous driving system using a vision-language planning (VLP) model. Image data is obtained from a vehicle-mounted camera, encompassing details about agents situated within the external environment. Via image processing, the system identifies these agents within the environment. A Bird's Eye View (BEV) representation of the surroundings is then generated, encapsulating the spatiotemporal information linked to the vehicle and the recognized agents. Execution of the VLP machine learning model begins by extracting vision-based planning features from the BEV, and receiving or generating textual information characterizing various attributes of the vehicle within the environment. Text-based planning features are extracted from this textual information. To enhance model performance, a contrastive learning model is engaged to establish similarities between the vision-based and text-based planning features, and a predicted trajectory is output based on the similarities.