Contrastive Pre-Training for Autonomous Driving Edge Cases

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing machine learning models for autonomous vehicle operation require extensive training time due to the limited availability of real-world data, especially for edge cases that do not occur frequently, such as collisions and off-road trajectories.

Innovation Solution

Pre-training machine learning components using a perturbed dataset generated from real-world driving logs, which includes modifications to vehicle trajectories, lane additions/removals, and actor inclusion/exclusion, to learn high-level features, followed by inserting these pre-trained components into the primary model for training.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If real-world driving data is used for training, then model reliability is improved, but training time increases due to limited data availability

Engineering Contradiction:
Improvemodel reliabilityVSAvoidtraining time
Core Design Contradiction:
ReliabilityVSLoss of time

Solution Approach 1:

The patent applies preliminary action by pre-training model components on synthetic perturbed data before final training on real-world data. This advance preparation allows the model to learn from expanded datasets including edge cases beforehand, reducing the time needed for subsequent training while maintaining reliability through progressive learning stages.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent uses copying by creating perturbed versions of real driving data that simulate edge cases and rare scenarios. These synthetic copies expand the training dataset without requiring additional real-world data collection, allowing the model to learn from diverse scenarios while maintaining the core characteristics of authentic driving situations.

Inventive Principle:
Principle #26Copying

2Adaptability or versatility

If more training data is collected to cover edge cases, then model adaptability is improved, but data availability remains limited

Engineering Contradiction:
Improvemodel adaptabilityVSAvoiddata availability
Core Design Contradiction:
Adaptability or versatilityVSQuantity of substance

Solution Approach 1:

The patent applies parameter changes by systematically modifying data parameters such as trajectory perturbations, lane configuration changes, and actor position variations. These parameter transformations generate diverse training scenarios from limited real data, enabling the model to learn adaptability across multiple conditions without requiring proportional increases in actual data collection.

Inventive Principle:
Principle #35Parameter changes

Solution Approach 2:

The patent uses segmentation by dividing the training process into distinct stages: pre-training on perturbed synthetic data and fine-tuning on real-world data. This segmented approach allows each stage to focus on specific learning objectives, with the first stage building general adaptability and the second stage refining performance on authentic scenarios.

Inventive Principle:
Principle #1Segmentation

Data Source

PatentUS20260028047A1Pre-Training Machine Learning Models with Contrastive Learning
Publication Date: 2026.01.29 MOTIONAL AD LLC
  • US20260028047A1 patent drawing
  • US20260028047A1 patent drawing
  • US20260028047A1 patent drawing

AI summary

Provided are methods for pre-training machine learning models with contrastive learning. Some methods described also include generating, with at least one processor, a perturbed dataset from a real dataset. The method includes pre-training, with the at least one processor, at least one component of a machine leaning model to perform an alternative task, wherein the machine learning model performs a primary task. The method also includes inserting, with the at least one processor, the pre-trained at least one component into the machine learning model that performs the primary task. Additionally, the method includes training, with the at least one processor, the machine learning model comprising the pre-trained at least one component to perform the primary task. Systems and computer program products are also provided.