Contrastive Pre-Training for Autonomous Driving Edge Cases
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing machine learning models for autonomous vehicle operation require extensive training time due to the limited availability of real-world data, especially for edge cases that do not occur frequently, such as collisions and off-road trajectories.
Innovation Solution
Pre-training machine learning components using a perturbed dataset generated from real-world driving logs, which includes modifications to vehicle trajectories, lane additions/removals, and actor inclusion/exclusion, to learn high-level features, followed by inserting these pre-trained components into the primary model for training.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If real-world driving data is used for training, then model reliability is improved, but training time increases due to limited data availability
Solution Approach 1:
The patent applies preliminary action by pre-training model components on synthetic perturbed data before final training on real-world data. This advance preparation allows the model to learn from expanded datasets including edge cases beforehand, reducing the time needed for subsequent training while maintaining reliability through progressive learning stages.
Solution Approach 2:
The patent uses copying by creating perturbed versions of real driving data that simulate edge cases and rare scenarios. These synthetic copies expand the training dataset without requiring additional real-world data collection, allowing the model to learn from diverse scenarios while maintaining the core characteristics of authentic driving situations.
2Adaptability or versatility
If more training data is collected to cover edge cases, then model adaptability is improved, but data availability remains limited
Solution Approach 1:
The patent applies parameter changes by systematically modifying data parameters such as trajectory perturbations, lane configuration changes, and actor position variations. These parameter transformations generate diverse training scenarios from limited real data, enabling the model to learn adaptability across multiple conditions without requiring proportional increases in actual data collection.
Solution Approach 2:
The patent uses segmentation by dividing the training process into distinct stages: pre-training on perturbed synthetic data and fine-tuning on real-world data. This segmented approach allows each stage to focus on specific learning objectives, with the first stage building general adaptability and the second stage refining performance on authentic scenarios.
Data Source
AI summary
Provided are methods for pre-training machine learning models with contrastive learning. Some methods described also include generating, with at least one processor, a perturbed dataset from a real dataset. The method includes pre-training, with the at least one processor, at least one component of a machine leaning model to perform an alternative task, wherein the machine learning model performs a primary task. The method also includes inserting, with the at least one processor, the pre-trained at least one component into the machine learning model that performs the primary task. Additionally, the method includes training, with the at least one processor, the machine learning model comprising the pre-trained at least one component to perform the primary task. Systems and computer program products are also provided.


