Autonomous Vehicle Training Data Augmentation for Rare Scenario Coverage
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing autonomous vehicle training datasets suffer from gaps in feature coverage, missing rare events, and imbalanced data distribution, leading to inadequate preparation for diverse real-world scenarios, which can result in unexpected behavior and inefficiencies in model training.
Innovation Solution
A method for generating synthetic training data by analyzing initial datasets to identify deficiencies, creating targeted prompts for generative models to produce realistic scenarios, and validating the data to minimize hallucinations, ensuring comprehensive and relevant training data for autonomous vehicles.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If real-world data collection is used to train autonomous vehicle models, then the training data reflects actual driving conditions, but the data coverage is incomplete and lacks rare events
Solution Approach 1:
The patent merges real-world data collection with synthetic data generation to create a hybrid training dataset. Real data provides authentic driving conditions while synthetic data generated by generative models fills gaps in feature coverage and adds rare events, combining the strengths of both approaches to resolve the contradiction between realism and comprehensiveness
Solution Approach 2:
The patent uses generative models to create synthetic copies of real-world driving scenarios. These synthetic data copies replicate actual driving conditions while allowing for controlled variation and inclusion of rare events that would be difficult to capture in real-world collection, thereby expanding feature coverage while maintaining realism
2Adaptability or versatility
If more diverse training scenarios are collected to improve model adaptability, then the model can handle more situations, but the data collection time and cost increase
Solution Approach 1:
Instead of physically collecting diverse scenarios in the real world, the patent uses generative models to synthesize training data that replicates diverse driving scenarios. This copying approach achieves high scenario diversity without the time and cost penalties of actual field data collection, as synthetic data can be generated rapidly once the model is trained
Solution Approach 2:
The patent performs preliminary data collection to train the generative model, then uses this model to generate additional diverse scenarios. This preliminary action creates a reusable system where the bulk of diverse scenario generation happens through fast synthetic data production rather than repeated time-consuming field collection
3Adaptability or versatility
If synthetic data is generated to fill data gaps, then the training data coverage improves, but the risk of hallucination and unrealistic data increases
Solution Approach 1:
The patent combines synthetic data generation with real-world data validation and filtering. Synthetic data expands coverage while real data constraints and validation procedures ensure realism, merging the advantages of both approaches to achieve comprehensive coverage without excessive hallucination
Solution Approach 2:
The patent implements feedback mechanisms where generated synthetic data is evaluated against real-world distributions and constraints. This feedback loop identifies and corrects hallucinations, ensuring that synthetic data maintains realism while achieving improved feature coverage
Data Source
AI summary
In variants, a method for generating synthetic data can include: determining an initial dataset, determining characteristics of the initial dataset, generating a set of prompts based on the characteristics, prompting a model to generate the synthetic data using the set of prompts, and training an AV model using the synthetic data.


