Autonomous Vehicle Training Data Augmentation for Rare Scenario Coverage

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing autonomous vehicle training datasets suffer from gaps in feature coverage, missing rare events, and imbalanced data distribution, leading to inadequate preparation for diverse real-world scenarios, which can result in unexpected behavior and inefficiencies in model training.

Innovation Solution

A method for generating synthetic training data by analyzing initial datasets to identify deficiencies, creating targeted prompts for generative models to produce realistic scenarios, and validating the data to minimize hallucinations, ensuring comprehensive and relevant training data for autonomous vehicles.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If real-world data collection is used to train autonomous vehicle models, then the training data reflects actual driving conditions, but the data coverage is incomplete and lacks rare events

Engineering Contradiction:
Improvetraining data realismVSAvoidfeature coverage
Core Design Contradiction:
ReliabilityVSAdaptability or versatility

Solution Approach 1:

The patent merges real-world data collection with synthetic data generation to create a hybrid training dataset. Real data provides authentic driving conditions while synthetic data generated by generative models fills gaps in feature coverage and adds rare events, combining the strengths of both approaches to resolve the contradiction between realism and comprehensiveness

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The patent uses generative models to create synthetic copies of real-world driving scenarios. These synthetic data copies replicate actual driving conditions while allowing for controlled variation and inclusion of rare events that would be difficult to capture in real-world collection, thereby expanding feature coverage while maintaining realism

Inventive Principle:
Principle #26Copying

2Adaptability or versatility

If more diverse training scenarios are collected to improve model adaptability, then the model can handle more situations, but the data collection time and cost increase

Engineering Contradiction:
Improvescenario diversityVSAvoiddata collection time
Core Design Contradiction:
Adaptability or versatilityVSLoss of time

Solution Approach 1:

Instead of physically collecting diverse scenarios in the real world, the patent uses generative models to synthesize training data that replicates diverse driving scenarios. This copying approach achieves high scenario diversity without the time and cost penalties of actual field data collection, as synthetic data can be generated rapidly once the model is trained

Inventive Principle:
Principle #26Copying

Solution Approach 2:

The patent performs preliminary data collection to train the generative model, then uses this model to generate additional diverse scenarios. This preliminary action creates a reusable system where the bulk of diverse scenario generation happens through fast synthetic data production rather than repeated time-consuming field collection

Inventive Principle:
Principle #10Preliminary action

3Adaptability or versatility

If synthetic data is generated to fill data gaps, then the training data coverage improves, but the risk of hallucination and unrealistic data increases

Engineering Contradiction:
Improvedata coverageVSAvoiddata realism
Core Design Contradiction:
Adaptability or versatilityVSReliability

Solution Approach 1:

The patent combines synthetic data generation with real-world data validation and filtering. Synthetic data expands coverage while real data constraints and validation procedures ensure realism, merging the advantages of both approaches to achieve comprehensive coverage without excessive hallucination

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The patent implements feedback mechanisms where generated synthetic data is evaluated against real-world distributions and constraints. This feedback loop identifies and corrects hallucinations, ensuring that synthetic data maintains realism while achieving improved feature coverage

Inventive Principle:
Principle #23Feedback

Data Source

PatentUS12361689B1System and method for augmenting autonomous vehicle training data
Publication Date: 2025.07.15 GATIK AI INC
  • US12361689B1 patent drawing
  • US12361689B1 patent drawing
  • US12361689B1 patent drawing

AI summary

In variants, a method for generating synthetic data can include: determining an initial dataset, determining characteristics of the initial dataset, generating a set of prompts based on the characteristics, prompting a model to generate the synthetic data using the set of prompts, and training an AV model using the synthetic data.