Autonomous Driving Scenario Simulation for Edge-Case AI Training Data

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing methods for generating training data for autonomous driving AI models are limited in securing data for edge cases, which are exceptional or extreme situations difficult to obtain, leading to performance limitations.

Innovation Solution

A method and apparatus that generate synthetic training data through autonomous driving simulation, using metadata and ground truth data from sensors to derive scenarios and create synthetic ground truth data for AI models, expanding the operational design domain.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If data is acquired through real-world autonomous driving, then the data reflects actual driving situations, but it is difficult to secure sufficient data for edge cases and requires enormous economic costs

Engineering Contradiction:
Improvedata quality for edge casesVSAvoiddata quantity for edge cases
Core Design Contradiction:
ReliabilityVSQuantity of substance

Solution Approach 1:

The patent uses simulation technology to create virtual copies of real-world driving scenarios, particularly edge cases that are difficult to obtain in reality. By generating synthetic training data through simulated environments, the system can produce unlimited quantities of edge case data without requiring actual physical occurrences of these rare situations.

Inventive Principle:
Principle #26Copying

Solution Approach 2:

The system performs preliminary identification of edge cases through uncertainty assessment before actual data collection. By predicting which scenarios are likely to be edge cases in advance, the simulation can be configured to generate these specific scenarios beforehand, ensuring sufficient data availability for training.

Inventive Principle:
Principle #10Preliminary action

2Measurement precision

If uncertainty-based selection is used to identify additional training data, then useful training examples can be found, but edge cases that are difficult to obtain remain insufficient

Engineering Contradiction:
Improveuncertainty assessment accuracyVSAvoiddata coverage for edge cases
Core Design Contradiction:
Measurement precisionVSAdaptability or versatility

Solution Approach 1:

The patent introduces simulation technology as an intermediary between uncertainty assessment and data collection. The uncertainty assessment identifies potential edge cases, and the simulation acts as a mediator to generate the actual training data for these identified cases, bridging the gap between identification and data acquisition.

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The system uses uncertainty assessment results as feedback to guide the simulation data generation process. By continuously monitoring model uncertainty and using this information to generate targeted synthetic data, the system creates a feedback loop that progressively improves coverage of edge cases.

Inventive Principle:
Principle #23Feedback

3Productivity

If more real-world data is collected to cover all driving situations, then model performance can be improved, but economic costs increase enormously

Engineering Contradiction:
Improvemodel performanceVSAvoideconomic cost
Core Design Contradiction:
ProductivityVSLoss of energy

Solution Approach 1:

Instead of collecting expensive real-world data for all possible driving situations, the patent creates virtual copies through simulation. This allows comprehensive coverage of diverse driving scenarios including rare edge cases without the proportional increase in data collection costs.

Inventive Principle:
Principle #26Copying

Solution Approach 2:

The system changes the parameter of data acquisition from physical collection to virtual generation. By transforming the data source from real-world sensors to simulation environments, the system maintains data diversity and model performance while dramatically reducing economic costs.

Inventive Principle:
Principle #35Parameter changes

Data Source

PatentUS20250272455A1Method and Apparatus for Generating Training Data for Artificial Intelligence Model
Publication Date: 2025.08.28 HYUNDAI MOTOR CO LTD
  • US20250272455A1 patent drawing
  • US20250272455A1 patent drawing
  • US20250272455A1 patent drawing

AI summary

A method of generating training data for an artificial intelligence model may comprise: providing metadata and ground truth data generated from original data; generating accumulated data obtained by accumulating training results less than or equal to a standard indicator target value by comparing artificial intelligence model training results with a target indicator based on the ground truth data; deriving an autonomous driving scenario using the accumulated data; and generating synthetic ground truth data for training the artificial intelligence model by performing a simulation according to the autonomous driving scenario.