Synthetic Training Data Generation for Autonomous Vehicle Control

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional methods for testing and training autonomous vehicle control software are expensive, time-consuming, and struggle to capture low-probability events, making it difficult to obtain comprehensive training data.

Innovation Solution

A computer-implemented method generates training data by transforming synthesized environmental representations, using semantic information to create a set of transformed images that include low-probability events, allowing for accelerated and automated labeling, and improving the testing of autonomous vehicle control software.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If conventional data collection methods are used to obtain labelled training data for autonomous vehicle control software, then the data can be used for testing and validation, but the process is massively expensive and time-consuming

Engineering Contradiction:
Improvetraining data qualityVSAvoiddata collection time
Core Design Contradiction:
ReliabilityVSLoss of time

Solution Approach 1:

The patent creates synthetic copies of real-world driving environments through simulation. Instead of collecting actual labeled data from physical tests, the system generates virtual representations of road scenes, vehicles, and environmental conditions that replicate real-world scenarios. This copying approach eliminates the time-consuming field data collection while maintaining training data quality for control software validation.

Inventive Principle:
Principle #26Copying

Solution Approach 2:

The patent performs preliminary synthesis of training data before actual testing begins. By pre-generating diverse driving scenarios, environmental conditions, and edge cases in virtual environments, the system prepares comprehensive training datasets in advance. This preliminary action eliminates the need for time-consuming data collection during the testing phase while ensuring reliability through pre-validated synthetic data.

Inventive Principle:
Principle #10Preliminary action

2Adaptability or versatility

If conventional data collection methods are used to capture low-probability events for training data, then comprehensive scenario coverage can be achieved, but the cost and time requirements become prohibitive

Engineering Contradiction:
Improvescenario coverageVSAvoiddata production ease
Core Design Contradiction:
Adaptability or versatilityVSEase of manufacture

Solution Approach 1:

The patent uses synthetic copying to reproduce rare and low-probability events that would be extremely difficult to capture in real-world data collection. The simulation system can deliberately generate edge cases such as unusual weather conditions, rare traffic scenarios, and exceptional environmental situations without the prohibitive costs and time requirements of waiting for these events to occur naturally during field testing.

Inventive Principle:
Principle #26Copying

Solution Approach 2:

The patent systematically varies environmental parameters, object positions, and scenario conditions in the simulation to generate diverse training data. By changing parameters such as weather conditions, lighting, road types, and vehicle configurations, the system achieves comprehensive scenario coverage including low-probability events while maintaining ease of data production through controlled virtual experimentation rather than expensive field deployment.

Inventive Principle:
Principle #35Parameter changes

3Measurement precision

If manual labeling of collected data is performed to create labelled training data, then accurate ground truth can be obtained, but the process becomes massively expensive and time-consuming

Engineering Contradiction:
Improvelabeling accuracyVSAvoiddata production rate
Core Design Contradiction:
Measurement precisionVSProductivity

Solution Approach 1:

The patent eliminates manual labeling by copying real-world scenarios into simulation environments where ground truth is automatically known. Since the synthetic data is generated from predefined models and parameters, the correct labels and annotations are inherently embedded in the generation process itself. This approach maintains measurement precision through accurate virtual representations while dramatically increasing productivity by removing the manual labeling bottleneck entirely.

Inventive Principle:
Principle #26Copying

Solution Approach 2:

The simulation system performs self-labeling by automatically generating annotated training data with inherent ground truth information. The synthetic environment inherently knows the correct positions, classifications, and relationships of all objects since it generates them according to defined rules and parameters. This self-service capability eliminates the need for external manual labeling while maintaining high accuracy, thereby大幅提升 data production rate.

Inventive Principle:
Principle #25Self-service

Data Source

PatentUS20240420482A1Method and Apparatus
Publication Date: 2024.12.19 OXA AUTONOMY LTD
  • US20240420482A1 patent drawing
  • US20240420482A1 patent drawing
  • US20240420482A1 patent drawing

AI summary

A computer-implemented method of generating training data, the method comprising: providing a representation of an environment, wherein the representation of the environment has a defined structure and/or a defined geometry; andgenerating the training data comprising a set of transformed representations, including a first transformed representation, of the environment by transforming the representation of the environment to the set of transformed representations, including the first transformed representation, of the environment;wherein providing the representation of the environment comprises synthesizing, at least in part, an image of the environment using semantic information.