Synthetic Data Generation for Autonomous Driving Neural Networks

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

The high cost and labor-intensive process of obtaining and preparing large-scale training data for autonomous driving neural networks, particularly for robust performance across various environments and conditions, is a significant challenge.

Innovation Solution

An apparatus and method that utilize processors to obtain localization information, extract landmark points from a landmark map, generate ground truth images, and refine these images to create training data, which can be used to train neural networks, reducing the need for manual labeling and enhancing data robustness across different conditions.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If manual labeling is used to obtain training data, then the quality and accuracy of ground truth data can be ensured, but the time consumption and labor cost increase significantly

Engineering Contradiction:
Improveground truth data qualityVSAvoiddata preparation time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The patent creates synthetic training data by copying and transforming real-world scenarios through simulation. Virtual environments replicate real driving conditions, allowing generation of large-scale training data without manual labeling while maintaining ground truth accuracy through programmed scene parameters and object properties.

Inventive Principle:
Principle #26Copying

Solution Approach 2:

The system performs automatic data generation and labeling through automated pipelines that synthesize training data with embedded ground truth information. The simulation framework automatically generates annotations, eliminating the need for manual labeling while ensuring data quality through controlled generation parameters.

Inventive Principle:
Principle #25Self-service

2Reliability

If large-scale training data is acquired through deep learning techniques, then the neural network performance improves, but the manpower and effort required increase significantly

Engineering Contradiction:
Improveneural network performanceVSAvoiddata acquisition efficiency
Core Design Contradiction:
ReliabilityVSProductivity

Solution Approach 1:

The patent generates large-scale training data through synthetic replication of driving scenarios. By copying real-world scene elements into virtual environments and applying various transformations, the system produces extensive training datasets automatically, achieving both large scale and high productivity simultaneously.

Inventive Principle:
Principle #26Copying

Solution Approach 2:

The system pre-generates training data with embedded ground truth information before neural network training. Synthetic data is created in advance with known annotations, allowing efficient large-scale data preparation that improves network performance without requiring significant manual effort during the training process.

Inventive Principle:
Principle #10Preliminary action

3Adaptability or versatility

If additional data is acquired for specific environments to improve robustness, then the neural network works better in those conditions, but the cost for additional acquisition is constantly required

Engineering Contradiction:
Improveenvironmental robustnessVSAvoiddata acquisition cost
Core Design Contradiction:
Adaptability or versatilityVSEase of manufacture

Solution Approach 1:

The patent achieves environmental adaptability by dynamically changing simulation parameters to represent different weather conditions, lighting scenarios, and environmental settings. By adjusting virtual scene parameters rather than collecting real-world data for each condition, the system improves robustness across diverse environments without incurring additional acquisition costs.

Inventive Principle:
Principle #35Parameter changes

Solution Approach 2:

The synthetic data generation system creates multi-functional training data that can be used across multiple environmental conditions. A single simulated scene can be transformed into various weather and lighting conditions through parameter adjustments, providing universal training data that improves neural network robustness across different environments without requiring separate data acquisition campaigns.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Data Source

PatentUS20230360381A1Method and apparatus with data labeling
Publication Date: 2023.11.09 SAMSUNG ELECTRONICS CO LTD
  • US20230360381A1 patent drawing
  • US20230360381A1 patent drawing
  • US20230360381A1 patent drawing

AI summary

An apparatus and method with data labeling are provided. An apparatus includes one or more processors configured to obtain localization information related to an object, based on the localization information, extract a landmark point from a landmark map including coordinates of a landmark, generate a ground truth image based on the extracted landmark point, and generate training data by refining the ground truth image.