Domain Adaptation for Synthetic Perception Model Training

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current autonomous driving systems face challenges in training perception models due to the reality gap between simulated and real-world data, where sensor models in simulation do not perfectly mimic real-world sensors, and real-world content is difficult to incorporate into simulation, leading to poor generalization of machine learning models from virtual to real domains.

Innovation Solution

The use of domain-adaptation theory to minimize the reality gap through spatial prior sampling and a discriminator that allows models to learn domain-invariant representations, enabling the training of perception models using synthetic data that performs well in real-world scenarios by generating pseudo labels and controlling sensor and architecture discrepancies.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If data is collected from real-world sensors, then the training data reflects actual sensor characteristics and environmental diversity, but the financial and time costs scale with the amount of data being collected and labeled

Engineering Contradiction:
Improvetraining data qualityVSAvoiddata collection and labeling time
Core Design Contradiction:
ReliabilityVSLoss of time

Solution Approach 1:

The patent creates synthetic copies of real-world sensor data through simulation environments. Instead of collecting and labeling actual real-world data, the system generates synthetic training data that replicates the characteristics, noise patterns, and environmental conditions of real sensors, thereby eliminating the time-consuming data collection and labeling process while maintaining training data quality

Inventive Principle:
Principle #26Copying

Solution Approach 2:

The system varies multiple parameters in the simulation environment including sensor extrinsics, intrinsics, weather conditions, lighting, and object properties to generate diverse synthetic training data. This parameter variation allows the model to learn from a wide range of scenarios without requiring physical data collection for each condition

Inventive Principle:
Principle #35Parameter changes

2Reliability

If data is collected from real-world sensors, then the training data captures actual sensor noise and environmental variations, but changing sensors may require redoing the effort to a large extent

Engineering Contradiction:
Improvesensor data authenticityVSAvoidsensor configuration flexibility
Core Design Contradiction:
ReliabilityVSAdaptability or versatility

Solution Approach 1:

The simulation framework is designed to be universal and configurable for different sensor types and configurations. By parameterizing sensor models with adjustable extrinsics, intrinsics, and noise characteristics, the system can adapt to various sensor setups without requiring new data collection campaigns, making the training process versatile across different autonomous vehicle sensor configurations

Inventive Principle:
Principle #6Universality (Multi-functionality)

3Measurement precision

If data is collected from real-world sensors, then human annotators can label visible objects, but labeling ambiguous, far away, or occluded objects may be hard or even impossible for humans

Engineering Contradiction:
Improveobject labeling accuracyVSAvoidannotation process complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The system uses synthetic data generation to create perfect ground truth labels for all objects in the simulation environment, including occluded, distant, and ambiguous objects. Since the simulator has complete knowledge of the virtual world state, it can generate accurate labels without requiring human annotators to interpret difficult-to-see objects, thereby achieving high measurement precision without the complexity of manual annotation

Inventive Principle:
Principle #26Copying

4Productivity

If synthetic data is used for training, then computational and time resources are reduced, but sensor models in simulation do not mimic real-world sensors perfectly, leading to poor generalization from virtual to real-world domain

Engineering Contradiction:
Improvetraining efficiencyVSAvoidmodel generalization performance
Core Design Contradiction:
ProductivityVSReliability

Solution Approach 1:

The patent applies domain adaptation techniques that focus on matching specific local characteristics of real sensor data, such as noise patterns, compression artifacts, and sensor-specific distortions. By targeting these local quality aspects rather than attempting to perfectly replicate the entire data distribution, the system improves model generalization to real-world sensors while maintaining the efficiency benefits of synthetic training data

Inventive Principle:
Principle #3Local quality

Solution Approach 2:

The system introduces domain adaptation layers and techniques as intermediaries between the synthetic training data and the real-world deployment. These intermediaries include domain randomization, style transfer, and adaptation modules that bridge the gap between simulation and reality, enabling models trained on synthetic data to generalize effectively to real sensors

Inventive Principle:
Principle #24Intermediary (Mediator)

Data Source

PatentUS20220391766A1Training perception models using synthetic data for autonomous systems and applications
Publication Date: 2022.12.08 NVIDIA CORP
  • US20220391766A1 patent drawing
  • US20220391766A1 patent drawing
  • US20220391766A1 patent drawing

AI summary

In various examples, systems and methods are disclosed that use a domain-adaptation theory to minimize the reality gap between simulated and real-world domains for training machine learning models. For example, sampling of spatial priors may be used to generate synthetic data that that more closely matches the diversity of data from the real-world. To train models using this synthetic data that still perform well in the real-world, the systems and methods of the present disclosure may use a discriminator that allows a model to learn domain-invariant representations to minimize the divergence between the virtual world and the real-world in a latent space. As such, the techniques described herein allow for a principled approach to learn neural-invariant representations and a theoretically inspired approach on how to sample data from a simulator that, in combination, allow for training of machine learning models using synthetic data.