Synthetic Image Augmentation for Autonomous Perception Training
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Autonomous vehicle perception systems face challenges in accurately detecting rare or infrequently encountered objects like traffic hazards, pedestrians, and bicycles due to limited diversity in training datasets, leading to unreliable object classification.
Innovation Solution
The use of a multi-camera network that augments real images with 3D synthetic models, incorporating realistic lighting and shadow information, to create a more diverse training dataset, reducing the domain gap between simulated and real-world scenarios, and improving the detection of traffic hazards with increased precision and accuracy.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Quantity of substance
If vehicles drive for hundreds of hours to collect images on interstate highways, then large quantities of images are gathered, but the diversity of objects and agents remains limited
Solution Approach 1:
The patent uses synthetic image generation to create copies of rare objects and agents that cannot be easily captured in real-world driving. By rendering 3D models of traffic hazards, pedestrians, bicycles, and other rare objects into synthetic images, the system supplements the limited real-world dataset with diverse synthetic counterparts, thereby increasing overall data diversity without requiring extended real-world collection periods.
Solution Approach 2:
The system performs preliminary data preparation by pre-rendering synthetic images of rare objects under various conditions (lighting, weather, angles) before they are needed for training. This advance preparation ensures that diverse object representations are already available when the model training begins, eliminating the need to wait for rare real-world occurrences during data collection.
2Reliability
If real-world images are used for training, then the model learns from actual scenarios, but rare and dangerous objects are insufficiently represented
Solution Approach 1:
The patent merges real-world images with synthetic images of rare objects to create a hybrid training dataset. The real images provide authentic background scenes and common objects, while the synthetic images contribute diverse representations of rare traffic hazards, pedestrians, and cyclists. This combination allows the model to learn from both actual driving scenarios and supplemented rare object examples, improving overall training reliability.
Solution Approach 2:
The system uses an intermediary synthetic data generation pipeline that bridges the gap between real-world data limitations and model training requirements. The synthetic image generation process acts as a mediator, translating 3D models of rare objects into realistic 2D images that can be seamlessly integrated with real-world photographs, thereby enabling the model to learn about rare objects without direct real-world exposure.
3Adaptability or versatility
If synthetic images are generated without realistic lighting, then data diversity increases, but the domain gap between synthetic and real images widens
Solution Approach 1:
The system dynamically adjusts rendering parameters such as lighting conditions, shadow intensity, and environmental factors during synthetic image generation to match the characteristics of real-world images. By varying these parameters across different synthetic samples, the system maintains data diversity while ensuring each synthetic image adheres to realistic physical constraints, thereby reducing the domain gap between synthetic and real data.
Data Source
AI summary
In various examples, systems and methods are disclosed that relate to data augmentation for training/updating perception models in autonomous or semi-autonomous systems and applications. For example, a system may receive data associated with a set of frames that are captured using a plurality of cameras positioned in fixed relation relative to the machine; generate a panoramic view based at least on the set of frames; provide data associated with the panoramic view to a model to cause the model to generate a high dynamic range (HDR) panoramic view; determine lighting information associated with a light distribution map based at least on the HDR panoramic view; determine a virtual scene; and render an asset and a shadow on at least one of the frames, based at least on the virtual scene and the light distribution map, the shadow being a shadow corresponding to the asset.


