Synthetic Image Generation for CNN Object Detection
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current convolutional neural networks (CNNs) used in autonomous vehicles face challenges in accurately detecting objects, especially motorcycles, due to underrepresentation in training datasets, leading to high false positives and poor performance in diverse environmental conditions.
Innovation Solution
The method involves generating simulated images using photorealistic rendering software to address underrepresented object configurations and noise factors, and processing these images with a generative adversarial network (GAN) to create synthetic images that include a variety of noise factors, which are then used to augment the training dataset, thereby improving the CNN's ability to correctly identify objects in real-world scenarios.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If real-world images are used for training CNN, then the training dataset reflects actual environmental conditions, but underrepresented object configurations and noise factors lead to poor detection accuracy
Solution Approach 1:
The patent creates synthetic copies of real-world images through photorealistic rendering and GAN processing. These synthetic images replicate actual environmental conditions, object configurations, and noise factors while providing unlimited diversity. The copying approach allows the training dataset to include rare and underrepresented scenarios without requiring extensive real-world data collection.
Solution Approach 2:
The patent performs preliminary generation of synthetic training images before CNN training. By pre-generating diverse object configurations and environmental conditions through rendering engines and GANs, the system prepares a comprehensive training dataset in advance, ensuring coverage of underrepresented cases before the actual training process begins.
2Measurement precision
If more diverse training data is collected from real-world scenarios, then detection accuracy improves, but data collection time and resources increase significantly
Solution Approach 1:
Instead of collecting diverse real-world images over extended periods, the patent copies existing images through synthetic generation processes. The photorealistic rendering engine and GANs create unlimited variations of training scenarios instantaneously, eliminating the time-consuming data collection process while maintaining diversity and realism.
Solution Approach 2:
The patent replaces the mechanical process of physical data collection (traveling to various locations, capturing images under different conditions) with computational synthesis. The rendering engine and GANs substitute for physical data gathering, generating diverse training scenarios through algorithms rather than field collection.
3Quantity of substance
If synthetic images are generated to augment training data, then training dataset diversity improves, but image realism and quality may deteriorate
Solution Approach 1:
The patent introduces a GAN (generative adversarial network) as an intermediary between the photorealistic rendering engine and the CNN training process. The GAN refines the synthetically generated images by learning from real-world image distributions, adding realistic noise factors and environmental details that bridge the gap between synthetic generation and photorealism.
Solution Approach 2:
The patent replaces direct photorealistic rendering with a two-stage process involving GAN refinement. Instead of relying solely on rendering engines to produce realistic images, the system uses GANs to transform synthetic images into photorealistic outputs, substituting algorithmic refinement for direct physical simulation.
4Productivity
If the CNN is trained with limited real-world data, then training speed is faster, but false positives increase in underrepresented scenarios
Solution Approach 1:
The patent performs preliminary generation of comprehensive synthetic training data that covers underrepresented scenarios before CNN training. By pre-populating the training dataset with diverse synthetic images including rare object configurations and environmental conditions, the system ensures the CNN learns from adequate examples during training, reducing false positives without sacrificing training speed.
Solution Approach 2:
The patent creates synthetic copies of underrepresented scenarios to supplement limited real-world training data. These copied and synthesized images provide the CNN with sufficient examples of rare cases during training, enabling the model to generalize better and reduce false positives while maintaining efficient training throughput.
Data Source
AI summary
A computer, including a processor and a memory, the memory including instructions to be executed by the processor to determine a second convolutional neural network (CNN) training dataset by determining an underrepresented object configuration and an underrepresented noise factor corresponding to an object in a first CNN training dataset, generate one or more simulated images including the object corresponding to the underrepresented object configuration in the first CNN training dataset by inputting ground truth data corresponding to the object into a photorealistic rendering engine and generate one or more synthetic images including the object corresponding to the underrepresented noise factor in the first CNN training dataset by processing the simulated images with a generative adversarial network (GAN) to determine a second CNN training dataset. The instructions can include further instructions to train a CNN to using the first and the second CNN training datasets.


