Synthetic Image Generation for CNN Object Detection

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current convolutional neural networks (CNNs) used in autonomous vehicles face challenges in accurately detecting objects, especially motorcycles, due to underrepresentation in training datasets, leading to high false positives and poor performance in diverse environmental conditions.

Innovation Solution

The method involves generating simulated images using photorealistic rendering software to address underrepresented object configurations and noise factors, and processing these images with a generative adversarial network (GAN) to create synthetic images that include a variety of noise factors, which are then used to augment the training dataset, thereby improving the CNN's ability to correctly identify objects in real-world scenarios.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If real-world images are used for training CNN, then the training dataset reflects actual environmental conditions, but underrepresented object configurations and noise factors lead to poor detection accuracy

Engineering Contradiction:
Improveobject detection accuracyVSAvoidtraining data coverage
Core Design Contradiction:
Measurement precisionVSQuantity of substance

Solution Approach 1:

The patent creates synthetic copies of real-world images through photorealistic rendering and GAN processing. These synthetic images replicate actual environmental conditions, object configurations, and noise factors while providing unlimited diversity. The copying approach allows the training dataset to include rare and underrepresented scenarios without requiring extensive real-world data collection.

Inventive Principle:
Principle #26Copying

Solution Approach 2:

The patent performs preliminary generation of synthetic training images before CNN training. By pre-generating diverse object configurations and environmental conditions through rendering engines and GANs, the system prepares a comprehensive training dataset in advance, ensuring coverage of underrepresented cases before the actual training process begins.

Inventive Principle:
Principle #10Preliminary action

2Measurement precision

If more diverse training data is collected from real-world scenarios, then detection accuracy improves, but data collection time and resources increase significantly

Engineering Contradiction:
Improvedetection accuracyVSAvoiddata collection time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

Instead of collecting diverse real-world images over extended periods, the patent copies existing images through synthetic generation processes. The photorealistic rendering engine and GANs create unlimited variations of training scenarios instantaneously, eliminating the time-consuming data collection process while maintaining diversity and realism.

Inventive Principle:
Principle #26Copying

Solution Approach 2:

The patent replaces the mechanical process of physical data collection (traveling to various locations, capturing images under different conditions) with computational synthesis. The rendering engine and GANs substitute for physical data gathering, generating diverse training scenarios through algorithms rather than field collection.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

3Quantity of substance

If synthetic images are generated to augment training data, then training dataset diversity improves, but image realism and quality may deteriorate

Engineering Contradiction:
Improvetraining data diversityVSAvoidimage quality
Core Design Contradiction:
Quantity of substanceVSManufacturing precision

Solution Approach 1:

The patent introduces a GAN (generative adversarial network) as an intermediary between the photorealistic rendering engine and the CNN training process. The GAN refines the synthetically generated images by learning from real-world image distributions, adding realistic noise factors and environmental details that bridge the gap between synthetic generation and photorealism.

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The patent replaces direct photorealistic rendering with a two-stage process involving GAN refinement. Instead of relying solely on rendering engines to produce realistic images, the system uses GANs to transform synthetic images into photorealistic outputs, substituting algorithmic refinement for direct physical simulation.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

4Productivity

If the CNN is trained with limited real-world data, then training speed is faster, but false positives increase in underrepresented scenarios

Engineering Contradiction:
Improvetraining speedVSAvoidfalse positive rate
Core Design Contradiction:
ProductivityVSReliability

Solution Approach 1:

The patent performs preliminary generation of comprehensive synthetic training data that covers underrepresented scenarios before CNN training. By pre-populating the training dataset with diverse synthetic images including rare object configurations and environmental conditions, the system ensures the CNN learns from adequate examples during training, reducing false positives without sacrificing training speed.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent creates synthetic copies of underrepresented scenarios to supplement limited real-world training data. These copied and synthesized images provide the CNN with sufficient examples of rare cases during training, enabling the model to generalize better and reduce false positives while maintaining efficient training throughput.

Inventive Principle:
Principle #26Copying

Data Source

PatentUS11645360B2Neural network image processing
Publication Date: 2023.05.09 FORD GLOBAL TECH LLC
  • US11645360B2 patent drawing
  • US11645360B2 patent drawing
  • US11645360B2 patent drawing

AI summary

A computer, including a processor and a memory, the memory including instructions to be executed by the processor to determine a second convolutional neural network (CNN) training dataset by determining an underrepresented object configuration and an underrepresented noise factor corresponding to an object in a first CNN training dataset, generate one or more simulated images including the object corresponding to the underrepresented object configuration in the first CNN training dataset by inputting ground truth data corresponding to the object into a photorealistic rendering engine and generate one or more synthetic images including the object corresponding to the underrepresented noise factor in the first CNN training dataset by processing the simulated images with a generative adversarial network (GAN) to determine a second CNN training dataset. The instructions can include further instructions to train a CNN to using the first and the second CNN training datasets.