Synthetic Image Generation for Object Detection Training

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Training object detection systems requires extensive resources and time to collect and label millions of images under various conditions, such as different lighting and weather conditions, making the process inefficient and resource-intensive.

Innovation Solution

A method involving a device that renders a machine model into a simulated environment with varying characteristics and motion data, capturing images and generating bounding box data to create a dataset for training an object detection system, allowing for efficient generation of synthetic training data that mimics real-world conditions.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If real images are collected under various conditions to train object detection system, then training data quality is improved, but resource consumption and time required increase significantly

Engineering Contradiction:
Improvetraining data qualityVSAvoidtime required
Core Design Contradiction:
ReliabilityVSLoss of time

Solution Approach 1:

The patent uses synthetic image generation to create copies of training data through computer-generated simulations rather than collecting real images. A rendering engine generates synthetic images of objects under various conditions (lighting, weather, angles) by copying and transforming 3D model data, eliminating the need for physical photo collection while maintaining training data quality

Inventive Principle:
Principle #26Copying

Solution Approach 2:

The patent performs preliminary actions by pre-rendering synthetic images covering a wide range of conditions before actual training begins. The system pre-generates training datasets with diverse lighting, weather, and positional variations in advance, so that when training is needed, the data is already prepared and available immediately

Inventive Principle:
Principle #10Preliminary action

2Reliability

If real images are collected under various conditions to train object detection system, then training data quality is improved, but resource consumption increases significantly

Engineering Contradiction:
Improvetraining data qualityVSAvoidresource consumption
Core Design Contradiction:
ReliabilityVSQuantity of substance

Solution Approach 1:

The patent replaces physical resource-intensive image collection with digital synthetic image generation. Instead of deploying cameras and personnel to capture real images under various conditions, the system copies 3D object models and renders them programmatically, reducing resource consumption from physical field operations to computational rendering

Inventive Principle:
Principle #26Copying

Solution Approach 2:

The patent creates a universal synthetic data generation system that can produce training images for multiple object types and conditions using a single rendering pipeline. The same rendering engine handles different objects, lighting conditions, weather scenarios, and camera angles, eliminating the need for separate data collection efforts for each scenario

Inventive Principle:
Principle #6Universality (Multi-functionality)

3Productivity

If synthetic images are generated to reduce resource consumption, then resource efficiency is improved, but image realism may deteriorate

Engineering Contradiction:
Improveresource efficiencyVSAvoidimage realism
Core Design Contradiction:
ProductivityVSReliability

Solution Approach 1:

The patent uses parameter changes to enhance synthetic image realism by adjusting rendering parameters such as lighting models, shadow calculations, reflection properties, and atmospheric effects. The rendering engine varies these parameters to simulate different weather conditions, times of day, and environmental factors, making synthetic images indistinguishable from real photographs

Inventive Principle:
Principle #35Parameter changes

Solution Approach 2:

The patent introduces an intermediary rendering engine that bridges the gap between simple 3D models and photorealistic images. This intermediary layer applies complex graphical algorithms for lighting, shading, and environmental interaction to transform basic model data into highly realistic synthetic images that maintain both visual quality and resource efficiency

Inventive Principle:
Principle #24Intermediary (Mediator)

4Adaptability or versatility

If manual image collection is used to ensure diverse conditions, then condition coverage is improved, but automation level deteriorates

Engineering Contradiction:
Improvecondition coverageVSAvoidautomation level
Core Design Contradiction:
Adaptability or versatilityVSExtent of automation

Solution Approach 1:

The patent uses dynamics to automatically vary image conditions through programmable parameters. The rendering system dynamically adjusts lighting angles, weather conditions, object positions, and camera perspectives based on predefined ranges and randomization algorithms, automatically generating diverse training conditions without manual intervention for each scenario

Inventive Principle:
Principle #15Dynamics

Data Source

PatentUS11132826B2Artificial image generation for training an object detection system
Publication Date: 2021.09.28 CATERPILLAR INC
  • US11132826B2 patent drawing
  • US11132826B2 patent drawing
  • US11132826B2 patent drawing

AI summary

A device is disclosed. The device may obtain a machine model of a machine and associated motion data for the machine model of the machine. The device may render the machine model into a rendered environment associated with a set of characteristics. The device may capture a set of images of the machine model and the rendered environment based on rendering the machine model into the rendered environment. The device may determine bounding box data for the set of images of the machine model and the rendered environment based on the position of the machine model within the rendered environment relative to an image capture orientation within the rendered environment. The device may provide the set of images of the machine model and the rendered environment and the bounding box data as a data set for object detection.