Synthesized Image Generation for Object Detection Learning Data

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing image recognition technologies face challenges in accurately detecting object regions with limited learning data, often leading to erroneous detections due to the difficulty in covering all object features and learning object features with few textures.

Innovation Solution

A method involving the generation of synthesized images within closed regions of an input image, along with corresponding labels, to create learning data that enhances the detection accuracy of object regions, utilizing a combination of image synthesis and deep learning techniques to improve multi-task detection.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If multi-task detection is used to detect unspecified objects, then the versatility of object detection is improved, but the detection precision deteriorates due to difficulty in covering all object features with limited learning data

Engineering Contradiction:
Improveversatility of object detectionVSAvoiddetection precision
Core Design Contradiction:
Adaptability or versatilityVSMeasurement precision

Solution Approach 1:

The system performs preliminary actions by generating synthesized learning data before actual detection tasks. The image generation unit creates artificial images with various object textures and appearances in advance, which are then used to train the detection model. This preliminary data preparation enables the model to be pre-adapted to diverse object features, improving both versatility and precision when facing unspecified objects in real detection scenarios.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The system creates copies of existing object images through synthesis to generate diverse training data. The image generation unit produces multiple variations of object images by manipulating textures, colors, and appearances while preserving the underlying object structures. These synthesized copies serve as additional learning data, enabling the detection model to recognize a broader range of object features without requiring extensive real-world data collection for each object type.

Inventive Principle:
Principle #26Copying

2Device complexity

If learning data is limited, then the complexity of data preparation is reduced, but the detection precision deteriorates due to inability to cover all object features

Engineering Contradiction:
Improvecomplexity of data preparationVSAvoiddetection precision
Core Design Contradiction:
Device complexityVSMeasurement precision

Solution Approach 1:

The system creates synthetic copies of existing images to expand limited learning data. The image generation unit synthesizes new training images by combining existing image elements with varied textures and appearances, effectively multiplying the information content of the original limited dataset without requiring proportional increases in data collection efforts.

Inventive Principle:
Principle #26Copying

Solution Approach 2:

The system changes parameters of existing images during synthesis, such as texture patterns, color distributions, and appearance characteristics. By manipulating these parameters while preserving the fundamental object structures, the system generates diverse training variations from limited source data, enabling the model to learn robust object features without increasing data preparation complexity.

Inventive Principle:
Principle #35Parameter changes

3Measurement precision

If domain randomization is used to synthesize learning data, then the detection precision for specific objects is improved, but the ability to learn object features with few textures deteriorates

Engineering Contradiction:
Improvedetection precision for specific objectVSAvoidability to learn object features
Core Design Contradiction:
Measurement precisionVSAdaptability or versatility

Solution Approach 1:

The system applies local quality by selectively randomizing only certain aspects of image synthesis while preserving others. Specifically, textures and appearances are randomized within local regions or for specific object parts, while the fundamental object structures and geometries remain consistent. This targeted approach enables the model to learn both specific object features and generalizable texture variations, resolving the trade-off between specific detection precision and overall feature learning capability.

Inventive Principle:
Principle #3Local quality

Data Source

PatentUS20230237777A1Information processing apparatus, learning apparatus, image recognition apparatus, information processing method, learning method, image recognition method, and non-transitory-computer-readable storage medium
Publication Date: 2023.07.27 CANON KK
  • US20230237777A1 patent drawing
  • US20230237777A1 patent drawing
  • US20230237777A1 patent drawing

AI summary

An information processing apparatus comprises a first generation unit configured to generate a synthesized image in which a second image is synthesized in a closed region in a first image, and a second generation unit configured to generate learning data, the learning data including a label and the synthesized image, the label indicating an object region including a region corresponding to the closed region in the synthesized image.