Object Detection Learning Data Generation from Difficult Regions
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing systems face challenges in generating learning data for object detection devices to accurately identify various objects in diverse environments without missed or false detections, requiring significant manual effort and simulation systems.
Innovation Solution
An information processing apparatus that detects objects in images, tracks them chronologically, estimates regions of difficulty, and generates images by superimposing objects based on estimation accuracy, using machine learning techniques to enhance detection precision.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If multiple images of difficult scenes are collected and manual labeling is performed to generate learning data, then the detection device can be optimized for specific environments, but the manual work effort increases significantly
Solution Approach 1:
The patent uses CG (computer graphics) images as copies of real scenes to generate learning data. Instead of manually labeling photographs of difficult scenes, the system creates synthetic images that replicate challenging detection scenarios, thereby eliminating manual work effort while maintaining the ability to optimize detection precision for specific environments
Solution Approach 2:
The system automatically generates and labels learning data through CG image synthesis without requiring human intervention for data collection and annotation. The automated process serves itself by generating the necessary training data on-demand, resolving the contradiction between detection precision and manual effort
2Ease of manufacture
If simulation systems are used to generate learning data of various images, then the generation of learning data becomes easier, but it remains difficult to generate images that cover all possible object positions and environments without missed or false detection
Solution Approach 1:
The patent employs a detection device that performs multiple functions: it detects objects in real images, identifies difficult regions, estimates parameters, and generates CG learning data. This multi-functional approach allows the same system to both use and create training data, improving adaptability across environments while maintaining ease of data generation
Solution Approach 2:
The system uses detection results from real images to identify difficult regions, which then inform the generation of targeted CG images. This feedback loop ensures that learning data is generated specifically for challenging scenarios, improving environmental coverage while keeping the generation process systematic and manageable
3Measurement precision
If detection devices are optimized for specific locations using collected images, then detection precision improves for those locations, but the device performance deteriorates in other environments
Solution Approach 1:
The detection device is designed to operate across multiple environments by incorporating CG image generation capabilities that can synthesize training data for various conditions. This allows the device to maintain detection precision across different locations rather than being optimized for a single environment
Solution Approach 2:
The system performs preliminary analysis of real detection images to identify difficult regions and estimate parameters before generating CG images. This preliminary action creates a foundation for generating comprehensive learning data that prepares the detection device for various environmental conditions, improving both precision and adaptability
Data Source
AI summary
An information processing apparatus includes at least one memory storing instructions and at least one processor. Upon execution of the stored instructions, the at least one processor causes the information processing apparatus to detect an object in a captured image, track the object in a chronologically captured image, based on a result of the detection of the object, estimate, from the captured image, a region in which detection of the object in the captured image is difficult for the detection unit per type of the object, based on the result of the detection of the object and a result of the tracking of the object, and generate an image acquired by superimposing, on a predetermined background image, a predetermined object image that corresponds to the type of the object, based on a result of the estimation.


