Composite Image Generation for Person-Object Interaction Detection
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing technologies struggle to accurately identify which object a person is interacting with, especially in scenes with similar objects or numerous background objects, leading to decreased detection accuracy.
Innovation Solution
A generation program that generates composite image data by arranging extracted objects near persons in specific positions, using contrastive learning to train a machine learning model to enhance object identification accuracy.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If detailed detection rules are set in accordance with camera arrangement and person orientation to improve detection accuracy, then detection accuracy is improved, but device complexity increases
Solution Approach 1:
The patent creates composite images by copying and arranging extracted object images at specific positions relative to person images. This synthetic data generation approach replaces the need for complex detection rules, allowing the model to learn object-person relationships from simplified visual patterns rather than intricate rule-based logic.
Solution Approach 2:
The patent performs preliminary extraction of person and object images from original images, then generates composite images in advance with objects positioned at specific distances and orientations relative to persons. This pre-processing creates training data that simplifies the detection task, eliminating the need for complex real-time rule evaluation.
2Device complexity
If machine learning model is trained without detailed detection rules to simplify the system, then device complexity is reduced, but object identification accuracy decreases in scenes with similar or numerous objects
Solution Approach 1:
The patent extracts only the essential visual features of persons and objects from the original images, separating them from the complex background context. By creating composite images that isolate these extracted elements and their spatial relationships, the model learns to identify objects based on their visual characteristics and position relative to persons, rather than relying on complex rule-based detection.
Solution Approach 2:
The composite images serve as an intermediary representation between the original complex scenes and the simplified detection task. These synthetic images contain only the necessary information (person, object, and their spatial relationship) while eliminating distracting background elements, allowing the model to achieve high accuracy without complex rules.
3Measurement precision
If composite image data is generated with objects arranged at specific positions to enhance training effectiveness, then object identification accuracy is improved, but manufacturing precision requirements increase
Solution Approach 1:
The patent applies different processing qualities to different parts of the image. The person and object regions are extracted and repositioned with specific spatial relationships, while the background is either removed or blurred. This selective processing approach achieves the necessary precision for object identification without requiring perfect reconstruction of the entire image, reducing overall manufacturing precision requirements.
Data Source
AI summary
A non-transitory computer-readable recording medium has stored therein a generation program that causes a computer to execute a process including acquiring an image that includes a person extracting an object that is used by the person included in the image by analyzing the acquired image generating a composite image in which the extracted object is arranged at a position that satisfies a predetermined condition on a basis of a position of the object that is used by the person included in the acquired image and generating, by using the generated composite image, a machine learning model that has been trained to identify the person who uses the object.


