Composite Image Generation for Person-Object Interaction Detection

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing technologies struggle to accurately identify which object a person is interacting with, especially in scenes with similar objects or numerous background objects, leading to decreased detection accuracy.

Innovation Solution

A generation program that generates composite image data by arranging extracted objects near persons in specific positions, using contrastive learning to train a machine learning model to enhance object identification accuracy.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If detailed detection rules are set in accordance with camera arrangement and person orientation to improve detection accuracy, then detection accuracy is improved, but device complexity increases

Engineering Contradiction:
Improvedetection accuracyVSAvoiddetection rule complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent creates composite images by copying and arranging extracted object images at specific positions relative to person images. This synthetic data generation approach replaces the need for complex detection rules, allowing the model to learn object-person relationships from simplified visual patterns rather than intricate rule-based logic.

Inventive Principle:
Principle #26Copying

Solution Approach 2:

The patent performs preliminary extraction of person and object images from original images, then generates composite images in advance with objects positioned at specific distances and orientations relative to persons. This pre-processing creates training data that simplifies the detection task, eliminating the need for complex real-time rule evaluation.

Inventive Principle:
Principle #10Preliminary action

2Device complexity

If machine learning model is trained without detailed detection rules to simplify the system, then device complexity is reduced, but object identification accuracy decreases in scenes with similar or numerous objects

Engineering Contradiction:
Improvesystem complexityVSAvoidobject identification accuracy
Core Design Contradiction:
Device complexityVSMeasurement precision

Solution Approach 1:

The patent extracts only the essential visual features of persons and objects from the original images, separating them from the complex background context. By creating composite images that isolate these extracted elements and their spatial relationships, the model learns to identify objects based on their visual characteristics and position relative to persons, rather than relying on complex rule-based detection.

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The composite images serve as an intermediary representation between the original complex scenes and the simplified detection task. These synthetic images contain only the necessary information (person, object, and their spatial relationship) while eliminating distracting background elements, allowing the model to achieve high accuracy without complex rules.

Inventive Principle:
Principle #24Intermediary (Mediator)

3Measurement precision

If composite image data is generated with objects arranged at specific positions to enhance training effectiveness, then object identification accuracy is improved, but manufacturing precision requirements increase

Engineering Contradiction:
Improveobject identification accuracyVSAvoidimage processing precision
Core Design Contradiction:
Measurement precisionVSManufacturing precision

Solution Approach 1:

The patent applies different processing qualities to different parts of the image. The person and object regions are extracted and repositioned with specific spatial relationships, while the background is either removed or blurred. This selective processing approach achieves the necessary precision for object identification without requiring perfect reconstruction of the entire image, reducing overall manufacturing precision requirements.

Inventive Principle:
Principle #3Local quality

Data Source

PatentUS20260065652A1Non-transitory computer-readable recording medium, generation method, and information processing apparatus
Publication Date: 2026.03.05 FUJITSU LTD
  • US20260065652A1 patent drawing
  • US20260065652A1 patent drawing
  • US20260065652A1 patent drawing

AI summary

A non-transitory computer-readable recording medium has stored therein a generation program that causes a computer to execute a process including acquiring an image that includes a person extracting an object that is used by the person included in the image by analyzing the acquired image generating a composite image in which the extracted object is arranged at a position that satisfies a predetermined condition on a basis of a position of the object that is used by the person included in the acquired image and generating, by using the generated composite image, a machine learning model that has been trained to identify the person who uses the object.