Visual Prompt Generation for Accurate Image Object Detection
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Generating a suitable prompt for text-based object detection in images is burdensome for users, as the accuracy of object detection heavily depends on the prompt.
Innovation Solution
An information processing apparatus and method that generates a suitable prompt by obtaining a visually expressing text group, which includes a plurality of texts that visually represent the detection target, and provides this prompt to a detection model for accurate object detection.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If a user manually creates a prompt for object detection, then the prompt can be tailored to specific needs, but the user burden and time required increase significantly
Solution Approach 1:
The system automatically generates prompts by extracting visual expressions from images using an expression extraction model, eliminating the need for users to manually create prompts. The model self-services by analyzing the image content and generating appropriate detection prompts automatically
Solution Approach 2:
The expression extraction model pre-extracts visual expressions from the image before the detection process. This preliminary action of extracting and preparing prompts in advance reduces the user burden during the actual detection task
2Ease of operation
If a simple prompt is used for object detection, then the operation is easy and quick, but the detection accuracy deteriorates
Solution Approach 1:
The system changes the parameters of the prompt by generating multiple visual expressions that describe different aspects of the object (color, shape, texture, pattern). This transforms a simple single-word prompt into a comprehensive multi-parameter description, improving detection accuracy while maintaining ease of use
3Measurement precision
If multiple visual expressions are extracted from an image, then the prompt becomes more comprehensive and accurate, but the processing time and computational resources increase
Solution Approach 1:
The system extracts only the necessary visual expressions needed for accurate detection, avoiding unnecessary over-extraction. By using the expression extraction model to identify and select relevant visual features, the system achieves sufficient accuracy without excessive processing time
Data Source
AI summary
An information processing apparatus includes: a text group obtaining section which obtains a visually expressing text group that includes a plurality of texts which visually express a detection target, with reference to input data that specifies the detection target; a prompt generating section which generates a prompt with reference to the visually expressing text group; and a providing section which provides, to a detection model, the prompt that has been generated by the prompt generating section, the detection model being a model into which a prompt and an image are inputted and which detects, from the image, a detection target that is specified by the prompt.


