Visual Prompt Generation for Accurate Image Object Detection

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Generating a suitable prompt for text-based object detection in images is burdensome for users, as the accuracy of object detection heavily depends on the prompt.

Innovation Solution

An information processing apparatus and method that generates a suitable prompt by obtaining a visually expressing text group, which includes a plurality of texts that visually represent the detection target, and provides this prompt to a detection model for accurate object detection.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If a user manually creates a prompt for object detection, then the prompt can be tailored to specific needs, but the user burden and time required increase significantly

Engineering Contradiction:
Improveobject detection accuracyVSAvoiduser burden
Core Design Contradiction:
Measurement precisionVSEase of operation

Solution Approach 1:

The system automatically generates prompts by extracting visual expressions from images using an expression extraction model, eliminating the need for users to manually create prompts. The model self-services by analyzing the image content and generating appropriate detection prompts automatically

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The expression extraction model pre-extracts visual expressions from the image before the detection process. This preliminary action of extracting and preparing prompts in advance reduces the user burden during the actual detection task

Inventive Principle:
Principle #10Preliminary action

2Ease of operation

If a simple prompt is used for object detection, then the operation is easy and quick, but the detection accuracy deteriorates

Engineering Contradiction:
Improveoperation simplicityVSAvoidobject detection accuracy
Core Design Contradiction:
Ease of operationVSMeasurement precision

Solution Approach 1:

The system changes the parameters of the prompt by generating multiple visual expressions that describe different aspects of the object (color, shape, texture, pattern). This transforms a simple single-word prompt into a comprehensive multi-parameter description, improving detection accuracy while maintaining ease of use

Inventive Principle:
Principle #35Parameter changes

3Measurement precision

If multiple visual expressions are extracted from an image, then the prompt becomes more comprehensive and accurate, but the processing time and computational resources increase

Engineering Contradiction:
Improvedetection accuracyVSAvoidprocessing time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The system extracts only the necessary visual expressions needed for accurate detection, avoiding unnecessary over-extraction. By using the expression extraction model to identify and select relevant visual features, the system achieves sufficient accuracy without excessive processing time

Inventive Principle:
Principle #16Partial or excessive action

Data Source

PatentUS20250225762A1Information processing apparatus, information processing method, and recording medium
Publication Date: 2025.07.10 NEC CORP
  • US20250225762A1 patent drawing
  • US20250225762A1 patent drawing
  • US20250225762A1 patent drawing

AI summary

An information processing apparatus includes: a text group obtaining section which obtains a visually expressing text group that includes a plurality of texts which visually express a detection target, with reference to input data that specifies the detection target; a prompt generating section which generates a prompt with reference to the visually expressing text group; and a providing section which provides, to a detection model, the prompt that has been generated by the prompt generating section, the detection model being a model into which a prompt and an image are inputted and which detects, from the image, a detection target that is specified by the prompt.