Image Recognition Support Apparatus for Automatic State Labeling

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing image recognition systems face difficulties in accurately detecting the state or situation of objects in images, particularly in disaster scenarios, due to the lack of detailed labels and manual effort required for display adjustments as objects change in appearance or number.

Innovation Solution

An image recognition support apparatus that uses an object detection model and an image language model to generate expanded image queries and language queries, calculating similarity to assign attribute detail labels and display switching labels, enabling automatic detection and display of object states without extensive manual labeling.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If detailed labels for object states are prepared manually as learning data, then detection precision of object states is improved, but loss of time and productivity deteriorate due to extensive manual labeling effort

Engineering Contradiction:
Improvedetection precision of object statesVSAvoidtime for manual labeling
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The system automatically generates state labels by processing images through the object detection model and image language model, eliminating the need for manual labeling. The apparatus self-generates learning data by detecting objects and their states from images, thereby resolving the contradiction between detection precision and manual labeling time.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The system performs preliminary generation of state labels and learning data before actual detection tasks. By pre-processing images to generate object-state pairs automatically, the system prepares ready-to-use learning data without manual intervention, improving both detection precision and reducing time loss.

Inventive Principle:
Principle #10Preliminary action

2Adaptability or versatility

If all possible state labels are prepared in advance as learning data, then completeness of object state detection is improved, but device complexity deteriorates due to the need to maintain extensive label sets

Engineering Contradiction:
Improvecompleteness of object state detectionVSAvoidcomplexity of label management
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The system dynamically generates state labels based on the actual content of images rather than relying on a static, pre-defined label set. The image language model processes images to generate relevant state descriptions, allowing the system to adapt to any object state without maintaining an exhaustive label database, thereby reducing device complexity while maintaining completeness.

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

The image language model serves multiple functions: it generates state labels, creates learning data, and adapts to various object types and states. This multi-functional approach eliminates the need for separate label management systems for different object categories, reducing overall device complexity while maintaining comprehensive detection capability.

Inventive Principle:
Principle #6Universality (Multi-functionality)

3Measurement precision

If manual adjustment of display conditions is performed when object appearance changes, then accuracy of display adaptation is improved, but productivity deteriorates due to continuous manual intervention

Engineering Contradiction:
Improveaccuracy of display adaptationVSAvoidproductivity of display adjustment
Core Design Contradiction:
Measurement precisionVSProductivity

Solution Approach 1:

The system continuously monitors changes in object appearance and automatically adjusts display conditions based on detected state changes. The object detection model provides feedback on object states, and the display control unit automatically adapts the display accordingly, eliminating the need for manual intervention while maintaining accurate display adaptation.

Inventive Principle:
Principle #23Feedback

Solution Approach 2:

The apparatus automatically detects object state changes and adjusts display parameters without human intervention. The system self-manages the adaptation process by comparing detected states with display conditions and making necessary adjustments, thereby maintaining accuracy while significantly improving productivity.

Inventive Principle:
Principle #25Self-service

4Reliability

If extensive learning data with detailed labels is collected, then reliability of object state detection is improved, but loss of substance deteriorates due to large data storage requirements

Engineering Contradiction:
Improvereliability of object state detectionVSAvoiddata storage capacity
Core Design Contradiction:
ReliabilityVSLoss of substance

Solution Approach 1:

The system generates learning data on-demand through automatic processing rather than storing extensive pre-collected labeled datasets. By performing preliminary generation of object-state pairs when needed, the system maintains reliable detection performance without requiring large storage capacities for extensive pre-collected data.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The apparatus self-generates learning data by processing images through the detection models, eliminating the need to store extensive pre-collected labeled datasets. This on-demand generation approach maintains detection reliability while minimizing data storage requirements, as the system creates only the necessary learning data for current tasks.

Inventive Principle:
Principle #25Self-service

Data Source

PatentUS20240029399A1Image recognition support apparatus and image recognition support method
Publication Date: 2024.01.25 HITACHI LTD
  • US20240029399A1 patent drawing
  • US20240029399A1 patent drawing
  • US20240029399A1 patent drawing

AI summary

An image recognition support apparatus includes: an image acquisition unit that acquires an image; an image recognition unit that detects an object included in the image using an object detection model; and a detection result processing unit that generates one or more expanded image queries indicating a partial image of the image including the object, and sets a combination having a high similarity calculated using an image language model trained on a relationship between an image and an attribute including a state or situation among combinations of an expanded image query and an expanded language query indicating one or more language labels, as a detected object and an attribute detail label of the object.