State Determination Apparatus Using Visual Question Answering for Anomaly Detection

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

There is a need for an efficient method to detect equipment or dangerous states in manufacturing or maintenance sites based on images captured by cameras, ensuring compliance with safety manuals, which existing technologies have not adequately addressed.

Innovation Solution

A state determination apparatus using a processor that acquires images, generates questions and expected answers, and employs a trained visual question answering model to determine the similarity between estimated and expected answers, identifying anomalies and presenting remedial measures.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If a trained visual question answering model is used to detect equipment or dangerous states, then detection accuracy is improved, but device complexity increases

Engineering Contradiction:
Improvedetection accuracyVSAvoiddevice complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The visual question answering model is trained in advance with question-answer pairs related to equipment states and safety conditions. This preliminary training enables the model to perform complex detection tasks during operation without requiring complex real-time processing systems, thereby improving detection accuracy while maintaining manageable device complexity.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent introduces an answer evaluation unit that acts as an intermediary between the VQA model and the final detection result. This unit compares the model's estimated answers with expected answers to determine equipment or dangerous states, simplifying the overall system architecture while maintaining high detection accuracy through structured comparison logic.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Reliability

If safety manuals are updated to reflect new safety requirements, then reliability is improved, but device complexity increases due to system reconfiguration

Engineering Contradiction:
Improvesafety complianceVSAvoidsystem reconfiguration
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The system dynamically adapts to updated safety manuals by allowing modification of question-answer pairs in the training data. When safety requirements change, the VQA model can be retrained with updated questions and expected answers, enabling the system to maintain high reliability without requiring complex structural reconfiguration. The flexible data structure allows easy updates to safety criteria.

Inventive Principle:
Principle #15Dynamics

3Measurement precision

If comprehensive image analysis is performed to detect all potential dangers, then detection accuracy is improved, but processing time increases

Engineering Contradiction:
Improvedetection accuracyVSAvoidprocessing time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The patent segments the image analysis process by dividing it into distinct question-answer evaluation steps. Instead of performing monolithic comprehensive analysis, the system breaks down detection into multiple focused VQA tasks, each addressing specific safety aspects. This segmentation enables parallel processing of different safety checks, improving overall detection accuracy while reducing total processing time through efficient task distribution.

Inventive Principle:
Principle #1Segmentation

Data Source

PatentUS12086210B2State determination apparatus and image analysis apparatus
Publication Date: 2024.09.10 KK TOSHIBA
  • US12086210B2 patent drawing
  • US12086210B2 patent drawing
  • US12086210B2 patent drawing

AI summary

According to one embodiment, a state determination apparatus includes a processor. The processor acquires a targeted image. The processor acquires a question concerning the targeted image and an expected answer to the question. The processor generates an estimated answer estimated with respect to the question concerning the targeted image using a trained model trained to estimate an answer based on a question concerning an image. The processor determines a state of a target for determination in accordance with a similarity between the expected answer and the estimated answer.