State Determination Apparatus Using Visual Question Answering for Anomaly Detection
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
There is a need for an efficient method to detect equipment or dangerous states in manufacturing or maintenance sites based on images captured by cameras, ensuring compliance with safety manuals, which existing technologies have not adequately addressed.
Innovation Solution
A state determination apparatus using a processor that acquires images, generates questions and expected answers, and employs a trained visual question answering model to determine the similarity between estimated and expected answers, identifying anomalies and presenting remedial measures.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If a trained visual question answering model is used to detect equipment or dangerous states, then detection accuracy is improved, but device complexity increases
Solution Approach 1:
The visual question answering model is trained in advance with question-answer pairs related to equipment states and safety conditions. This preliminary training enables the model to perform complex detection tasks during operation without requiring complex real-time processing systems, thereby improving detection accuracy while maintaining manageable device complexity.
Solution Approach 2:
The patent introduces an answer evaluation unit that acts as an intermediary between the VQA model and the final detection result. This unit compares the model's estimated answers with expected answers to determine equipment or dangerous states, simplifying the overall system architecture while maintaining high detection accuracy through structured comparison logic.
2Reliability
If safety manuals are updated to reflect new safety requirements, then reliability is improved, but device complexity increases due to system reconfiguration
Solution Approach 1:
The system dynamically adapts to updated safety manuals by allowing modification of question-answer pairs in the training data. When safety requirements change, the VQA model can be retrained with updated questions and expected answers, enabling the system to maintain high reliability without requiring complex structural reconfiguration. The flexible data structure allows easy updates to safety criteria.
3Measurement precision
If comprehensive image analysis is performed to detect all potential dangers, then detection accuracy is improved, but processing time increases
Solution Approach 1:
The patent segments the image analysis process by dividing it into distinct question-answer evaluation steps. Instead of performing monolithic comprehensive analysis, the system breaks down detection into multiple focused VQA tasks, each addressing specific safety aspects. This segmentation enables parallel processing of different safety checks, improving overall detection accuracy while reducing total processing time through efficient task distribution.
Data Source
AI summary
According to one embodiment, a state determination apparatus includes a processor. The processor acquires a targeted image. The processor acquires a question concerning the targeted image and an expected answer to the question. The processor generates an estimated answer estimated with respect to the question concerning the targeted image using a trained model trained to estimate an answer based on a question concerning an image. The processor determines a state of a target for determination in accordance with a similarity between the expected answer and the estimated answer.


