Text-Based Anomaly Detection for Robust Surveillance Imaging
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Unsupervised anomaly detection in surveillance footage is challenging due to variations in imaging environments such as camera position, angle, and lighting conditions, leading to reduced accuracy in detecting anomalies.
Innovation Solution
An anomaly detection apparatus that converts image data into text using a trained model, calculates statistics on the text, and uses a text information dictionary to determine anomalies based on the degree of text appearance, providing robustness against environmental changes.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Ease of manufacture
If unsupervised anomaly detection is performed using only normal images for training, then the need for anomaly images and annotations during training is eliminated and unknown anomalies can be detected, but detection accuracy deteriorates when there are changes in imaging environment such as camera position, angle, or lighting conditions
Solution Approach 1:
The patent introduces text descriptions as an intermediary layer between images and anomaly detection. Images are converted to text representations, and anomaly detection is performed on text rather than directly on images. This intermediary text representation is invariant to imaging environment changes, thereby maintaining detection accuracy while preserving the simplicity of unsupervised training using only normal images.
2Loss of information
If image data is directly analyzed for anomaly detection, then detailed visual information is preserved, but sensitivity to environmental variations increases leading to reduced detection accuracy
Solution Approach 1:
Text descriptions serve as an intermediary that preserves essential semantic information while being invariant to environmental variations. The text representation captures the content and context of images without being sensitive to camera position, angle, or lighting conditions, thus improving reliability while maintaining reasonable information retention.
Solution Approach 2:
The patent transforms the representation parameters from pixel-based image data to text-based semantic descriptions. This parameter change from visual features to linguistic features fundamentally alters the detection space, making it insensitive to environmental parameter changes while preserving the ability to detect anomalies in content.
Data Source
AI summary
According to one embodiment, an anomaly detection apparatus includes a processor. The processor acquires a first sample that is a subject for anomaly detection. The processor generates, using a trained model, a first text from the first sample. The first text represents a content of the first sample. The processor determines whether the first sample has an anomaly based on a statistic associated with all or a part of the first text in a dictionary. The dictionary associates all or a part of a second text representing a content of a second sample included in a training data set with a statistic related to a degree of appearance of all or a part of the second text in the training data set. The processor outputs a determination result.


