Text-Based Anomaly Detection for Robust Surveillance Imaging

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Unsupervised anomaly detection in surveillance footage is challenging due to variations in imaging environments such as camera position, angle, and lighting conditions, leading to reduced accuracy in detecting anomalies.

Innovation Solution

An anomaly detection apparatus that converts image data into text using a trained model, calculates statistics on the text, and uses a text information dictionary to determine anomalies based on the degree of text appearance, providing robustness against environmental changes.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Ease of manufacture

If unsupervised anomaly detection is performed using only normal images for training, then the need for anomaly images and annotations during training is eliminated and unknown anomalies can be detected, but detection accuracy deteriorates when there are changes in imaging environment such as camera position, angle, or lighting conditions

Engineering Contradiction:
Improvetraining process simplicityVSAvoidanomaly detection accuracy
Core Design Contradiction:
Ease of manufactureVSMeasurement precision

Solution Approach 1:

The patent introduces text descriptions as an intermediary layer between images and anomaly detection. Images are converted to text representations, and anomaly detection is performed on text rather than directly on images. This intermediary text representation is invariant to imaging environment changes, thereby maintaining detection accuracy while preserving the simplicity of unsupervised training using only normal images.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Loss of information

If image data is directly analyzed for anomaly detection, then detailed visual information is preserved, but sensitivity to environmental variations increases leading to reduced detection accuracy

Engineering Contradiction:
Improvevisual information retentionVSAvoiddetection robustness to environmental changes
Core Design Contradiction:
Loss of informationVSReliability

Solution Approach 1:

Text descriptions serve as an intermediary that preserves essential semantic information while being invariant to environmental variations. The text representation captures the content and context of images without being sensitive to camera position, angle, or lighting conditions, thus improving reliability while maintaining reasonable information retention.

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The patent transforms the representation parameters from pixel-based image data to text-based semantic descriptions. This parameter change from visual features to linguistic features fundamentally alters the detection space, making it insensitive to environmental parameter changes while preserving the ability to detect anomalies in content.

Inventive Principle:
Principle #35Parameter changes

Data Source

PatentUS20260073138A1Anomaly detection apparatus, method, and storage medium
Publication Date: 2026.03.12 KK TOSHIBA
  • US20260073138A1 patent drawing
  • US20260073138A1 patent drawing
  • US20260073138A1 patent drawing

AI summary

According to one embodiment, an anomaly detection apparatus includes a processor. The processor acquires a first sample that is a subject for anomaly detection. The processor generates, using a trained model, a first text from the first sample. The first text represents a content of the first sample. The processor determines whether the first sample has an anomaly based on a statistic associated with all or a part of the first text in a dictionary. The dictionary associates all or a part of a second text representing a content of a second sample included in a training data set with a statistic related to a degree of appearance of all or a part of the second text in the training data set. The processor outputs a determination result.