Primitive Visual Knowledge Extraction from CCTV Video

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current technologies for intelligent CCTV systems face challenges in managing massive image data, including false alarms, storage, processing, and analysis, due to the lack of efficient methods for extracting and analyzing important information from continuous video feeds.

Innovation Solution

An apparatus and method that divide images into scenes, extract representative shots, identify key frames, generate primitive visual knowledge by classifying objects and actions, and store this information in a database for easy retrieval and visualization, allowing managers to focus on critical events.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Loss of information

If continuous video feeds are stored and processed in full detail, then complete information is preserved, but storage requirements and processing time increase significantly

Engineering Contradiction:
Improveinformation completenessVSAvoiddata volume
Core Design Contradiction:
Loss of informationVSQuantity of substance

Solution Approach 1:

The system extracts only essential visual information from continuous video feeds by identifying key frames that contain significant events or changes. Instead of storing and processing all video data, the system selectively extracts representative frames and their associated metadata, dramatically reducing data volume while preserving critical information for analysis and retrieval.

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The system creates simplified representations (copies) of video content in the form of metadata structures that capture essential information about objects, actions, and events. These metadata copies enable efficient storage and processing while maintaining the ability to retrieve and visualize important visual information without handling the full video data.

Inventive Principle:
Principle #26Copying

2Reliability

If all video data is processed for analysis, then comprehensive event detection is achieved, but processing time and computational resources increase

Engineering Contradiction:
Improveevent detection accuracyVSAvoidprocessing time
Core Design Contradiction:
ReliabilityVSLoss of time

Solution Approach 1:

The system segments video data into discrete scenes and identifies key frames within each scene that contain significant events. By dividing the continuous video stream into manageable segments and focusing processing efforts only on key frames rather than all frames, the system achieves comprehensive event detection with significantly reduced processing time and computational resources.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The system performs preliminary analysis to identify key frames and extract metadata before detailed processing is applied. By pre-identifying which frames contain important information and organizing data into structured metadata formats in advance, the system prepares the data for efficient processing and retrieval, reducing the time required for comprehensive analysis.

Inventive Principle:
Principle #10Preliminary action

3Loss of information

If detailed metadata is generated for all video content, then precise event information is available, but storage and retrieval complexity increase

Engineering Contradiction:
Improveevent information precisionVSAvoidsystem complexity
Core Design Contradiction:
Loss of informationVSDevice complexity

Solution Approach 1:

The system employs a universal metadata structure that can represent multiple types of visual information (objects, actions, events, scenes) in a unified format. This standardized metadata framework enables precise event information to be captured and stored efficiently, while the consistent structure simplifies retrieval and processing operations compared to handling multiple specialized data formats.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Data Source

PatentUS9898666B2Apparatus and method for providing primitive visual knowledge
Publication Date: 2018.02.20 ELECTRONICS & TELECOMM RES INST
  • US9898666B2 patent drawing
  • US9898666B2 patent drawing
  • US9898666B2 patent drawing

AI summary

An apparatus and method for providing primitive visual knowledge are disclosed. The method of providing primitive visual knowledge includes receiving an image in a form of a digital image sequence, dividing the received image into scenes, extracting a representative shot from each of the scenes, extracting objects from frames which compose the representative shot, extracting action verbs based on a mutual relationship between the extracted objects, selecting a frame best expressing the mutual relationship with the objects, which are the basis for the extracting of the action verbs, as a key frame, generating the primitive visual knowledge based on the selected key frame, storing the generated primitive visual knowledge in a database, and visualizing the primitive visual knowledge stored in the database to provide the primitive visual knowledge to a manager.