Multi-Frame Video Behavior Recognition for Efficient CCTV Event Detection

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

The challenge of effectively detecting event occurrences in videos from CCTV footage without wasting resources on constant human monitoring, as existing methods are inefficient and costly.

Innovation Solution

A video-based behavior recognition device that synthesizes channel frames from multiple channels using a multi-frame convolution neural network to generate a behavior recognition result, utilizing a gray synthesized frame and weighted values for enhanced detection.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If constant human monitoring of CCTV videos is performed, then event detection capability is improved, but resource consumption and cost increase

Engineering Contradiction:
Improveevent detection capabilityVSAvoidresource consumption
Core Design Contradiction:
Measurement precisionVSLoss of energy

Solution Approach 1:

The patent replaces the mechanical system of human monitoring with an automated computer-based system that uses image processing and neural networks to detect events in CCTV videos, eliminating the need for continuous human observation while maintaining detection capability

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Solution Approach 2:

The system enables self-service by automatically analyzing video frames, detecting events, and generating alerts without requiring human intervention, allowing the monitoring system to serve itself through automated intelligence

Inventive Principle:
Principle #25Self-service

2Measurement precision

If traditional multi-channel video processing is used, then comprehensive behavior analysis is achieved, but computational complexity increases

Engineering Contradiction:
Improvebehavior analysis accuracyVSAvoidcomputational complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent segments the complex video processing task into distinct components: frame extraction, channel frame generation for different color channels, highlight information extraction, and neural network processing, allowing each component to be optimized independently

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent processes video data across multiple color channels (RGB) and temporal dimensions simultaneously, creating channel frames that capture information from different spectral and temporal perspectives, thereby analyzing behavior comprehensively without exponentially increasing computational burden

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

3Measurement precision

If highlight information is extracted and synthesized from multiple channels, then event detection accuracy is improved, but processing time increases

Engineering Contradiction:
Improveevent detection accuracyVSAvoidprocessing time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The patent performs preliminary extraction of highlight information from each color channel before synthesis, preparing processed data in advance that can be quickly combined and analyzed, reducing the time required for final event detection

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent merges highlight information from multiple color channels into a synthesized channel frame that consolidates relevant event indicators, allowing the system to process integrated information more efficiently than analyzing separate channels independently

Inventive Principle:
Principle #5Merging (Combining)

Data Source

PatentUS12573200B2Video-based behavior recognition device and operation method therefor
Publication Date: 2026.03.10 SOGANG UNIV RES & BUSINESS DEV FOUND
  • US12573200B2 patent drawing
  • US12573200B2 patent drawing
  • US12573200B2 patent drawing

AI summary

The present disclosure provides a video-based behavior recognition device comprising a synthesized channel frame provision unit generating highlight information by comparing channel frames corresponding to the respective channels among a plurality of channels and synthesizing the channel frames and the highlight information to provide a synthesized channel frame, a neural network unit providing a middle frame on the basis of the synthesized channel frame and a multi-frame convolution neural network, and a behavior recognition result provision unit providing a behavior recognition result on the basis of the middle frame and a weighted value generated according to the middle frame, and a method operating thereof. In the present disclosure, the behavior recognition result is provided on the basis of the multi-frame convolution neural network and the synthesized channel frame synthesized from the channel frames provided for the respective channels, and, thereby, an event occurrence in a video is more effectively detected.