Human Behavior Recognition Using Sliding Windows and Multiple Cameras

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Single cameras often have blind angles, leading to inaccuracies in capturing and anticipating human behavior due to incomplete video capture, which limits the effectiveness of video surveillance in preventing events.

Innovation Solution

A method and apparatus for human behavior recognition using multiple cameras to capture human behavior videos, where start and end points of human motions are extracted as sliding windows, and pre-trained models determine whether these windows represent motion sections, allowing for accurate anticipation of motion categories without human intervention.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Device complexity

If a single camera is used to capture human behavior video, then the device complexity is reduced, but the measurement precision and reliability of human behavior recognition deteriorate due to blind angles

Engineering Contradiction:
Improvecamera system complexityVSAvoidhuman behavior recognition accuracy
Core Design Contradiction:
Device complexityVSMeasurement precision

Solution Approach 1:

The patent combines multiple camera views into a unified human behavior recognition system. Video streams from multiple cameras are merged and processed together, allowing the system to overcome blind angles and improve recognition accuracy by synthesizing information from different perspectives.

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The human behavior recognition system is designed to process video inputs from multiple cameras simultaneously, making it multi-functional in terms of input sources. The same recognition algorithm can analyze behavior patterns regardless of which camera captures them, enhancing overall system reliability.

Inventive Principle:
Principle #6Universality (Multi-functionality)

2Measurement precision

If multiple cameras are used to capture human behavior video, then the measurement precision and reliability of human behavior recognition improve by covering blind angles, but the device complexity increases

Engineering Contradiction:
Improvehuman behavior recognition accuracyVSAvoidcamera system complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent segments the monitoring space into multiple camera fields of view, with each camera responsible for capturing specific zones. This segmentation allows comprehensive coverage while maintaining manageable system complexity by assigning specific detection responsibilities to individual cameras.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The system introduces an intermediary processing layer that receives video streams from multiple cameras, performs synchronization and integration, and outputs unified behavior recognition results. This intermediary layer manages the complexity of multiple inputs by providing a standardized processing interface.

Inventive Principle:
Principle #24Intermediary (Mediator)

3Measurement precision

If manual annotation is used to create training samples for motion sections, then the measurement precision of motion category classification improves, but the loss of time and productivity deteriorate

Engineering Contradiction:
Improvemotion category classification accuracyVSAvoidtraining data preparation time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The system implements self-service through automated labeling of training samples. The behavior time span discriminating model automatically identifies and labels motion sections in video data, eliminating the need for manual annotation while maintaining high classification accuracy through iterative training and validation.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The patent performs preliminary action by pre-processing video data to automatically generate candidate motion sections before formal training. This preliminary segmentation and filtering of training samples reduces the overall time required for data preparation while ensuring high-quality input for the classification model.

Inventive Principle:
Principle #10Preliminary action

4Measurement precision

If the sliding window size is increased to capture more motion information, then the measurement precision of motion detection improves, but the loss of time increases due to longer processing windows

Engineering Contradiction:
Improvemotion detection accuracyVSAvoidprocessing time per motion section
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The patent implements dynamic sliding window adjustment where the window size adapts based on the detected motion characteristics. For fast movements, smaller windows are used to reduce processing time, while for slower, more complex motions, larger windows capture sufficient detail. This dynamic adjustment optimizes both accuracy and processing efficiency.

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

The system changes the temporal parameter of the sliding window based on the specific motion context. By adjusting window duration and step size according to motion speed and complexity, the system achieves high detection accuracy without excessive processing time penalties.

Inventive Principle:
Principle #35Parameter changes

Data Source

PatentUS11321966B2Method and apparatus for human behavior recognition, and storage medium
Publication Date: 2022.05.03 BAIDU ONLINE NETWORK TECH (BEIJIBG) CO LTD
  • US11321966B2 patent drawing
  • US11321966B2 patent drawing
  • US11321966B2 patent drawing

AI summary

A method and an apparatus for human behavior recognition, and a storage medium, the method includes: obtaining a human behavior video captured by a camera; extracting a start point and an end point of a human motion from the human behavior video, where the human motion between the start point and the end point corresponds to a sliding window; determining whether the sliding window is a motion section; and if the sliding window is a motion section, anticipating a motion category of the motion section using a pre-trained motion classifying model. Thus, accurate anticipation of a motion in a human behavior video captured by a camera is realized without human intervention.