Video Content Aggregation Using Training Model Filtering

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

The challenge lies in efficiently aggregating and reducing the volume of video content from multiple surveillance cameras in smart home systems, as users face a cumbersome task in reviewing extensive footage, with existing systems failing to effectively filter out ordinary events, leading to information overload.

Innovation Solution

A method and system that utilize a training model to categorize and filter video content, applying image processing techniques and user-generated input to discard non-essential footage, allowing users to review only significant events, with the ability to adjust and refine the training model based on user feedback and characteristics of the premises.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If video content from multiple cameras is aggregated and presented to users, then complete surveillance coverage is achieved, but users face information overload and cumbersome review tasks

Engineering Contradiction:
Improvesurveillance coverageVSAvoiduser review task
Core Design Contradiction:
ReliabilityVSEase of operation

Solution Approach 1:

The system extracts and separates significant events from the aggregated video content using training models. Ordinary events are filtered out and removed from the presentation to users, while only significant events are retained and displayed. This extraction process resolves the contradiction by maintaining complete surveillance coverage through aggregation while eliminating information overload in user presentation.

Inventive Principle:
Principle #2Taking out (Extraction)

2Reliability

If all video content is presented to users for review, then no important events are missed, but users spend excessive time reviewing footage

Engineering Contradiction:
Improveevent detection completenessVSAvoidvideo review time
Core Design Contradiction:
ReliabilityVSLoss of time

Solution Approach 1:

The system performs preliminary classification of video content using training models before presenting it to users. The training models pre-identify and categorize significant events, filtering out ordinary events in advance. This preliminary action ensures that no important events are missed while dramatically reducing the time users need to spend reviewing footage, as only pre-filtered significant events are presented.

Inventive Principle:
Principle #10Preliminary action

3Reliability

If existing systems present all video footage to users, then complete monitoring is maintained, but user experience deteriorates due to information overload

Engineering Contradiction:
Improvemonitoring completenessVSAvoiduser experience
Core Design Contradiction:
ReliabilityVSEase of operation

Solution Approach 1:

The system uses feedback from user interactions with significant events to continuously refine and improve the training models. When users review or interact with presented events, this feedback is used to adjust the training models, making them more accurate at identifying significant events. This feedback loop maintains monitoring completeness while progressively improving user experience by reducing information overload through more precise event filtering.

Inventive Principle:
Principle #23Feedback

Data Source

PatentUS11462051B2Method and system for aggregating video content
Publication Date: 2022.10.04 AT&T INTELLECTUAL PROPERTY I L P
  • US11462051B2 patent drawing
  • US11462051B2 patent drawing
  • US11462051B2 patent drawing

AI summary

Aspects of the subject disclosure may include, for example, systems and methods aggregating video content and adjusting the aggregate video content according to a training model. The adjusted aggregate video content comprises a first subset of the images and does not comprise a second subset of the images. The first subset of the images is determined by the training model based on a plurality of categories corresponding to a plurality of events. The illustrative embodiments also include presenting the adjusted aggregate video content and receiving identifications for the first subset of the images in the aggregate video content. Further, the illustrative embodiments include adjusting the training model according to the identifications and providing the adjusted training model to a network device. Other embodiments are disclosed.