Video Synopsis Generation via Object Interaction Analysis

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing video synopsis generation methods often produce unwanted information, as they focus solely on object activities and appearances, failing to customize the synopsis according to user needs by not considering interactions between objects and the background.

Innovation Solution

A system and method that detect a source object, track its motion, generate a source object tube, and determine interaction with the background or other objects using pre-learned models, allowing for the generation of a video synopsis that highlights specific interactions based on user queries.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Device complexity

If video synopsis is generated based only on object activities and appearances, then the generation process is simple, but the synopsis includes many unnecessary objects and lacks customization

Engineering Contradiction:
Improvesynopsis generation processVSAvoidnecessary interaction information
Core Design Contradiction:
Device complexityVSLoss of information

Solution Approach 1:

The synopsis generation process is segmented into multiple analysis levels: low-level analysis (object detection, appearance extraction, activity recognition) and high-level analysis (interaction determination, scene understanding). This segmentation allows the system to process different types of information separately and combine them to generate customized synop ses that include both object activities and their interactions with the background or other objects.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

An interaction determination model serves as an intermediary between low-level object detection and high-level scene understanding. This intermediary component analyzes relationships between objects and the background, determining interactions such as 'person sitting on chair' or 'car driving on road'. The interaction information acts as a bridge that connects basic object activities to meaningful scene context, enabling customized synopsis generation.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Quantity of substance

If video synopsis includes all objects in the video, then comprehensive coverage is achieved, but user convenience and relevance are reduced

Engineering Contradiction:
Improvenumber of objects in synopsisVSAvoiduser convenience
Core Design Contradiction:
Quantity of substanceVSEase of operation

Solution Approach 1:

The system applies local quality by selectively including different types of information in the synopsis based on user needs. Instead of uniformly including all objects, the system identifies and emphasizes objects and interactions that are locally important to the user's query. For example, if a user queries about 'sitting' activities, the system prioritizes objects involved in sitting interactions while de-emphasizing unrelated objects, creating a customized synopsis with relevant local quality.

Inventive Principle:
Principle #3Local quality

3Measurement precision

If interaction analysis is added to object-based analysis, then scene understanding is improved, but system complexity increases

Engineering Contradiction:
Improvescene understanding accuracyVSAvoidanalysis system structure
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The system performs preliminary action by pre-processing video content into structured representations before interaction analysis. Objects are detected and tracked in advance, creating object tubes that contain temporal information. Background regions are segmented and labeled beforehand. These preliminary structures serve as inputs to the interaction determination model, reducing the complexity of real-time interaction analysis while maintaining high scene understanding accuracy.

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS11620335B2Method for generating video synopsis through scene understanding and system therefor
Publication Date: 2023.04.04 KOREA INST OF SCI & TECH
  • US11620335B2 patent drawing
  • US11620335B2 patent drawing
  • US11620335B2 patent drawing

AI summary

Embodiments relate to a method for generating a video synopsis including receiving a user query; performing an object based analysis of a source video; and generating a synopsis video in response to a video synopsis generation request from a user, and a system therefor. The video synopsis generated by the embodiments reflects the user's desired interaction.