Automated Video Summarization via Scene Partitioning and Similarity
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
The existing methods for generating video summaries are time-consuming and impractical for large volumes of video content, as they require manual processing and do not efficiently allow for the creation of representative and interesting summaries that can motivate viewers to watch or purchase the original video.
Innovation Solution
A system that automatically summarizes videos by partitioning them into scenes, determining similarities between scenes, and using dynamic-programming techniques to select representative scenes based on similarity scores and time constraints, while also considering audio breaks and feature vectors extracted from frames.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If manual processing is used to generate video summaries, then the quality and representativeness of summaries can be ensured, but the time consumption and processing cost increase significantly
Solution Approach 1:
The patent replaces manual mechanical processing with an automated computer-based system that uses feature extraction, shot boundary detection, and dynamic programming algorithms to generate video summaries automatically, eliminating the need for human operators while maintaining summary quality
Solution Approach 2:
The system performs self-service by automatically analyzing video content, detecting shot boundaries, computing scene similarities, and selecting representative frames without external human intervention, enabling the system to generate summaries independently and efficiently
2Productivity
If automated methods are used to generate video summaries, then processing efficiency is improved, but the accuracy and representativeness of summaries may deteriorate
Solution Approach 1:
The patent segments the video into shots and scenes based on detected shot boundaries, allowing the automated system to process and analyze distinct video segments separately, which improves both processing efficiency and the accuracy of summary generation by focusing on representative segments
Solution Approach 2:
The system changes parameters by extracting multiple features (color histograms, motion vectors, audio characteristics) and using dynamic programming to optimize scene selection based on similarity metrics, enabling automated methods to achieve high summary accuracy while maintaining processing efficiency
3Measurement precision
If detailed feature analysis is performed on all frames, then the representativeness of selected scenes is improved, but the computational complexity and processing time increase
Solution Approach 1:
The patent extracts only the most relevant features (color histograms, motion vectors, audio breaks) from video frames rather than analyzing all frame data, and extracts features only at shot boundaries and scene transitions, reducing computational complexity while maintaining scene representativeness
Solution Approach 2:
The system performs detailed feature analysis only on selected key frames at shot boundaries and scene transitions rather than all frames, applying partial action to the most critical segments, which reduces computational complexity while preserving the accuracy of scene representation
Data Source
AI summary
One embodiment of the present invention provides a system that automatically produces a summary of a video. During operation, the system partitions the video into scenes and then determines similarities between the scenes. Next, the system selects representative scenes from the video based on the determined similarities, and combines the selected scenes to produce the summary for the video.


