Digest Video Production via Scene Evaluation and Frame Extraction
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing video processing technologies are limited in producing digest video data that includes numerous scenes within a desired playback time, as the number of scenes is constrained by the desired playback time.
Innovation Solution
A video processing device that splits original video data into scenes, calculates frame and scene evaluation levels, determines scene playback time, and extracts frame images to produce digest video data, allowing for flexible inclusion of multiple scenes based on their importance and desired playback time.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Quantity of substance
If the number of scenes in digest video data is increased, then the comprehensiveness of video summary is improved, but the playback time exceeds the desired playback time
Solution Approach 1:
The patent applies local quality by differentiating scene importance through evaluation levels. Different scenes are assigned different weights based on their characteristics (e.g., presence of key objects, motion intensity, scene transitions). This allows the system to selectively allocate playback time to important scenes while compressing or omitting less important ones, thereby including more scenes within the desired playback time constraint.
Solution Approach 2:
The patent changes the parameter of scene representation by extracting key frames or representative frames from each scene rather than using all frames. The number of frames extracted per scene is dynamically adjusted based on the scene's evaluation level and the remaining playback time budget. This parameter change enables denser scene inclusion while maintaining quality.
2Loss of information
If all scenes are included in digest video data, then the completeness of video summary is improved, but the playback time becomes unmanageably long
Solution Approach 1:
The patent extracts essential information from each scene by identifying and retaining only the most representative frames based on evaluation criteria. Instead of including all frames from all scenes, the system extracts key moments that capture the essence of each scene. This extraction process maintains information completeness while dramatically reducing the total playback time.
Solution Approach 2:
The patent segments the video content into discrete scenes and further segments each scene into selectable frame units. This hierarchical segmentation allows flexible recombination of scene segments to create a digest that covers all original scenes but with reduced temporal duration, achieving both completeness and time efficiency.
3Manufacturing precision
If important scenes are given more playback time, then the quality of important scene representation is improved, but the total number of scenes that can be included decreases
Solution Approach 1:
The patent dynamically changes the parameter of frames per scene based on the scene's evaluation level. Important scenes (high evaluation level) are allocated more frames and thus more playback time, while less important scenes receive fewer frames. This parameter adjustment resolves the contradiction by making the quality distribution adaptive rather than uniform, allowing more total scenes to be included while preserving quality for important ones.
Data Source
AI summary
The video processing device calculates evaluation levels of frame images in video data. Scene evaluation levels are then calculated from the frame evaluation levels in each scene. The playback times for the digest video of each scene are determined from the scene evaluation levels. Collections of frame image data with the playback time are extracted from each scene and are combined to produce digest video of the desired playback time.


