Intelligent Media Content Segmentation Using Temporal Continuity
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional video content analysis methods are inadequate in segmenting and analyzing media content effectively, particularly in identifying and grouping shots, sub-scenes, and scenes, which hampers the understanding and navigation of complex video programs.
Innovation Solution
A system and method for intelligent media content segmentation and analysis, utilizing a computing device with modules for shot identification, video grouping, and time code management, which automatically or manually annotates and groups frames into shots, sub-scenes, and scenes based on similarity and temporal continuity, enabling efficient media content hierarchy creation and display.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If conventional video content analysis methods are used, then basic motion and object recognition can be achieved, but effective segmentation and analysis of media content into narrative units is inadequate
Solution Approach 1:
The patent divides media content into hierarchical narrative units (shots, sub-scenes, scenes, acts) based on visual and audio features. This multi-level segmentation enables precise identification of narrative structures while maintaining adaptability to different content types through automated analysis of transitions, camera movements, and audio cues.
Solution Approach 2:
The system analyzes multiple parameters simultaneously including visual features (color, motion, camera angle), audio features (dialogue, music, sound effects), and temporal characteristics. By monitoring changes in these parameters, the system achieves both precise segmentation and versatile content analysis across different media types.
2Productivity
If automated shot identification and grouping is implemented, then productivity is improved, but system complexity increases
Solution Approach 1:
The system performs preliminary analysis of visual and audio features to pre-identify potential shot boundaries and narrative units before final grouping. This preliminary action enables rapid automated processing while the modular architecture manages complexity by separating feature extraction, boundary detection, and hierarchical grouping into distinct processing stages.
3Loss of information
If detailed annotation and supplemental information are provided, then content accessibility is enhanced, but information processing requirements increase
Solution Approach 1:
The system extracts and stores only the most relevant contextual information for each narrative unit, such as key visual descriptors, audio elements, and narrative relationships. This selective extraction preserves essential context while minimizing data volume by excluding redundant information.
Solution Approach 2:
The annotation structure is organized hierarchically with shots nested in sub-scenes, sub-scenes nested in scenes, and scenes nested in acts. This nested organization allows detailed information to be stored at appropriate levels of granularity, reducing overall data requirements by sharing common information at higher hierarchical levels.
Data Source
AI summary
There is provided a system including a non-transitory memory storing an executable code and a hardware processor executing the executable code to receive a media content including a plurality of frames, divide the media content into a plurality of shots, each of the plurality of shots including a plurality of frames of the media content based on a first similarity between the plurality of frames, determine a plurality of sequential shots of the plurality of shots to be part of a first sub-scene of a plurality of sub-scenes of a scene based on a timeline continuity of the plurality of sequential shots, identify each of the plurality of shots of the media content and each of the plurality of sub-scenes with a corresponding beginning time code and a corresponding ending time code.


