Scene Clip Extraction via Pre-generated Time Metadata
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Users face difficulties in accurately locating and sharing specific video scenes from media data due to linear playback and the complexity of controlling time axes, leading to increased time spent searching and operational troubles.
Innovation Solution
A system and method that utilize scene time information to extract and align media clips, involving media supply equipment, a metadata server, and a scene server, allowing end devices to input capture time information to extract target scene times and form media division data, simplifying sharing and reducing operational complexity.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If users manually search for video scenes using time axis control in linear playback mode, then they can locate specific scenes, but the time required increases significantly and operational complexity increases
Solution Approach 1:
The system performs preliminary analysis of media data to pre-generate scene time information and scene clips before user requests. Scene time information is extracted in advance using scene detection algorithms, and scene clips are pre-processed and stored with their corresponding time stamps. When users want to locate a scene, the system can quickly retrieve pre-processed scene information rather than requiring manual searching through the entire media data, significantly reducing the time to locate specific scenes while maintaining high accuracy.
2Ease of operation
If users manually control time axis to find specific playback time points, then they can access desired clips, but operational complexity and difficulty increase
Solution Approach 1:
The system extracts and separates scene time information from the original media data, creating independent scene metadata that can be queried and accessed separately. Instead of requiring users to interact with the complex time axis control mechanism of the entire media file, the system extracts specific scene time points and presents them as simple, discrete access points. Users can directly select from extracted scene information without needing to understand or manipulate the underlying complex time control mechanisms.
Solution Approach 2:
The system introduces scene time information as an intermediary layer between users and the original media data. This intermediary contains pre-processed scene detection results, time stamps, and scene descriptions that simplify user interaction. Instead of directly controlling the complex media player time axis, users interact with the simplified scene time information interface, which then translates user requests into precise time point access in the original media data, reducing operational complexity.
3Ease of operation
If users share media clips using traditional media capture software, then they can capture specific scenes, but software acquisition and operation become troublesome
Solution Approach 1:
The system creates and stores scene clips as separate, self-contained data copies extracted from the original media data. These scene clips include the actual video/audio segments along with their metadata (time stamps, scene descriptions). Instead of requiring users to use media capture software to截取 and share clips, the system has already created ready-to-share scene clip copies that can be directly transmitted and distributed. Users can share these pre-extracted scene clips without needing additional capture software, simplifying the sharing process.
Data Source
AI summary
A system and method for constructing a scene clip, and a non-statutory record medium thereof are provided. The system includes media supply equipment, a metadata server, a scene server, and an end device. The media supply equipment is used for providing media data. The metadata server is used for providing scene time information corresponding to playback scenes of the media data. A first end device acquires the media data and the scene time information, and extracts, according to capture time information input when playing the media data, at least one piece of target scene time from each piece of the scene time information. The scene server acquires the media data and the target scene time, and according to an alignment result of the target scene time and each piece of the scene time information, extracts local scene clips from the media data to form a piece of media division data.


