Video Digest Event Extraction with Contextual Segmentation
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing video digest generation techniques fail to effectively extract and include relevant event segments before and after a specific event in a digest video, making it difficult for viewers to understand the context of the event.
Innovation Solution
An information processing device and method that acquire video materials and digest videos, detect coincident segments, and generate training data to create an event segment detection model, which identifies and clips event segments from the video material for inclusion in the digest video, ensuring contextual understanding.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If only the main event scene is extracted from video material, then the extraction precision is high, but the contextual understanding is lost
Solution Approach 1:
The patent segments the video material into multiple regions: the main event scene and surrounding contextual scenes. By dividing the video content into these functional segments, the system can extract both the precise main event and the necessary context, resolving the contradiction between extraction precision and context preservation.
Solution Approach 2:
The system performs preliminary action by pre-defining event templates that specify not only the main event characteristics but also the required contextual information before and after the event. This preliminary structuring of event requirements ensures that both precise event detection and context preservation are achieved systematically.
2Measurement precision
If event templates include detailed scene requirements, then the event detection accuracy improves, but the template complexity increases
Solution Approach 1:
The event template is segmented into distinct components: event definition, pre-event context requirements, post-event context requirements, and detection parameters. This segmentation allows the template to maintain high detection accuracy while organizing complexity into manageable, structured sections that can be processed systematically.
3Reliability
If the entire video material is processed to ensure complete event coverage, then the event coverage is complete, but the processing time increases
Solution Approach 1:
The system performs preliminary action by pre-processing the video material to identify and mark potential event regions based on basic features before applying the full event detection algorithm. This preliminary filtering reduces the amount of video data that requires intensive processing, maintaining complete event coverage while significantly reducing overall processing time.
Solution Approach 2:
The video processing is segmented into multiple stages: preliminary scanning for event candidates, detailed detection using event templates, and contextual scene extraction. This multi-stage segmentation allows the system to process only relevant portions of the video in detail, ensuring complete event coverage while minimizing total processing time through efficient resource allocation across different processing stages.
Data Source
AI summary
In an information processing device, an acquisition means acquires a plurality of videos including a video material and a digest video. A coincident segment detection means detects each coincident segment where the video material and the digest video match with each other in content. A training data generation means generates training data from the video material based on the coincident segment.


