Video Cut Grouping Using Temporal Context and Feature Extraction
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing video processing technologies fail to appropriately group video cuts, leading to an unclear understanding of the video structure, especially when cuts are not similarly grouped.
Innovation Solution
A moving image processing apparatus and method that determines the similarity between video cuts using feature amounts generated from extraction images, grouping similar cuts together and generating feature amounts for dissimilar cuts to enhance understanding of the video structure by prioritizing images with later time codes.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If cuts are grouped based on traditional similarity methods, then the grouping process is simple, but the cut structure understanding becomes unclear when cuts are not appropriately grouped
Solution Approach 1:
The patent changes the parameter used for cut grouping from traditional visual similarity to temporal position-based similarity. By using time code information and comparing temporal positions of cuts, the system achieves more accurate cut structure understanding while maintaining a relatively simple grouping process. The feature amount generation unit generates feature amounts based on temporal positions rather than complex visual analysis.
2Loss of information
If all images in each cut group are processed equally, then the processing is straightforward, but temporal context information is lost
Solution Approach 1:
The patent applies local quality by differentiating the processing of images based on their temporal positions. The image extraction unit preferentially extracts images with later time codes from each cut group, giving different weights to different images within the same cut group. This ensures that temporal context information is preserved while maintaining efficient processing by not treating all images equally.
3Reliability
If extraction images are selected without considering time codes, then the selection process is simple, but the temporal sequence and context of video content is not preserved
Solution Approach 1:
The patent applies preliminary action by pre-establishing the time code information for each image before the extraction process. The image extraction unit uses this pre-existing time code information to preferentially select images with later time codes, ensuring temporal sequence preservation without requiring complex real-time analysis during extraction.
Data Source
AI summary
There is provided a moving image processing apparatus, including a similarity determination unit configured to determine a degree of similarity between a subsequent cut and first and second cut groups based on feature amounts generated from extraction images of the first cut group included in a moving image and feature amounts generated from extraction images of the second cut group included in the moving image, a cut grouping unit configured to group the subsequent cut into a similar cut group similar to the subsequent cut, which is one of the first and second cut groups, when the subsequent cut is similar to the first or second cut group, and group the subsequent cut into a third cut group when the subsequent cut is not similar to either of the first and second cut groups, a feature amount generation unit configured to compare extraction images extracted from the third cut group with the extraction images extracted from the first and second cut groups when the subsequent cut is not similar to either of the first and second cut groups, and generate feature amounts of the third cut group, and an image extraction unit configured to preferentially extract an image with a later time code of the moving image from images included in each cut group, thereby obtaining extraction images of each cut group.


