Semantic Frame Extraction From Video Using Text Similarity
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing techniques fail to appropriately extract important frame images from moving images, as these images often have minimal pixel value changes within scenes, leading to misidentification of scene breaks.
Innovation Solution
A method involving the acquisition of semantic vectors from frame images and text representing the moving image content, followed by similarity calculation and extraction of frame images based on predetermined conditions, to identify and extract relevant frame images.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If frame images with great pixel value changes are used to detect scene breaks, then scene break detection is improved, but important frame images in the middle of scenes are missed
Solution Approach 1:
The patent transforms frame images into semantic vectors, changing the parameter representation from raw pixel values to semantic features. This allows the system to evaluate frame importance based on semantic content rather than pixel change magnitude, resolving the contradiction between scene break detection and important frame extraction.
Solution Approach 2:
The patent introduces semantic vectors as an intermediary between frame images and the evaluation process. These vectors serve as a mediator that captures the semantic meaning of frames, enabling accurate identification of important frames without relying on pixel value changes that may miss semantically significant content.
2Measurement precision
If semantic vector comparison with text is performed, then frame image extraction accuracy is improved, but processing complexity increases
Solution Approach 1:
The patent replaces traditional mechanical image processing methods (pixel value comparison, edge detection) with semantic vector-based processing. This substitution uses mathematical operations on vector representations rather than complex image analysis algorithms, reducing computational complexity while improving accuracy.
Solution Approach 2:
By changing the representation parameters from raw image data to compressed semantic vectors, the patent reduces the dimensionality and complexity of processing while maintaining or improving extraction accuracy. The semantic vectors capture essential information in a more efficient format.
Data Source
AI summary
A non-transitory computer readable storage medium includes a program that causes a hardware processor on a computer to perform: acquiring a plurality of first semantic vectors generated based on a plurality of frame images of a moving image and at least one second semantic vector generated based on a text representing content of the moving image; calculating a similarity between each of the plurality of first semantic vectors and each of the at least one second semantic vector; and specifying, from among the plurality of first semantic vectors, the first semantic vector for which the similarity satisfying a predetermined condition has been calculated, and extracting, from among the plurality of frame images, the frame image used for generating the specified first semantic vector.


