Video Preview Generation via Image Recognition and Frame Filtering
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing methods for generating video preview content are inadequate in quality, particularly in discerning header and tail text content and rapid changes in video content, often requiring manual intervention to achieve high-quality previews.
Innovation Solution
A method and device for generating video preview content that involves parsing a video to obtain image frames, filtering out slice headers and tails using image recognition, and generating preview content based on filtered image frames, which can include calculating similarity between adjacent frames and determining a screening number for optimal preview content.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Manufacturing precision
If manual intervention is used to generate high-quality video previews, then the quality of preview content is improved, but the time consumption and operational complexity increase
Solution Approach 1:
The system performs self-service by automatically analyzing video content through image recognition and similarity calculation algorithms. The device autonomously identifies slice headers and tails, filters repetitive content, and generates preview images without requiring manual intervention, thereby maintaining high quality while reducing time consumption
Solution Approach 2:
Manual mechanical operations are replaced with automated computational systems. The patent substitutes manual video analysis with computer-based image recognition technology and automated similarity comparison algorithms, eliminating the need for human operators while achieving consistent high-quality results
2Measurement precision
If manual intervention is applied to discern header and tail text content, then the accuracy of content identification is improved, but the operational complexity and time required increase
Solution Approach 1:
Manual text identification is replaced with automated optical character recognition (OCR) technology. The system uses computer-based image recognition to automatically discern and identify text content in video frames, achieving high accuracy without requiring human operators to manually examine and transcribe text
Solution Approach 2:
An intermediary computational layer is introduced between the video content and the final preview generation. The patent employs image recognition algorithms as an intermediary that automatically analyzes and identifies text content, serving as a bridge that translates visual text information into actionable data for preview creation
3Loss of information
If all image frames are used for preview generation, then the completeness of video representation is improved, but the file size and processing time increase
Solution Approach 1:
The system extracts only the essential and non-repetitive image frames from the video sequence. By identifying and removing slice headers, tails, and duplicate content through image recognition and similarity comparison, the patent extracts a optimized subset of frames that represent the video content completely but with reduced quantity, improving generation efficiency
Solution Approach 2:
The patent applies parameter changes by dynamically adjusting the selection criteria for image frames based on similarity thresholds and content importance. By changing the parameters of frame selection (using similarity calculations and text content analysis), the system optimizes which frames are included in the preview, maintaining completeness while reducing overall frame count
Data Source
AI summary
The present disclosure relates to a method and device for generating video preview content, a computer device and a storage medium. The method for generating video preview content includes parsing a video to be processed, to obtain all image frames of the video to be processed and generate a list of ordered image frames; processing the list of the ordered image frames by image recognition to filter out a slice header and a slice tail of the video; and generating preview content of the video based on a list of the filtered image frames.


