Video Clip Selection Model Using Frame-Title Relevance
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
The increasing volume of video resources leads to a manpower shortage and high costs due to the inefficiency of existing methods for selecting exciting video clips, which either require manual selection or full-supervised model training, resulting in repetitive and costly processes.
Innovation Solution
A method and apparatus that determine the excitement of video clips by inputting feature sequences and title information into a pre-established prediction model, using a combination of fully connected networks, GRU modules, and attention modules to calculate the relevance between video frames and titles, thereby automating the selection of the most exciting clip.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If manual selection method is used to select exciting video clips, then the selection accuracy can be ensured, but the productivity is low and manpower shortage occurs due to the increasing volume of video resources
Solution Approach 1:
The system enables automatic video clip selection through self-service mechanisms. The prediction model automatically processes video frames, extracts features, and selects exciting clips without human intervention. The tagger only needs to input the video and title, while the system autonomously completes the entire selection process, eliminating manual workload and大幅提升 productivity.
Solution Approach 2:
The patent replaces the mechanical manual selection process with an automated prediction model system. Instead of human taggers manually watching and selecting clips, the system uses deep learning models (GRU, attention mechanisms) to automatically analyze video content and identify exciting segments, substituting human mechanical work with automated computational processes.
2Measurement precision
If full-supervised model training is used to select exciting video clips, then the selection accuracy can be improved, but the device complexity and costs increase due to the need for operator marking and deep learning training
Solution Approach 1:
The system performs preliminary action by pre-processing video frames and extracting features before the actual selection process. The prediction model is pre-trained and ready to use, so when a new video is input, the system can immediately process it without requiring new training. This preliminary preparation eliminates the need for complex per-video training operations.
Solution Approach 2:
The patent extracts only the essential features from video frames using the prediction model, rather than processing the entire video content manually. The system extracts key visual features and temporal patterns that are most relevant to identifying exciting clips, eliminating unnecessary processing steps and reducing complexity while maintaining accuracy.
3Measurement precision
If deep learning training is performed for each training video, then the measurement precision of exciting clip detection can be achieved, but the loss of time increases due to repetitive training processes
Solution Approach 1:
The prediction model is designed with universality to handle multiple different videos without requiring separate training for each one. The model learns general patterns of exciting clips that can be applied across diverse video content. This multi-functional capability allows the same model to accurately process various types of videos, eliminating repetitive training time while maintaining detection accuracy.
Data Source
AI summary
Embodiments of the present disclosure relate to a method and apparatus for selecting a video clip, a server and a medium. The method may include: determining at least two video clips from a video; for each video clip, perform following excitement determination steps: inputting a feature sequence of a video frame in the video clip and title information of the video into a pre-established prediction model to obtain a relevance between the inputted video frame and a title of the video; and determining an excitement of the video clip, based on the relevance between the video frame in the video clip and the title; and determining a target video clip from the video clips, based on the excitement of each of the video clips.


