Video Clip Selection Model Using Frame-Title Relevance

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

The increasing volume of video resources leads to a manpower shortage and high costs due to the inefficiency of existing methods for selecting exciting video clips, which either require manual selection or full-supervised model training, resulting in repetitive and costly processes.

Innovation Solution

A method and apparatus that determine the excitement of video clips by inputting feature sequences and title information into a pre-established prediction model, using a combination of fully connected networks, GRU modules, and attention modules to calculate the relevance between video frames and titles, thereby automating the selection of the most exciting clip.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If manual selection method is used to select exciting video clips, then the selection accuracy can be ensured, but the productivity is low and manpower shortage occurs due to the increasing volume of video resources

Engineering Contradiction:
Improveselection accuracyVSAvoidvideo processing efficiency
Core Design Contradiction:
Measurement precisionVSProductivity

Solution Approach 1:

The system enables automatic video clip selection through self-service mechanisms. The prediction model automatically processes video frames, extracts features, and selects exciting clips without human intervention. The tagger only needs to input the video and title, while the system autonomously completes the entire selection process, eliminating manual workload and大幅提升 productivity.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The patent replaces the mechanical manual selection process with an automated prediction model system. Instead of human taggers manually watching and selecting clips, the system uses deep learning models (GRU, attention mechanisms) to automatically analyze video content and identify exciting segments, substituting human mechanical work with automated computational processes.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

2Measurement precision

If full-supervised model training is used to select exciting video clips, then the selection accuracy can be improved, but the device complexity and costs increase due to the need for operator marking and deep learning training

Engineering Contradiction:
Improveexciting clip identification accuracyVSAvoidmodel training complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The system performs preliminary action by pre-processing video frames and extracting features before the actual selection process. The prediction model is pre-trained and ready to use, so when a new video is input, the system can immediately process it without requiring new training. This preliminary preparation eliminates the need for complex per-video training operations.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent extracts only the essential features from video frames using the prediction model, rather than processing the entire video content manually. The system extracts key visual features and temporal patterns that are most relevant to identifying exciting clips, eliminating unnecessary processing steps and reducing complexity while maintaining accuracy.

Inventive Principle:
Principle #2Taking out (Extraction)

3Measurement precision

If deep learning training is performed for each training video, then the measurement precision of exciting clip detection can be achieved, but the loss of time increases due to repetitive training processes

Engineering Contradiction:
Improveexciting clip detection accuracyVSAvoidtraining time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The prediction model is designed with universality to handle multiple different videos without requiring separate training for each one. The model learns general patterns of exciting clips that can be applied across diverse video content. This multi-functional capability allows the same model to accurately process various types of videos, eliminating repetitive training time while maintaining detection accuracy.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Data Source

PatentUS11490168B2Method and apparatus for selecting video clip, server and medium
Publication Date: 2022.11.01 BEIJING BAIDU NETCOM SCI & TECH CO LTD
  • US11490168B2 patent drawing
  • US11490168B2 patent drawing
  • US11490168B2 patent drawing

AI summary

Embodiments of the present disclosure relate to a method and apparatus for selecting a video clip, a server and a medium. The method may include: determining at least two video clips from a video; for each video clip, perform following excitement determination steps: inputting a feature sequence of a video frame in the video clip and title information of the video into a pre-established prediction model to obtain a relevance between the inputted video frame and a title of the video; and determining an excitement of the video clip, based on the relevance between the video frame in the video clip and the title; and determining a target video clip from the video clips, based on the excitement of each of the video clips.