Video Preview Extraction Using Attention Weights

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Users face challenges in selecting and browsing video content due to insufficient illustration provided by static video titles and cover pictures, often leading to missed desired resources and wasted time and network resources, as only big-budget videos typically have video previews, while others lack them.

Innovation Solution

A method and apparatus for automatically extracting video previews using a pre-trained video classification model comprising a Convolutional Neural Network, time sequence neural network, attention module, and fully-connected layer, which calculates weights of video frames to select continuous frames as the preview, reducing manual clipping and manpower costs.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If manual clipping is used to create video previews, then the quality of video preview representation is improved, but the manpower costs and time consumption increase significantly

Engineering Contradiction:
Improvevideo preview representation qualityVSAvoidtime consumption for manual clipping
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The patent replaces the manual mechanical clipping process with an automated deep learning-based system. The video classification model, comprising CNN for feature extraction, LSTM for temporal modeling, and attention mechanisms, automatically identifies and extracts representative video frames without human intervention, thereby eliminating time consumption while maintaining preview quality.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Solution Approach 2:

The system enables self-service by automatically generating video previews using the trained deep learning model. The model processes input videos autonomously, selecting appropriate frames based on learned patterns and attention weights, thus performing the preview generation task independently without requiring manual clipping operations.

Inventive Principle:
Principle #25Self-service

2Productivity

If only big-budget videos have video previews, then the resource allocation is optimized for high-value content, but the coverage of video content with previews is limited

Engineering Contradiction:
Improvevideo preview generation efficiencyVSAvoidcoverage of videos with previews
Core Design Contradiction:
ProductivityVSAdaptability or versatility

Solution Approach 1:

The patent achieves universality by developing a general-purpose video preview generation system that can handle diverse video types and content categories. The deep learning model, trained on comprehensive video datasets, applies the same automated processing pipeline to various video genres, enabling cost-effective preview generation across all video types rather than仅限于 high-budget productions.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The system utilizes parameter changes by adjusting the deep learning model's attention weights and frame selection criteria based on video characteristics. The model dynamically adapts its processing parameters to suit different video types, content lengths, and visual characteristics, thereby achieving efficient and versatile preview generation across the entire video library without requiring manual intervention for each video type.

Inventive Principle:
Principle #35Parameter changes

Data Source

PatentUS11302103B2Method and apparatus for extracting video preview, device and computer storage medium
Publication Date: 2022.04.12 BAIDU ONLINE NETWORK TECH (BEIJIBG) CO LTD
  • US11302103B2 patent drawing
  • US11302103B2 patent drawing
  • US11302103B2 patent drawing

AI summary

The present disclosure provides a method and apparatus for extracting a video preview, a device and a computer storage medium. The method comprises: inputting a video into a video classification model obtained by pre-training; obtaining weights of respective video frames output by an attention module in the video classification model; extracting continuous N video frames whose total weight value satisfies a preset requirement, as the video preview of the target video, N being a preset positive integer. It is possible to, in the manner provided by the present disclosure, automatically extract continuous video frames from the video as the video preview, without requiring manual clipping, and with manpower costs being reduced.