Machine Learning Video Frame Selection for Aesthetic Thumbnails

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing messaging systems lack an efficient method to select a representative video frame for use as a thumbnail, often relying on manual selection or simplistic algorithms that do not account for aesthetic quality or user preferences.

Innovation Solution

A messaging system utilizing a machine learning model trained with aesthetic visual analysis data and self-training on user-preferred video frames to rank and select a representative frame based on predefined preferences, reducing the need for extensive manual labeling.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Ease of operation

If manual selection or simplistic algorithms are used to select video frames, then the system is easy to operate, but the aesthetic quality and user preference alignment of selected frames deteriorates

Engineering Contradiction:
Improveease of frame selectionVSAvoidaesthetic quality of selected frame
Core Design Contradiction:
Ease of operationVSManufacturing precision

Solution Approach 1:

The machine learning model performs self-training by using its own predictions on unlabeled video frames to generate pseudo-labels, which are then used to further improve the model. This self-service mechanism allows the system to automatically enhance its frame selection capability without requiring continuous manual intervention or extensive pre-labeled training data, thereby maintaining ease of operation while improving aesthetic quality.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The system transforms the frame selection process from a simple algorithmic choice to a sophisticated machine learning-based decision process. By changing the parameter of selection methodology from deterministic rules to probabilistic machine learning models trained on aesthetic preferences, the system achieves higher aesthetic quality while maintaining operational simplicity through automation.

Inventive Principle:
Principle #35Parameter changes

2Measurement precision

If extensive manual labeling is performed to train the machine learning model, then the measurement precision of user preferences improves, but the loss of time and productivity deteriorates

Engineering Contradiction:
Improveprecision of user preference measurementVSAvoidtime for manual labeling
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The machine learning model performs self-training by using its own predictions on unlabeled video frames to generate pseudo-labels, which are then used to further improve the model. This self-service mechanism allows the system to automatically enhance its frame selection capability without requiring continuous manual intervention or extensive pre-labeled training data, thereby maintaining ease of operation while improving aesthetic quality.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The system performs preliminary training with a small set of manually labeled frames to establish initial user preference measurements. This preliminary action creates a foundation that enables subsequent self-training on large volumes of unlabeled data, significantly reducing the total time required compared to manually labeling all training frames while maintaining high measurement precision.

Inventive Principle:
Principle #10Preliminary action

3Measurement precision

If a machine learning model is trained with extensive labeled data, then the selection accuracy of representative frames improves, but the device complexity and resource requirements worsen

Engineering Contradiction:
Improveaccuracy of frame selectionVSAvoidcomplexity of training system
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The machine learning model performs self-training by using its own predictions on unlabeled video frames to generate pseudo-labels, which are then used to further improve the model. This self-service mechanism allows the system to automatically enhance its frame selection capability without requiring continuous manual intervention or extensive pre-labeled training data, thereby maintaining ease of operation while improving aesthetic quality.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

Instead of requiring complete manual labeling of all training data, the system uses partial manual labeling combined with self-generated pseudo-labels. This partial action approach achieves high selection accuracy by leveraging the model's own capabilities to generate sufficient training signals, reducing both data labeling burden and system complexity while maintaining or improving frame selection accuracy.

Inventive Principle:
Principle #16Partial or excessive action

Data Source

PatentUS12354355B2Machine learning-based selection of a representative video frame within a messaging application
Publication Date: 2025.07.08 SNAP INC
  • US12354355B2 patent drawing
  • US12354355B2 patent drawing
  • US12354355B2 patent drawing

AI summary

Aspects of the present disclosure involve a system comprising a medium storing a program and method for machine-learning based selection of a representative video frame. The program and method provide for receiving a set of video frames; determining a first subset of frames by removing frames outside of an image quality threshold; determining a second subset by removing frames outside of an image stillness threshold; computing feature data for each frame in the second subset; providing, for each frame in the second subset, the feature data to a machine learning model (MLM), the MLM being configured to output a score for each frame in the second subset of frames based on the feature data, the MLM having been trained with a first set of images labeled based on aesthetics, and with a second set of images labeled based on image quality; and selecting a frame based on output scores.