Video Search Preview Image Selection via Key Frame Content Analysis

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional media-hosting systems return search results that are not indicative of the content within a video, often providing irrelevant preview images, leading to inefficient and time-consuming searches for users as they struggle to find relevant content.

Innovation Solution

The system identifies and provides relevant preview images by selecting key frames that include specific content features based on confidence values, allowing users to easily find video content related to their search queries through content-based and non-content-based methods, including machine learning techniques for content feature recognition.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Loss of information

If conventional media-hosting systems return search results using video titles or categories, then the search process is simple to implement, but the search results are not indicative of the actual content within videos, leading to poor user experience

Engineering Contradiction:
Improvecontent information in search resultsVSAvoidsearch result generation system
Core Design Contradiction:
Loss of informationVSDevice complexity

Solution Approach 1:

The system performs preliminary content analysis on video frames before search queries are submitted. Key frames are extracted and pre-tagged with content features, so that when a search query arrives, the system can quickly match queries against pre-analyzed content rather than analyzing everything from scratch. This resolves the contradiction by preparing information in advance, reducing both information loss and query-time complexity.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The system introduces key frame images and content tags as intermediary elements between the full video content and the search query. Instead of directly comparing queries to entire videos or simple metadata, the system uses these intermediate representations that capture essential content features, enabling efficient and accurate content-based search without requiring complex full-video analysis for each query.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Loss of information

If conventional systems assign the first frame or manually selected frame as preview image, then the implementation is simple, but the preview image is rarely relevant to the user's search query related to particular content

Engineering Contradiction:
Improverelevance information in preview imageVSAvoidpreview image selection system
Core Design Contradiction:
Loss of informationVSDevice complexity

Solution Approach 1:

Instead of using a single uniform frame selection approach for all videos and queries, the system applies local quality by selecting different key frames based on their content relevance to specific search queries. Each preview image is chosen to highlight the particular content features that match the user's search intent, making the preview locally optimized for relevance rather than globally uniform.

Inventive Principle:
Principle #3Local quality

Solution Approach 2:

The system changes the selection parameter from fixed (first frame or manual selection) to dynamic (relevance-based selection). By evaluating content features against search query terms and selecting frames that maximize relevance scores, the system transforms the preview image selection from a static process to a dynamic, query-adaptive process that optimizes for information relevance.

Inventive Principle:
Principle #35Parameter changes

3Productivity

If users manually view multiple videos to find relevant content, then no additional system complexity is required, but the search process becomes inefficient and time-consuming

Engineering Contradiction:
Improvesearch efficiencyVSAvoidcontent analysis system
Core Design Contradiction:
ProductivityVSDevice complexity

Solution Approach 1:

The system performs content analysis and extracts key features from video frames in advance, before users submit search queries. This preliminary processing creates an indexed representation of video content that enables rapid search matching, dramatically improving search efficiency without requiring users to manually view videos. The computational complexity is shifted to a pre-processing stage rather than being incurred during each search operation.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The system replaces the mechanical approach of manual video viewing with an automated content-based retrieval system. Instead of relying on users to physically watch and evaluate multiple videos, the system uses computer vision and machine learning to automatically analyze video content, extract meaningful features, and retrieve relevant results, substituting human manual effort with automated intelligent processing.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Data Source

PatentUS11461392B2Providing relevant cover frame in response to a video search query
Publication Date: 2022.10.04 ADOBE INC
  • US11461392B2 patent drawing
  • US11461392B2 patent drawing
  • US11461392B2 patent drawing

AI summary

The present disclosure is directed towards methods and systems for providing relevant video scenes in response to a video search query. The systems and methods identify a plurality of key frames of a media object and detect one or more content features represented in the plurality of key frames. Based on the one or more detect content features, the systems and methods associate tags indicating the detected content features with the plurality of key frames of the media object. The systems and methods, in response to receiving a search query including search terms, compare the search terms with the tags of the selected key frames, identify a selected key frame that depicts at least one content feature related to the search terms, and provide a preview image of the media item depicting the at least one content feature.