Video Entity Annotation via Frame-Level Probability Calibration

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Users face difficulties in identifying relevant video content items from large search results due to the lack of precise information about the content within media hosting services, even with additional metadata like images, authors, and popularity indicators.

Innovation Solution

A computer-implemented method for annotating video frames with entities and their associated probabilities of existence, using a video hosting system that selects features correlated with entities, determines a classifier, and applies an aggregation calibration function to identify and rank relevant video frames based on entity presence.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Loss of information

If metadata such as images, authors, and popularity indicators are provided to help users assess video content relevance, then the amount of information available to users increases, but the ability of users to accurately determine whether video content contains relevant information remains insufficient

Engineering Contradiction:
Improveinformation completenessVSAvoidcontent relevance accuracy
Core Design Contradiction:
Loss of informationVSMeasurement precision

Solution Approach 1:

The patent segments video content into individual frames and identifies specific entities within each frame. By analyzing video frames at a granular level rather than treating the entire video as a single unit, the system can provide precise information about what specific content appears in specific portions of the video, thereby improving both information completeness and relevance accuracy.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent replaces manual content assessment with automated computer vision and machine learning systems. The system uses trained classifiers and aggregation calibration functions to automatically identify entities in video frames and calculate probabilities of entity existence, substituting human judgment with algorithmic analysis to achieve more consistent and accurate relevance determination.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

2Quantity of substance

If hundreds or thousands of media content items are returned in search results, then the comprehensiveness of search results improves, but the difficulty for users to assess which items are most relevant increases

Engineering Contradiction:
Improvesearch result quantityVSAvoiduser assessment difficulty
Core Design Contradiction:
Quantity of substanceVSEase of operation

Solution Approach 1:

The patent highlights video frames that contain entities with high probabilities of existence, visually distinguishing relevant content from irrelevant content. This visual emphasis mechanism helps users quickly identify which portions of videos are relevant to their search queries, making it easier to assess relevance across large numbers of search results without having to manually review each video.

Inventive Principle:
Principle #32Color changes

3Measurement precision

If video frames are annotated with entities and probabilities of existence, then the precision of content identification improves, but the computational complexity of processing and analyzing video data increases

Engineering Contradiction:
Improveentity identification accuracyVSAvoidsystem processing complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent performs preliminary actions by pre-processing video data into frames, pre-training classifiers on labeled data, and pre-calculating feature correlations with entities before actual search queries are executed. This preprocessing and model training phase prepares the system in advance, so that when queries are received, the system can quickly apply the pre-trained models to annotate video frames without performing complex computations in real-time, thus reducing operational complexity.

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS12141199B2Feature-based video annotation
Publication Date: 2024.11.12 GOOGLE LLC
  • US12141199B2 patent drawing
  • US12141199B2 patent drawing
  • US12141199B2 patent drawing

AI summary

A system and methodology provide for annotating videos with entities and associated probabilities of existence of the entities within video frames. A computer-implemented method identifies an entity from a plurality of entities identifying characteristics of video items. The computer-implemented method selects a set of features correlated with the entity based on a value of a feature of a plurality of features, determines a classifier for the entity using the set of features, and determines an aggregation calibration function for the entity based on the set of features. The computer-implemented method selects a video frame from a video item, where the video frame having associated features, and determines a probability of existence of the entity based on the associated features using the classifier and the aggregation calibration function.