Automated Media Annotation via Visual Similarity Detection

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

The manual annotation of interactive metadata for streaming video services, such as Prime Video's X-Ray, is not scalable due to the need for human operators to watch and annotate each scene of millions of titles, making it inefficient and labor-intensive.

Innovation Solution

Automating the annotation process by identifying visually similar intervals in video frames to determine uneventful portions of media presentations, allowing human operators to skip or fast-forward through these sections, using facial recognition and histogram analysis to time-code meaningful events and reduce redundant event detection.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If manual annotation is used for each scene and cast member, then annotation accuracy is maintained, but productivity becomes extremely low due to the need for human operators to watch and annotate every title

Engineering Contradiction:
Improveannotation throughputVSAvoidmanual annotation requirement
Core Design Contradiction:
ProductivityVSExtent of automation

Solution Approach 1:

The system performs preliminary automated detection of cast members, scenes, and events before human operators review the content. Automated event detection, facial recognition, and scene segmentation are performed in advance to prepare annotated data for operator review, significantly reducing the manual workload while maintaining accuracy

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The annotation system performs self-service by automatically detecting and annotating cast members, scenes, and events without requiring continuous human intervention. The system independently processes video content, identifies relevant entities, and generates metadata that can be directly used or reviewed by operators

Inventive Principle:
Principle #25Self-service

2Productivity

If automated event detection is implemented, then productivity increases, but measurement precision may be compromised due to potential false detections

Engineering Contradiction:
Improveannotation speedVSAvoidevent detection accuracy
Core Design Contradiction:
ProductivityVSMeasurement precision

Solution Approach 1:

The system incorporates feedback mechanisms where automated event detections are reviewed and validated by human operators. Operators provide feedback on detected events, allowing the system to learn from corrections and improve its detection accuracy over time while maintaining high productivity

Inventive Principle:
Principle #23Feedback

Solution Approach 2:

The patent replaces manual mechanical review of every frame with automated computer-based detection systems using facial recognition algorithms, scene segmentation, and event detection models. This substitution enables high-speed processing while maintaining precision through sophisticated automated algorithms

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

3Reliability

If all video frames are processed to identify events, then completeness of metadata is ensured, but loss of time increases due to processing entire titles manually

Engineering Contradiction:
Improvemetadata completenessVSAvoidannotation time
Core Design Contradiction:
ReliabilityVSLoss of time

Solution Approach 1:

The system extracts and processes only the most relevant portions of video content for annotation, such as scenes with cast members, dialogue, or significant events. By identifying and focusing on these key segments rather than processing every frame uniformly, the system maintains metadata completeness while significantly reducing processing time

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The video content is segmented into distinct scenes and temporal segments based on visual characteristics and event occurrences. This segmentation allows the system to process and annotate only the relevant segments rather than treating the entire video as a single unit, reducing overall annotation time while maintaining completeness through comprehensive scene-level analysis

Inventive Principle:
Principle #1Segmentation

Data Source

PatentUS10897658B1Techniques for annotating media content
Publication Date: 2021.01.19 AMAZON TECH INC
  • US10897658B1 patent drawing
  • US10897658B1 patent drawing
  • US10897658B1 patent drawing

AI summary

Methods and apparatus are described for automating aspects of the annotation of a media presentation. Events are identified that relate to entities associated with the scenes of the media presentation. These events are time coded relative to the media timeline of the media presentation and might represent, for example, the appearance of a particular cast member or playback of a particular music track. The video frames of the media presentation are processed to identify visually similar intervals that may serve as or be used to identify contexts (e.g., scenes) within the media presentation. Relationships between the event data and the visually similar intervals or contexts are used to identify portions of the media presentation during which the occurrence of additional meaningful events is unlikely. This information may be surfaced to a human operator tasked with annotating the content as an indication that part of the media presentation may be skipped.