Video Object Tracking with Key Frame Analysis

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current video consumption and interaction are limited due to the lack of contextual metadata, particularly in identifying and tracking visual objects within video frames, which requires laborious frame-by-frame analysis.

Innovation Solution

A video editor that automatically or manually identifies objects in a video and tracks their location across frames, using object tracking systems and machine learning to update indicators and associate metadata such as product IDs with annotations like customer reviews and related videos.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If frame-by-frame analysis is used to identify and track visual objects, then object identification accuracy is improved, but time consumption and labor requirements increase significantly

Engineering Contradiction:
Improveobject identification accuracyVSAvoidtime consumption
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The system performs preliminary object identification and tracking by analyzing key frames or representative frames of the video, rather than every single frame. Objects identified in key frames are then tracked across subsequent frames using motion estimation and object persistence algorithms, significantly reducing the total number of frames that require full analysis while maintaining accurate object identification and continuous tracking throughout the video sequence

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The video processing is divided into segments based on key frames or significant events. Object identification is performed intensively on these segmented key frames, while intermediate frames use lighter tracking algorithms. This segmentation approach reduces overall processing time while preserving object identification accuracy at critical moments

Inventive Principle:
Principle #1Segmentation

2Adaptability or versatility

If contextual metadata is added to videos to improve user interaction and information availability, then user engagement and decision-making are improved, but system complexity and processing requirements increase

Engineering Contradiction:
Improveuser interaction capabilityVSAvoidsystem complexity
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The video processing system is enhanced with multi-functional capabilities that simultaneously perform object identification, tracking, metadata extraction, and annotation generation within a single integrated framework. This universal approach consolidates multiple separate processing functions into one system, reducing overall complexity while enabling rich contextual metadata generation for improved user interaction

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

An intermediary processing layer is introduced between video playback and user interaction, which automatically generates and attaches contextual metadata (object identifiers, tracking information, annotations) to video frames. This intermediary layer handles the complexity of metadata generation transparently, allowing users to benefit from enhanced interaction capabilities without directly managing system complexity

Inventive Principle:
Principle #24Intermediary (Mediator)

Data Source

PatentUS11170817B2Tagging tracked objects in a video with metadata
Publication Date: 2021.11.09 AMAZON TECH INC
  • US11170817B2 patent drawing
  • US11170817B2 patent drawing
  • US11170817B2 patent drawing

AI summary

Embodiments herein describe a video editor that can identify and track objects (e.g., products) in a video. The video editor identifies a particular object in one frame of the video and tracks the location of the object in the video. The video editor can update a position of an indicator that tracks the location of the object in the video. In addition, the video editor can identify an identification (ID) of the object which the editor can use to suggest annotations that provide additional information about the object. Once modified, the video is displayed on a user device, and when the viewer sees an object she can is interested in, she can pause the video which causes the indicator to appear. The user can select the indicator which prompts the user device to display the annotations corresponding to the object.