Object Tracking in Zoomed Video for Small Screens

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Users face difficulties in viewing fine details and facial expressions in video content on small-screen devices due to the content being optimized for larger screens, as the current methods of adjusting display settings often fail to keep objects of interest centered and adequately magnified.

Innovation Solution

The system allows users to select and track objects of interest within video content using touch input, gaze tracking, or audio commands, dynamically adjusting magnification and centering to maintain a consistent presentation size, even as the object moves within the frame, using algorithms like particle tracking and object recognition to ensure accurate tracking.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Ease of operation

If video content is displayed on small-screen devices, then the device portability and accessibility are improved, but the ability to view fine details and facial expressions deteriorates

Engineering Contradiction:
Improvedevice accessibilityVSAvoiddetail visibility
Core Design Contradiction:
Ease of operationVSMeasurement precision

Solution Approach 1:

The system dynamically adjusts the magnification level and displayed region based on user interactions (touches, gestures) and tracked object positions. The magnification level is not fixed but changes in response to user actions, allowing the same device to adapt between viewing entire scenes and examining fine details of specific objects or facial expressions.

Inventive Principle:
Principle #15Dynamics

2Device complexity

If a static region is selected for display, then the display settings are simplified, but the object of interest may move outside the selected region

Engineering Contradiction:
Improvedisplay setting complexityVSAvoidobject tracking reliability
Core Design Contradiction:
Device complexityVSReliability

Solution Approach 1:

The system continuously monitors the position of tracked objects within the displayed region and uses this feedback to determine when and how to adjust the displayed region. When an object approaches the boundary or leaves the region, the system responds by shifting or expanding the displayed region to keep the object of interest visible, creating a closed-loop control system.

Inventive Principle:
Principle #23Feedback

Solution Approach 2:

The displayed region transitions from a static selection to a dynamic area that automatically adjusts its position and size based on tracked object movement. This dynamic region follows the object of interest while maintaining appropriate framing, eliminating the need for manual repositioning by the user.

Inventive Principle:
Principle #15Dynamics

3Measurement precision

If magnification is increased to view fine details, then the detail visibility is improved, but the field of view is reduced

Engineering Contradiction:
Improvedetail visibilityVSAvoidfield of view
Core Design Contradiction:
Measurement precisionVSArea of stationary object

Solution Approach 1:

The system dynamically adjusts the magnification level based on user gestures (such as pinching motions) and automatically adapts the displayed region to match the new magnification level. This allows the field of view to be appropriately scaled with each magnification change, maintaining optimal framing at every level rather than using a fixed magnification setting.

Inventive Principle:
Principle #15Dynamics

Data Source

PatentUS10664140B2Object tracking in zoomed video
Publication Date: 2020.05.26 AMAZON TECH INC
  • US10664140B2 patent drawing
  • US10664140B2 patent drawing
  • US10664140B2 patent drawing

AI summary

A user can select an object represented in video content in order to set a magnification level with respect to that object. A portion of the video frames containing a representation of the object is selected to maintain a presentation size of the representation corresponding to the magnification level. The selection provides for a “smart zoom” feature enabling an object of interest, such as a face of an actor, to be used in selecting an appropriate portion of each frame to magnify, such that the magnification results in a portion of the frame being selected that includes the one or more objects of interest to the user. Pre-generated tracking data can be provided for some objects, which can enable a user to select an object and then have predetermined portion selections and magnifications applied that can provide for a smoother user experience than for dynamically-determined data.