Content-Based Video Zooming and Panning Automation

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing video editing processes for electronic devices are manual and subjective, leading to uninteresting playback of wide-field video data captured by cameras, as they lack automation in emphasizing content within the video.

Innovation Solution

The system automates video editing by using computer vision and machine learning algorithms to detect and analyze content in video data, applying content-based zooming and panning effects to emphasize events of interest, such as sporting events or human interactions, by determining framing windows and simulating panning and zooming between them.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Ease of operation

If manual video editing is used to emphasize content, then the video playback becomes more interesting, but the editing process becomes time-consuming and subjective

Engineering Contradiction:
Improvevideo playback interestVSAvoidediting time
Core Design Contradiction:
Ease of operationVSLoss of time

Solution Approach 1:

The system enables self-service video editing by automatically analyzing video content using computer vision algorithms to detect events of interest, determine framing windows, and generate edited video clips without requiring manual user intervention. The processor autonomously performs content analysis, event detection, and video synthesis tasks.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The patent replaces manual mechanical editing operations with automated computer vision-based content analysis. Instead of manually reviewing and editing video frames, the system uses machine learning algorithms to automatically detect events, determine framing windows, and generate edited clips, substituting human editorial work with computational processing.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

2Area of stationary object

If wide-field video data is captured to cover more area, then more content is visible, but the playback becomes static and uninteresting

Engineering Contradiction:
Improvevideo field of viewVSAvoidvideo playback engagement
Core Design Contradiction:
Area of stationary objectVSEase of operation

Solution Approach 1:

The system segments the wide-field video data by determining specific framing windows that focus on events of interest. Instead of displaying the entire wide field of view statically, the processor divides the video into relevant segments based on detected events, creating dynamic framed clips that maintain viewer engagement while preserving the comprehensive coverage capability.

Inventive Principle:
Principle #1Segmentation

3Ease of operation

If automated content-based zooming and panning is applied, then video playback becomes more engaging, but the device complexity increases

Engineering Contradiction:
Improvevideo playback engagementVSAvoidprocessing complexity
Core Design Contradiction:
Ease of operationVSDevice complexity

Solution Approach 1:

The patent replaces complex manual video editing operations with automated computer vision algorithms. The processor uses machine learning models to perform content analysis, event detection, and framing window determination, substituting human editorial complexity with computational processing that can be efficiently executed on modern devices.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Data Source

PatentUS9973711B2Content-based zooming and panning for video curation
Publication Date: 2018.05.15 AMAZON TECH INC
  • US9973711B2 patent drawing
  • US9973711B2 patent drawing
  • US9973711B2 patent drawing

AI summary

Devices, systems and methods are disclosed for identifying content in video data and creating content-based zooming and panning effects to emphasize the content. Contents may be detected and analyzed in the video data using computer vision, machine learning algorithms or specified through a user interface. Panning and zooming controls may be associated with the contents, panning or zooming based on a location and size of content within the video data. The device may determine a number of pixels associated with content and may frame the content to be a certain percentage of the edited video data, such as a close-up shot where a subject is displayed as 50% of the viewing frame. The device may identify an event of interest, may determine multiple frames associated with the event of interest and may pan and zoom between the multiple frames based on a size/location of the content within the multiple frames.