Machine Learning Aspect Ratio Conversion

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Traditional pan and scan methods for converting wide-screen motion picture aspect ratios to smaller aspect ratios for home distribution are labor-intensive and time-consuming, requiring manual decisions that are difficult to automate effectively.

Innovation Solution

A computer-implemented method using machine learning to determine aspect ratio conversion functions based on extracted visual and audio features from video frames, mimicking human framing decisions to convert image frames between aspect ratios, incorporating both rule-based and model-based approaches for predicting pan and scan framing decisions.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If manual pan and scan methods are used for aspect ratio conversion, then framing decisions can be made with human creativity and judgment, but the process becomes labor intensive and time consuming

Engineering Contradiction:
Improveframing decision accuracyVSAvoidconversion speed
Core Design Contradiction:
Measurement precisionVSProductivity

Solution Approach 1:

The system creates a digital copy of manual framing decisions by training machine learning models on manually converted video data. The models learn to replicate human framing choices without requiring actual human operators during conversion, thus preserving the quality of manual decisions while eliminating the time and labor costs.

Inventive Principle:
Principle #26Copying

Solution Approach 2:

The patent replaces the mechanical process of manual video framing with an automated machine learning system. The ML models automatically perform frame selection and aspect ratio conversion based on patterns learned from training data, substituting human manual operations with computational processes that are both accurate and efficient.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

2Productivity

If automated methods are used for aspect ratio conversion, then productivity increases and manual labor decreases, but the ability to preserve critical elements and creative intent may be compromised

Engineering Contradiction:
Improveconversion efficiencyVSAvoidpreservation of critical elements
Core Design Contradiction:
ProductivityVSReliability

Solution Approach 1:

The system performs preliminary action by pre-training machine learning models on extensive datasets of manually converted video content before actual conversion tasks. This pre-training establishes a foundation of learned patterns for identifying and preserving critical visual and audio elements, ensuring reliable performance during automated conversion operations.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent incorporates feedback mechanisms where the machine learning models are trained on ground truth data from manual conversions, allowing them to learn from examples of correct framing decisions. The models receive feedback during training on which elements should be preserved and which framing choices are optimal, enabling them to reliably replicate human judgment in automated operations.

Inventive Principle:
Principle #23Feedback

Data Source

PatentUS10992902B2Aspect ratio conversion with machine learning
Publication Date: 2021.04.27 DISNEY ENTERPRISES INC
  • US10992902B2 patent drawing
  • US10992902B2 patent drawing
  • US10992902B2 patent drawing

AI summary

Techniques are disclosed for converting image frames, such as the image frames of a motion picture, from one aspect ratio to another while predicting the pan and scan framing decisions that a human operator would make. In one configuration, one or more functions for predicting pan and scan framing decisions are determined, at least in part, via machine learning using training data that includes historical pan and scan conversions. The training data may be prepared by extracting features indicating visual and/or audio elements associated with particular shots, among other things. Function(s) may be determined, using machine learning, that take such extracted features as input and output predicted pan and scan framing decisions. Thereafter, the image frames of a received video may be converted between aspect ratios on a shot-by-shot basis, by extracting the same features and using the function(s) to make pan and scan framing predictions.