Machine Learning Aspect Ratio Conversion
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Traditional pan and scan methods for converting wide-screen motion picture aspect ratios to smaller aspect ratios for home distribution are labor-intensive and time-consuming, requiring manual decisions that are difficult to automate effectively.
Innovation Solution
A computer-implemented method using machine learning to determine aspect ratio conversion functions based on extracted visual and audio features from video frames, mimicking human framing decisions to convert image frames between aspect ratios, incorporating both rule-based and model-based approaches for predicting pan and scan framing decisions.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If manual pan and scan methods are used for aspect ratio conversion, then framing decisions can be made with human creativity and judgment, but the process becomes labor intensive and time consuming
Solution Approach 1:
The system creates a digital copy of manual framing decisions by training machine learning models on manually converted video data. The models learn to replicate human framing choices without requiring actual human operators during conversion, thus preserving the quality of manual decisions while eliminating the time and labor costs.
Solution Approach 2:
The patent replaces the mechanical process of manual video framing with an automated machine learning system. The ML models automatically perform frame selection and aspect ratio conversion based on patterns learned from training data, substituting human manual operations with computational processes that are both accurate and efficient.
2Productivity
If automated methods are used for aspect ratio conversion, then productivity increases and manual labor decreases, but the ability to preserve critical elements and creative intent may be compromised
Solution Approach 1:
The system performs preliminary action by pre-training machine learning models on extensive datasets of manually converted video content before actual conversion tasks. This pre-training establishes a foundation of learned patterns for identifying and preserving critical visual and audio elements, ensuring reliable performance during automated conversion operations.
Solution Approach 2:
The patent incorporates feedback mechanisms where the machine learning models are trained on ground truth data from manual conversions, allowing them to learn from examples of correct framing decisions. The models receive feedback during training on which elements should be preserved and which framing choices are optimal, enabling them to reliably replicate human judgment in automated operations.
Data Source
AI summary
Techniques are disclosed for converting image frames, such as the image frames of a motion picture, from one aspect ratio to another while predicting the pan and scan framing decisions that a human operator would make. In one configuration, one or more functions for predicting pan and scan framing decisions are determined, at least in part, via machine learning using training data that includes historical pan and scan conversions. The training data may be prepared by extracting features indicating visual and/or audio elements associated with particular shots, among other things. Function(s) may be determined, using machine learning, that take such extracted features as input and output predicted pan and scan framing decisions. Thereafter, the image frames of a received video may be converted between aspect ratios on a shot-by-shot basis, by extracting the same features and using the function(s) to make pan and scan framing predictions.


