Automatic Video Reframing Using Saliency and OCR for Social Formats
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing video reframing and transformation methods are time-consuming, costly, and not scalable for social media distribution, requiring manual editing and powerful rendering machines, and fail to automatically adapt to different aspect ratios and genres.
Innovation Solution
A method and system that automatically reframes and transforms videos using saliency models, OCR, and spatio-temporal analysis to identify regions of interest, eliminate black bands, and apply genre-specific transformations, enabling rapid conversion across aspect ratios and formats.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Manufacturing precision
If manual video reframing is performed by skilled editors, then video quality and contextual accuracy are improved, but time consumption and cost increase significantly
Solution Approach 1:
The system performs automatic video reframing using AI algorithms that analyze video content, detect regions of interest, and generate reframed outputs without human intervention. The automated workflow includes shot change detection, saliency model application, and genre-specific transformation rules that enable the system to serve itself in the reframing process.
Solution Approach 2:
The patent replaces manual mechanical editing operations with automated computational processes. AI models substitute for human editors in detecting shot changes, identifying salient regions, and applying transformation rules, thereby eliminating the need for skilled manual labor while maintaining reframing quality.
2Extent of automation
If AI-assisted editors are used for partial reframing, then some automation is achieved, but dependency on editing software and powerful rendering machines increases
Solution Approach 1:
The system provides a comprehensive automated reframing solution that handles multiple video formats, aspect ratios, and genres through a single unified platform. The service-based architecture delivers multi-functional capabilities including shot change detection, saliency analysis, text OCR, and genre-specific transformations without requiring users to install or configure complex local software.
Solution Approach 2:
The patent adopts a cloud-based service model where computational resources are provisioned on-demand rather than requiring permanent powerful rendering machines. Users access reframing capabilities through online services, eliminating the need for expensive local hardware investments while maintaining access to advanced AI models.
3Adaptability or versatility
If traditional reframing methods are used, then basic aspect ratio conversion is achieved, but adaptability to different genres and social media platforms is limited
Solution Approach 1:
The system applies different transformation rules and parameters based on the detected video genre and target platform. Genre-specific models customize the reframing approach for different content types (e.g., news, entertainment, sports), while platform-specific rules optimize outputs for various social media aspect ratios and formatting requirements, thereby achieving local adaptation without sacrificing overall processing efficiency.
Data Source
AI summary
The present invention provides a method and system for automatically reframing and transforming videos to different aspect ratios, such that the source video has a fixed resolution and multiple outputs of different aspect ratios are generated. The present invention automatically reframes and transforms horizontal videos into vertical, portrait, square, and/or landscape for social media distribution. The method comprises acquiring input video details; extracting audio; detecting shot change points and extracting frames; detecting salient regions; detecting text; generating and stabilizing viewports; applying genre-specific transformations and generating transformed images; recreating text; obtaining output frames; generating videos and uploading the reframed and transformed videos into cloud. The present invention helps to reframe and transform videos by maintaining the visibility of regions of interest. Furthermore, the present invention helps to create multiple content variants, and ready-to-distribute videos to the social media platforms by retaining contextually important text and moments and rapidly monetizing content.


