Machine Learning Video Conversion for Screen Adaptation
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current video transmission on smartphones is inefficient as horizontal videos do not fill the screen properly, resulting in unnecessary blank spaces when viewed vertically, and vertical videos become smaller when watched on horizontal devices.
Innovation Solution
An apparatus using machine learning to detect objects in horizontal videos, segment them, and rearrange them to create vertical videos that fit any screen size, including additional elements like subtitles or Picture In Picture videos, for optimal viewing on any device.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If horizontal videos are transmitted for TV or monitor viewing, then the video quality is appropriate for large screens, but the video does not fill the screen properly when viewed on smartphones held vertically, resulting in blank spaces
Solution Approach 1:
The system segments the horizontal video content into multiple vertical segments using object detection and recognition models. These segments are then reassembled in a vertical arrangement to create a new vertical video format that fills smartphone screens completely while maintaining the integrity of detected objects and scenes.
Solution Approach 2:
The system transforms the video from horizontal dimension (landscape orientation) to vertical dimension (portrait orientation) by detecting objects in the horizontal video, segmenting them accordingly, and reconstructing them in a vertical arrangement. This dimensional transformation allows the same content to adapt to different device orientations and screen formats.
2Area of stationary object
If vertical videos are recorded for smartphone viewing, then the video fills the screen when watched vertically, but the video becomes smaller with blank spaces when viewed on horizontal devices
Solution Approach 1:
The system creates a universal video solution by generating both horizontal and vertical video versions from the same source content. The transmission unit sends both formats simultaneously, allowing the receiving device to select the appropriate format based on its orientation and screen characteristics, thus achieving multi-device compatibility.
Solution Approach 2:
The system dynamically adapts video format based on device characteristics. Instead of using a fixed video orientation, the system provides both horizontal and vertical versions and allows the receiving device to dynamically select the appropriate format based on whether it is held vertically or horizontally, making the viewing experience adaptive to user behavior.
3Extent of automation
If machine learning-based object detection and video segmentation is implemented, then automatic conversion to vertical video format is achieved, but the device complexity increases
Solution Approach 1:
The system introduces an intermediary processing unit that contains the machine learning models for object detection and video segmentation. This intermediary component acts as a bridge between the received horizontal video and the generated vertical video, automating the conversion process while isolating the complexity within a dedicated module rather than distributing it throughout the entire system.
Data Source
AI summary
Provided is an apparatus for converting and transmitting horizontal and vertical videos based on machine learning. The apparatus for converting and transmitting horizontal and vertical videos based on machine learning includes a receiver unit which receives a horizontal video taken by a camera from the camera; an object detection unit which detects objects in the received horizontal video by using an object recognition model learned by machine learning; a video segmentation unit which generates a plurality of segmentation videos by using the result detected from the object detection unit; a video synthesis unit which arranges the plurality of segmentation videos to fit the size of a screen of a vertical video and generates a vertical video, wherein the segmentation videos may be videos segmented from the horizontal video including at least one object detected from the object detection unit, videos in which the additional videos are synthesized to videos segmented from the horizontal video to comprise at least one object detected from the object detection unit, or the additional videos.


