Nonlinear Video Retargeting for Aspect Ratio Adaptation
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing video retargeting techniques, such as linear downscaling, cropping, and panning, fail to account for the underlying content, leading to suboptimal video output on target platforms with different aspect ratios, as they do not preserve specific image information like face and body proportions.
Innovation Solution
The implementation of nonlinear warps in video retargeting, which adaptively deform pixel shapes based on saliency information and user-specified features to prioritize important image elements, generating a scalable bitstream that can be decoded for various target platforms, including 16:9 and 4:3 aspect ratios.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If linear downscaling is used for video retargeting, then the video content can be scaled to fit target device frames, but the aspect ratio and important image information such as face and body proportions are distorted
Solution Approach 1:
The patent applies different processing treatments to different regions of the video content based on their importance. Saliency detection identifies important regions (such as faces and bodies) that require preservation of proportions, while less important regions (such as backgrounds) can be freely scaled. This local differentiation resolves the contradiction by maintaining manufacturing precision for critical elements while achieving adaptability for the overall video output.
Solution Approach 2:
The video content is segmented into important and less important regions through saliency detection. This segmentation allows the system to apply different scaling strategies to different parts of the video, preserving aspect ratios for important elements while adapting the overall content to target platform requirements, thus resolving the contradiction between adaptability and information preservation.
2Adaptability or versatility
If cropping and panning techniques are used for video retargeting, then the video content can be adapted to different aspect ratios, but important image information may be removed or lost
Solution Approach 1:
By identifying important regions through saliency detection and applying local quality principles, the patent ensures that these regions are preserved and not cropped or panned away. The system adapts the video to different aspect ratios while maintaining the integrity of important image information, thus preventing information loss while achieving adaptability.
Solution Approach 2:
The patent performs preliminary saliency detection and important region identification before applying any cropping or panning operations. This preliminary action ensures that important image information is identified and protected from removal, allowing the system to adapt to different aspect ratios without losing critical content.
3Manufacturing precision
If nonlinear warps are used to preserve important image information, then face and body proportions are maintained, but less important information like backgrounds may be degraded
Solution Approach 1:
The patent applies nonlinear warps selectively to important regions identified through saliency detection, preserving face and body proportions with high manufacturing precision. Less important regions such as backgrounds are allowed to degrade or be simplified, as they contribute less to the overall quality perception. This local differentiation resolves the contradiction by concentrating precision where it matters most.
Solution Approach 2:
The patent changes the processing parameters applied to different regions of the video. Important regions undergo nonlinear warping with high precision parameters to preserve proportions, while less important regions are processed with lower precision parameters that allow for degradation. This parameter differentiation resolves the contradiction between preserving important information and managing overall information loss.
Data Source
Figure 1
Figure 2~3
Figure 4~5
AI summary
Systems, methods and a carrier medium are disclosed for performing scalable video coding. In one embodiment, non-linear functions are used to predict source video data using retargeted video data. Differences may be determined between the predicted video data and the source video data. The retargeted video data, the non-linear functions, and the differences may be jointly encoded into a scalable bitstream. The scalable bitstream may be transmitted and selectively decoded to produce output video for one of a plurality of predefined target platforms.