Video Retargeting via Sound Localization and Preserving Maps
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing video conferencing technologies fail to accurately adapt video streams to display different aspect ratios and effectively preserve important regions, such as active speakers and body language, leading to reduced field of view and quality of experience due to cropping or black borders.
Innovation Solution
The method employs sound localization to determine the active speaker and create a preserving map, allowing for nonlinear video retargeting that prioritizes important regions while maintaining aspect ratio and allowing for simultaneous display of original and retargeted videos.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Loss of information
If video is cropped to preserve important regions, then important information is preserved, but field of view is severely restricted and other parts of interest are lost
Solution Approach 1:
The patent applies local quality by creating a preserving map that assigns different importance weights to different regions of the video frame. Important regions (active speaker, body language) are assigned higher weights to be preserved, while less important regions are assigned lower weights and can be distorted or cropped. This allows selective preservation of critical information while maintaining overall field of view.
2Device complexity
If rectangular region of interest is used for cropping, then detection is simplified, but multiple persons of interest cannot be displayed simultaneously
Solution Approach 1:
The patent segments the video frame into multiple independent regions of interest rather than using a single rectangular crop. Each detected person or important object becomes a separate region with its own bounding box and importance weight. This allows the system to handle multiple persons of interest simultaneously while keeping the detection algorithm relatively simple.
3Ease of manufacture
If linear scaling is used to adapt aspect ratio, then implementation is simple, but black borders are introduced reducing field of view
Solution Approach 1:
The patent replaces static linear scaling with dynamic nonlinear scaling that adapts to content importance. The preserving map guides a warp function that dynamically adjusts the scaling factor for each region based on its importance weight. This allows the system to fill the entire display area without black borders while preserving important regions, making the field of view adaptable rather than fixed.
4Manufacturing precision
If content-aware resizing is applied to raster images, then quality is improved, but conversion to vector is required resulting in limited quality for natural videos
Solution Approach 1:
The patent substitutes the mechanical conversion process (raster to vector) with a direct raster image processing approach. Instead of converting raster video to vector format and then applying content-aware resizing, the method works directly on raster images using a preserving map to guide nonlinear scaling. This avoids the quality degradation associated with vector conversion while still achieving content-aware resizing on natural videos.
Data Source
AI summary
According to embodiments of the present invention, sound localization is used to determine the active speaker in a video conference. A network element uses the localization to determine which regions of the image that should be preserved and retargets the video accordingly. By providing a retargeted video where the speaker is more visible, a better user experience is achieved.


