Dynamic Face Cropping for Low-Bandwidth Video Streaming
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Streaming video applications face challenges in maintaining high-quality video content due to bandwidth demands, especially in remote working environments with inconsistent network connections, where traditional compression solutions can degrade video quality, particularly for facial reconstruction, and fixed cropping methods limit movement and appear unnatural.
Innovation Solution
The use of a dynamically resized bounding shape to track and encode the face or feature in video content, ensuring efficient encoding and reconstruction while maintaining high quality, using neural networks to adjust the bounding shape based on the face's position and aspect ratio, and embedding cropping information in SEI messages for seamless composition.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Loss of energy
If traditional video compression solutions are used, then bandwidth requirements are reduced, but video quality particularly facial reconstruction quality deteriorates
Solution Approach 1:
The patent divides the video frame into multiple depth layers (foreground, midground, background) and processes each layer separately with different compression techniques. The foreground layer containing facial regions uses higher quality encoding while background layers use more aggressive compression, thus reducing overall bandwidth while maintaining facial quality.
Solution Approach 2:
The patent applies different quality levels to different regions of the video frame based on their importance. Facial regions and regions containing speakers receive high-quality encoding, while less important background areas receive lower quality encoding. This localized quality adjustment reduces total bandwidth consumption while preserving critical video quality.
2Manufacturing precision
If high quality video is maintained, then video quality is preserved, but resource consumption increases significantly
Solution Approach 1:
The patent segments the video processing into multiple depth layers and processes only necessary portions at high quality. By identifying and prioritizing important regions (foreground, speaker regions) and processing them separately, the system maintains high quality where needed while reducing resource consumption in less critical areas.
Solution Approach 2:
The patent implements local quality adjustment by applying high-quality processing only to specific regions containing important visual information such as faces and speakers, while using lower quality processing for background regions. This selective approach maintains overall video quality perception while significantly reducing computational resources required.
3Productivity
If a fixed rectangular crop is used around the face, then encoding efficiency is improved, but the user's movement is limited and the display appears unnatural
Solution Approach 1:
The patent replaces fixed rectangular crops with dynamic depth masks that adapt to the user's movements and facial expressions. These depth masks are generated in real-time and adjust the foreground region boundaries dynamically, allowing natural movement while maintaining encoding efficiency by keeping the foreground region tightly bound to the subject.
Solution Approach 2:
The patent uses feedback from facial detection and tracking algorithms to continuously adjust the depth mask boundaries. The system monitors the subject's position, orientation, and movements, and updates the foreground region accordingly, ensuring the crop adapts to maintain both encoding efficiency and natural appearance during movement.
Data Source
AI summary
Systems and methods relate to facial video encoding and reconstruction, particularly in ultra-low bandwidth settings. In embodiments, a video conferencing or other streaming application uses automatically tracked feature cropping information. A bounding shape size—used to identify the cropped region—varies and is dynamically determined to maintain a proportion for feature reconstruction, such as resizing in the event of a zoom-in on a face (or other feature of interest) or a zoom-out. The tracking scheme may be used to smooth sudden movements, including lateral ones, to generate more natural transitions between frames. Tracking and cropping information (e.g., size and position of the cropped region) may be embedded within an encoded bitstream as supplemental enhancement information (“SEI”), for eventual decoding by a receiver and for compositing a decoded face at a proper location in the applicable stream.


