Scalable Video Encoding for Multi-View Conferencing
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current video conferencing systems face challenges in providing high-quality video content for active speakers while maintaining bandwidth efficiency, as they often employ extensive encoding techniques that decrease video quality for all participants, including non-active speakers.
Innovation Solution
The implementation of scalable video coding (SVC) in a multi-view camera system, where multiple cameras capture different image areas, and the video content is encoded with a lower SVC layer for non-active areas and a higher SVC layer for the active area, allowing for higher quality reconstruction of the active speaker's video while maintaining lower quality for others, thus optimizing bandwidth usage.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Loss of energy
If extensive encoding techniques are employed to reduce bandwidth, then bandwidth efficiency is improved, but video quality for all participants deteriorates
Solution Approach 1:
The patent applies different encoding qualities to different regions of the video content. The active speaker region is encoded at high quality using base layer + enhancement layers, while non-active regions are encoded at lower quality. This resolves the contradiction by making video quality local rather than uniform, ensuring high quality where needed (active speaker) while conserving bandwidth in less critical areas.
Solution Approach 2:
The video content is segmented into multiple regions based on speaker activity. The encoding process segments the bitstream into base layer data and enhancement layer data, where the base layer provides acceptable quality for all regions and enhancement layers provide quality improvements only for the active speaker region. This segmentation allows differential quality distribution.
2Manufacturing precision
If uniform high quality encoding is applied to all video content, then video quality for all participants is improved, but bandwidth consumption increases significantly
Solution Approach 1:
Instead of applying uniform high quality encoding across the entire video content, the system applies high quality encoding locally only to the active speaker region while using lower quality encoding for non-active regions. This resolves the contradiction by making quality distribution selective rather than uniform.
Solution Approach 2:
The system applies enhancement encoding partially only to the extent necessary for the active speaker region rather than excessively to the entire video content. The base layer provides sufficient quality for non-active regions without requiring enhancement layers, thus avoiding unnecessary bandwidth consumption.
3Manufacturing precision
If video cameras are positioned closer to active participants to increase resolution, then video resolution for participants is improved, but device positioning complexity and tracking requirements increase
Solution Approach 1:
The system uses a fixed multi-camera array where each camera captures a specific field of view. Rather than moving cameras to achieve high resolution, the system processes the fixed camera feeds to identify active speakers and applies enhancement encoding locally to those regions. This maintains fixed device positioning while achieving high resolution for active speakers through selective encoding rather than physical camera movement.
Data Source
AI summary
The present invention employs scalable video coding (SVC) in a multi-view camera system, which is particularly suited for video conferencing. Multiple cameras are oriented to capture video content of different image areas and generate corresponding original video streams that provide video content of the image areas. An active one of the image areas may be identified at any time by analyzing the audio content originating from the different image areas and selecting the image area that is associated with the most dominant speech activity.


