Scalable Video Encoding for Multi-View Conferencing

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current video conferencing systems face challenges in providing high-quality video content for active speakers while maintaining bandwidth efficiency, as they often employ extensive encoding techniques that decrease video quality for all participants, including non-active speakers.

Innovation Solution

The implementation of scalable video coding (SVC) in a multi-view camera system, where multiple cameras capture different image areas, and the video content is encoded with a lower SVC layer for non-active areas and a higher SVC layer for the active area, allowing for higher quality reconstruction of the active speaker's video while maintaining lower quality for others, thus optimizing bandwidth usage.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Loss of energy

If extensive encoding techniques are employed to reduce bandwidth, then bandwidth efficiency is improved, but video quality for all participants deteriorates

Engineering Contradiction:
Improvebandwidth consumptionVSAvoidvideo quality
Core Design Contradiction:
Loss of energyVSManufacturing precision

Solution Approach 1:

The patent applies different encoding qualities to different regions of the video content. The active speaker region is encoded at high quality using base layer + enhancement layers, while non-active regions are encoded at lower quality. This resolves the contradiction by making video quality local rather than uniform, ensuring high quality where needed (active speaker) while conserving bandwidth in less critical areas.

Inventive Principle:
Principle #3Local quality

Solution Approach 2:

The video content is segmented into multiple regions based on speaker activity. The encoding process segments the bitstream into base layer data and enhancement layer data, where the base layer provides acceptable quality for all regions and enhancement layers provide quality improvements only for the active speaker region. This segmentation allows differential quality distribution.

Inventive Principle:
Principle #1Segmentation

2Manufacturing precision

If uniform high quality encoding is applied to all video content, then video quality for all participants is improved, but bandwidth consumption increases significantly

Engineering Contradiction:
Improvevideo qualityVSAvoidbandwidth consumption
Core Design Contradiction:
Manufacturing precisionVSLoss of energy

Solution Approach 1:

Instead of applying uniform high quality encoding across the entire video content, the system applies high quality encoding locally only to the active speaker region while using lower quality encoding for non-active regions. This resolves the contradiction by making quality distribution selective rather than uniform.

Inventive Principle:
Principle #3Local quality

Solution Approach 2:

The system applies enhancement encoding partially only to the extent necessary for the active speaker region rather than excessively to the entire video content. The base layer provides sufficient quality for non-active regions without requiring enhancement layers, thus avoiding unnecessary bandwidth consumption.

Inventive Principle:
Principle #16Partial or excessive action

3Manufacturing precision

If video cameras are positioned closer to active participants to increase resolution, then video resolution for participants is improved, but device positioning complexity and tracking requirements increase

Engineering Contradiction:
Improvevideo resolutionVSAvoidcamera positioning and tracking
Core Design Contradiction:
Manufacturing precisionVSDevice complexity

Solution Approach 1:

The system uses a fixed multi-camera array where each camera captures a specific field of view. Rather than moving cameras to achieve high resolution, the system processes the fixed camera feeds to identify active speakers and applies enhancement encoding locally to those regions. This maintains fixed device positioning while achieving high resolution for active speakers through selective encoding rather than physical camera movement.

Inventive Principle:
Principle #3Local quality

Data Source

PatentUS8791978B2Scalable video encoding in a multi-view camera system
Publication Date: 2014.07.29 APPLE INC
  • US8791978B2 patent drawing
  • US8791978B2 patent drawing
  • US8791978B2 patent drawing

AI summary

The present invention employs scalable video coding (SVC) in a multi-view camera system, which is particularly suited for video conferencing. Multiple cameras are oriented to capture video content of different image areas and generate corresponding original video streams that provide video content of the image areas. An active one of the image areas may be identified at any time by analyzing the audio content originating from the different image areas and selecting the image area that is associated with the most dominant speech activity.