Split Frame Multistream Video Encoding for Conference Systems
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current video conferencing technologies face challenges in providing a seamless and efficient experience, particularly in switched video scenarios where self-view suppression and continuous presence configurations are desired, leading to increased media-processing resources and latency issues.
Innovation Solution
The implementation of a video conferencing system using a transcoder/MCU with shared and non-shared encoders, which encodes video streams to maintain self-view suppression and reduce media-processing resources by sharing common slice data among participants, allowing for efficient encoding and decoding of continuous presence views.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If a transcoder/MCU decodes individual streams and composes video streams into a single view for each participant (continuous presence configuration), then the conference experience is improved with visual representation of all participants, but the image processing and video encoding resources are significantly increased
Solution Approach 1:
The video frame is segmented into multiple slices, where each slice corresponds to a different participant's video stream. The transcoder/MCU decodes only the necessary slices for each participant rather than fully decoding entire video streams, reducing processing resources while maintaining continuous presence capability.
Solution Approach 2:
Different regions (slices) of the video frame are processed with different quality levels and encoding parameters. The transcoder/MCU applies selective encoding to specific slices based on participant importance and network conditions, optimizing resource usage while maintaining overall conference quality.
2Device complexity
If switched video is used where only the active speaker is visible to all participants, then the media-processing resources are reduced, but the conference lacks a group feel and visual representation of other participants
Solution Approach 1:
The video frame is divided into multiple slices, each containing video from different participants. This allows the system to transmit multiple participant views within a single encoded stream, providing visual representation of the group while maintaining efficient resource usage through shared encoding.
Solution Approach 2:
Multiple participant video streams are merged into a single composed video frame with multiple slices. The transcoder/MCU combines video from multiple participants into one encoded stream that all participants can receive, reducing media-processing resources while maintaining group visibility.
3Loss of time
If self-view suppression is implemented to avoid showing participants themselves, then the latency and distraction are reduced, but the device complexity increases due to additional processing requirements
Solution Approach 1:
The self-view slice is extracted and removed from the composed video frame before encoding. The transcoder/MCU identifies which slice corresponds to each participant and excludes it from their received stream, implementing self-view suppression efficiently within the slice-based encoding framework.
4Adaptability or versatility
If multiple video streams are transmitted to each participant for continuous presence, then the visual representation of all participants is improved, but the network bandwidth and decoding resources at the endpoint are increased
Solution Approach 1:
Multiple participant video streams are merged into a single composed video frame containing multiple slices. Instead of transmitting separate video streams to each participant, the system transmits one encoded stream per participant that contains all necessary video content in slice form, reducing network bandwidth and endpoint decoding resources.
Data Source
AI summary
Techniques for video conferencing include receiving a stream of video slices from a participant, designating the video slices as a primary sub-picture of a frame of video, encoding, with a first encoder, a first secondary sub-picture of the frame of video to obtain an encoded first secondary sub-picture of a frame of video, encoding, with a second encoder, a second secondary sub-picture of the frame of video to obtain an encoded first secondary sub-picture of a frame of video, combining the primary sub-picture with the encoded first secondary sub-picture to obtain a first video stream, combining the primary sub-picture with the encoded second secondary sub-picture to obtain a second video stream, and transmitting the first and second video streams to respective recipients.


