Videoconferencing Endpoint Selective Video Compositing
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
In multi-way videoconferencing, existing systems face challenges in efficiently managing video streams and resources, leading to increased clutter and reduced scalability, as they attempt to display all participants simultaneously, which limits the number of supported participants and degrades video quality.
Innovation Solution
A videoconferencing endpoint, acting as a Multipoint Control Unit (MCU), selectively composites video images based on criteria such as the last N talking participants, ignoring non-displayed endpoints to conserve resources and dynamically adjust the composite video image by adding or removing participants based on talker detection, allowing for more participants to be supported and improving video quality.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Loss of information
If all participants are displayed simultaneously in the composite video image, then complete participant visibility is achieved, but device complexity and resource consumption increase significantly
Solution Approach 1:
The patent extracts only the necessary video streams for display by identifying and selecting only the last N talking participants whose video streams are composited into the composite video image, while ignoring video streams from non-speaking participants. This reduces the number of decoders required at the master endpoint while maintaining essential participant visibility.
Solution Approach 2:
Instead of processing all participant video streams equally, the patent applies partial action by selectively processing only the video streams of active speakers (last N talkers). This partial processing approach reduces computational resources while maintaining the critical function of displaying speaking participants.
2Loss of information
If all participants are displayed in the composite video image, then complete participant visibility is achieved, but the video quality and resolution decrease due to resource constraints
Solution Approach 1:
The patent extracts and processes only the video streams of the last N talking participants, ignoring video streams from non-speaking participants. This selective extraction allows the system to allocate more computational resources to decoding and displaying the video streams of active speakers, thereby improving video quality and resolution for the displayed participants.
3Loss of information
If all participants are displayed simultaneously, then complete participant visibility is achieved, but visual clutter increases and reduces scalability
Solution Approach 1:
The patent removes non-essential video streams from the composite video image by displaying only the last N talking participants. This extraction of necessary information (speaking participants) while omitting non-essential information (non-speaking participants) reduces visual clutter and improves the scalability of the videoconferencing system.
Solution Approach 2:
The patent implements dynamic participant selection where the composite video image automatically updates to reflect the current set of last N talking participants. This dynamic adjustment ensures that the displayed participants change based on who is currently speaking, maintaining visual clarity and reducing clutter as participants join and leave the conference.
Data Source
AI summary
In some embodiments, a videoconferencing endpoint may be an MCU (Multipoint Control Unit) or may include embedded MCU functionality. In various embodiments, the endpoint may thus conduct a videoconference by receiving/compositing video and audio from multiple videoconference endpoints. The endpoint may select a subset of endpoints and form a composite video image from the subset of the videoconference endpoints to send to the other videoconference endpoints. In some embodiments, the subset of endpoints that are selected for compositing into the composite video image may be selected according to criteria such as the last N talking participants. In some embodiments, the master endpoint may request the non-talker endpoints to stop sending video to help conserve the resources on the master endpoint. In some embodiments, the master endpoint may ignore video from endpoints that are not being displayed.


