Video Call Host Module Using Motion Vector Ranking
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional video call systems face limitations in handling multiple participants due to high computational and data transmission bandwidth demands, restricting the number of participants and efficiency in video communication, especially for hearing-impaired users who rely on visual communication methods.
Innovation Solution
A video call host module with a processor and transceiver that decodes and ranks video data using motion indicators from motion vectors to select and display the most active participants, generating a mixed video stream that reduces bandwidth requirements and enhances participant management.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If all participant videos are decoded and transmitted in a video conference, then complete visual communication is provided, but computational power and network bandwidth are excessively consumed
Solution Approach 1:
The patent extracts only the essential visual information (motion vectors) from participant videos rather than transmitting complete video streams. The host system decodes motion vectors to identify active speakers, separating the critical communication information from redundant visual data, thereby reducing computational load while maintaining communication effectiveness.
Solution Approach 2:
The system performs partial decoding of video data by focusing only on motion vector information rather than complete frame decomposition. This partial action approach provides sufficient information to determine speaker activity without the excessive computational resources required for full video decoding and transmission of all participants.
2Reliability
If all participant videos are transmitted in a video conference, then complete visual communication is provided, but network bandwidth is excessively consumed
Solution Approach 1:
The patent extracts only the essential visual information (motion vectors) from participant videos rather than transmitting complete video streams. The host system decodes motion vectors to identify active speakers, separating the critical communication information from redundant visual data, thereby reducing computational load while maintaining communication effectiveness.
Solution Approach 2:
The system performs partial decoding of video data by focusing only on motion vector information rather than complete frame decomposition. This partial action approach provides sufficient information to determine speaker activity without the excessive computational resources required for full video decoding and transmission of all participants.
3Adaptability or versatility
If the display is divided into multiple segments for multiple participants, then all participants can be displayed, but the system cannot accommodate more participants beyond the display segments
Solution Approach 1:
The patent implements a dynamic participant display system where the number of displayed participants is not fixed by display segments but determined by real-time speaker activity detection. The host system continuously monitors motion vectors and dynamically adjusts which participants are displayed based on who is currently speaking, allowing the system to adapt to varying numbers of active participants without being constrained by a fixed segment-based layout.
4Measurement precision
If motion vectors are decoded for all participants, then accurate speaker detection is achieved, but computational resources are excessively consumed
Solution Approach 1:
The patent extracts only the essential visual information (motion vectors) from participant videos rather than transmitting complete video streams. The host system decodes motion vectors to identify active speakers, separating the critical communication information from redundant visual data, thereby reducing computational load while maintaining communication effectiveness.
Data Source
AI summary
A video call host module comprising a processor configured to decode video data corresponding to videos from endpoints and rank the videos based on motion indicators corresponding to each of the endpoints. The motion indicators are calculated from motion vectors corresponding to each of the videos. A predetermined number of highest-ranking videos are selected for display. A method of hosting a video call includes receiving encoded video data including motion vectors. Videos are ranked based on a motion indicator calculated from the motion vectors for each of the videos. Encoded video data is converted to decoded video data, and decoded video data corresponding to a predetermined number of the highest ranking videos is combined to create a single video.


