Push-to-talk floor control for multi-point video conferencing
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current video conferencing systems that prioritize loudness for determining which locations to display during a conference fail to effectively include participants using non-audible communication methods, such as sign language or gestures, and can lead to users raising their voices continuously to be heard.
Innovation Solution
A floor control algorithm that incorporates a push-to-talk button mechanism, allowing users to gain display priority without relying on loudness, by using a combination of loudness-based and gesture-based analysis to determine which video segments are displayed, ensuring that users communicating through non-audible methods can also participate.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Ease of operation
If loudness-based algorithm is used to determine displayed locations, then users who speak louder are prioritized, but participants using non-audible methods (sign language, gestures) cannot effectively cause their location to be displayed
Solution Approach 1:
The floor control mechanism is segmented into multiple independent channels: audio-based floor control for audible speakers and push-to-talk-based floor control for non-audible communicators. This segmentation allows each channel to operate independently, ensuring that sign language users and gesture communicators can gain floor control without needing to raise their voices, while audible speakers continue to be prioritized based on loudness. The segmentation resolves the contradiction by providing tailored access methods for different user types.
Solution Approach 2:
A push-to-talk button serves as an intermediary device that bridges the gap between non-audible communicators and the floor control system. When a user presses the button, it generates a floor control request that is processed by the conference system, allowing the user to gain display priority regardless of audio output. This intermediary mechanism enables participants using sign language or gestures to effectively participate in the conference without relying on audible speech.
2Productivity
If loudness-based algorithm is used to determine displayed locations, then audible speakers are prioritized, but users may raise their voices continually to be heard
Solution Approach 1:
The harmful feedback loop of continuous voice raising is extracted and eliminated by introducing an alternative floor control mechanism. The push-to-talk button provides a direct control method that does not depend on audio loudness, allowing users to gain floor control without needing to increase their voice volume. This extraction removes the causal link between speaking volume and display priority, thereby eliminating the harmful behavior of continuous voice raising while maintaining effective communication.
3Adaptability or versatility
If push-to-talk button mechanism is added, then non-audible communicators can gain floor control, but system complexity increases
Solution Approach 1:
The push-to-talk floor control mechanism is merged with the existing audio-based floor control system into a unified processing architecture. Both mechanisms share common components such as the floor control decision logic, video switching infrastructure, and conference bridge software. By merging these systems, the patent achieves compatibility with multiple communication methods while minimizing the increase in overall system complexity through resource sharing and integrated processing.
Data Source
AI summary
In one embodiment, a conference with multiple end points is provided. At the locations, multiple screens may be configured to display video from a portion of the multiple end points. Video from multiple locations is output onto the multiple screens, such as video streams from N different segments are output on N different screens. The video output may be determined based on a first dimension of the floor control algorithm. A push-to-talk input may then be received from a button. A video segment associated with the push-to-talk button is then determined and the video segment is output on one of the multiple screens in response to receiving the push-to-talk input. The push-to-talk input may be used by users that cannot actively participate in the first dimension of the floor control algorithm. For example, users using sign language cannot speak louder and thus by using the push to talk button or hand gestures can indicate their desire to be switched in as one of the displayed segments.


