Conference Server Active Speaker Identification
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
In centralized VoIP conference calls, it is difficult for participants to identify who is speaking, especially in large groups where users are unfamiliar with each other, due to the inclusion of all SSRC identifiers in the CSRC list, making it impossible to distinguish active from inactive participants.
Innovation Solution
A method where a conference call server receives data packets from terminals, determines active voice data providers, and transmits mixed voice data with distinguishable identifiers for active participants, allowing terminals to recognize and present the active speaker to users.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Loss of information
If all SSRC identifiers are included in the CSRC list of RTP packets, then complete participant information is provided to all terminals, but it becomes impossible to distinguish active speakers from inactive participants
Solution Approach 1:
The patent extracts only the SSRC identifiers of active speakers from the complete CSRC list and transmits them separately in RTCP UPDATE packets. This allows receiving terminals to distinguish active speakers without being overwhelmed by all participant identifiers, resolving the contradiction between providing complete information and enabling easy identification.
Solution Approach 2:
The patent segments the participant information transmission into two parts: the complete CSRC list is maintained for reference, while a separate, updated list of only active speakers is transmitted via RTCP UPDATE packets. This segmentation allows terminals to access both complete information and simplified identification as needed.
2Difficulty of detecting and measuring
If the conference call server processes and transmits updated SSRC information to all terminals, then active speaker identification is improved, but network traffic and processing load increase
Solution Approach 1:
Instead of transmitting complete participant lists repeatedly, the patent transmits only the necessary subset information (active speaker SSRCs) in RTCP UPDATE packets. This partial action approach provides the needed identification capability while minimizing network traffic compared to full information repetition.
3Ease of operation
If RTCP UPDATE packets are sent to all terminals with active speaker information, then real-time speaker identification is achieved, but communication overhead increases
Solution Approach 1:
The patent uses the existing RTCP protocol framework for UPDATE packets, which can serve multiple functions: maintaining participant lists, identifying active speakers, and providing synchronization information. This multi-functionality approach avoids creating entirely new protocol mechanisms while achieving real-time speaker identification.
Data Source
AI summary
The invention relates to a method for managing a packet switched, centralized conference call between a plurality of terminals 13. In order to enable an enhancement of the user comfort, it is proposed that the method comprises at a conference call server 12 receiving data packets from all terminals 13. Based on these data packets, then at least one terminal 13 currently providing voice data is determined. In a next step, the data received in the data packets is mixed, and the mixed data is inserted into new data packets together with at least one identifier associated to one of the terminals 13 which were determined to provide voice data, such that the at least one identifier can be distinguished from any other information in the data packets. Finally, the new data packets are transmitted to terminals 13 participating in the conference call. The invention relates equally to a corresponding server and to a corresponding terminal.


