Conference Call Audio Prioritization for Overlapping Speech
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conference calls often experience confusion and inefficiency due to overlapping speech from multiple participants, particularly when users speak simultaneously, leading to delays and the need for repetitive speech.
Innovation Solution
An apparatus and method that prioritize audio data transmission based on sound type classification, using machine-learned models to identify normal speech and restrict transmission of non-prioritized sound types, such as silence, background noise, and interjections, while allowing transmission of non-restricted sounds.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If audio data from multiple users is transmitted simultaneously in conference calls, then all participants can communicate, but overlapping speech causes confusion and inefficiency
Solution Approach 1:
The audio transmission is segmented by prioritizing specific users' audio data over others based on speech detection. The system divides the audio stream into prioritized and non-prioritized segments, allowing selective transmission that prevents overlapping speech confusion while maintaining communication capability.
Solution Approach 2:
The system changes the transmission parameter by introducing audio data prioritization based on speech detection. When speech is detected from a particular user, the system modifies the transmission parameters to prioritize that user's audio data, thereby eliminating overlapping speech issues without sacrificing overall communication.
2Reliability
If audio data prioritization is implemented based on speech detection, then communication clarity improves, but system complexity increases
Solution Approach 1:
The system performs self-service by automatically detecting speech and prioritizing audio data without requiring manual intervention. The speech detection mechanism autonomously identifies when a user is speaking and applies prioritization accordingly, reducing the need for complex manual control systems while maintaining clarity.
Solution Approach 2:
The system manages complexity by implementing prioritization through parameter changes in the audio transmission protocol. By modifying transmission parameters based on speech detection results, the system achieves clarity without requiring fundamental architectural changes, thus limiting the increase in system complexity.
3Loss of information
If all sound types are transmitted equally, then complete audio information is maintained, but non-speech sounds cause interference and reduce clarity
Solution Approach 1:
The system extracts and removes non-speech audio data from the transmission stream. By detecting speech and excluding non-speech sounds (such as background noise, music, or other audio) from the prioritized transmission, the system maintains audio information completeness for relevant content while eliminating interference that reduces clarity.
Solution Approach 2:
The system applies local quality by treating different audio data differently based on its content. Speech audio receives prioritized transmission quality, while non-speech audio is treated with lower priority or excluded. This localized differentiation maintains clarity for important audio information without sacrificing overall audio completeness.
Data Source
AI summary
Example embodiments relate to conference calls. In a method, there may be provided audio data and associated classification data from a plurality of user devices as part of a conference call, the plurality of user devices including at least a first user device and a second user device. The classification data may be indicative of one of a plurality of predetermined sound types represented by the associated audio data in a current time frame. The method may comprise determining that the audio data provided by the first user device is associated with a dominant speaker. The method may comprise controlling transmission of the audio data as part of the conference call, including preventing transmission of audio data provided by at least the second user device based on its associated classification data being indicative of a restricted sound type and the first user device being associated with a dominant speaker.


