Conference Audio Sequencing for Overlapping Speech Resolution
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional video conferencing technologies face challenges with overlapping speech, leading to unintelligible communication and decreased clarity, especially as the number of participants increases, due to technical issues like diminished network bandwidth and latency.
Innovation Solution
A conference server system that identifies overlapping speech by analyzing audio from multiple users and selectively mutes or prioritizes audio based on factors such as speaking order, user information, topic relevance, and speaking frequency, ensuring that only one user's audio is provided in real-time while storing and delaying the other user's audio for later provision.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Loss of information
If all user audio is transmitted in real-time during a communication session, then complete audio content is provided to all users, but overlapping speech causes unintelligible communication and decreased clarity
Solution Approach 1:
The system extracts and identifies overlapping speech segments from the audio stream, separating them from the main audio flow. By detecting when multiple users speak simultaneously, the system isolates these overlapping portions and handles them differently (by muting or delaying), allowing the non-overlapping portions to be transmitted clearly in real-time while preserving complete audio content for later provision.
Solution Approach 2:
The system performs preliminary analysis of incoming audio streams to detect overlapping speech before transmission. By proactively identifying overlapping segments in advance, the system can pre-process the audio by muting or delaying specific users' audio, preventing the transmission of unintelligible overlapping content while maintaining the ability to provide complete audio content afterward.
2Adaptability or versatility
If audio from multiple users is provided simultaneously, then all users can communicate freely, but network bandwidth consumption increases and latency increases
Solution Approach 1:
The system implements a nested structure where the audio transmission system operates at multiple levels: real-time transmission of non-overlapping audio, delayed transmission of overlapping audio, and post-session provision of complete audio content. This nested approach allows the system to maintain communication freedom while optimizing bandwidth usage by transmitting only necessary audio segments in real-time and deferring less critical transmissions.
3Ease of operation
If overlapping speech is allowed in communication sessions, then users can speak naturally without restrictions, but communication clarity and intelligibility decrease
Solution Approach 1:
The system implements feedback mechanisms by monitoring the audio stream for overlapping speech and automatically adjusting audio transmission in response. When overlapping speech is detected, the system provides feedback by muting or delaying specific users' audio, creating a controlled environment that maintains speech intelligibility while allowing users to speak naturally without manual intervention or restrictions.
Data Source
AI summary
Systems, methods, and non-transitory computer-readable media can be configured to determine first audio associated with a first user and second audio associated with a second user, the first user and the second user associated with a communication session. The second audio can be muted based on a determination that the first audio and the second audio overlap. The second audio can be provided based on completion of the first audio.


