Concurrent Speech Handling via Dynamic Delay and Priority Policies
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional conferencing systems face challenges in handling concurrent speech, leading to user frustration and confusion due to overlapping voices, where speakers' words may not be heard for extended periods or are frequently interrupted, reducing the effectiveness of communication.
Innovation Solution
A method and system that selectively adjust and output speech based on predetermined thresholds and participant priorities, allowing speech to be delayed or dropped to prevent overlap, ensuring orderly communication by prioritizing speech delivery.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Object-affected harmful factors
If speech is broadcast in serial first-in first-out manner, then speech overlap is eliminated, but system complexity and processing overhead increase significantly
Solution Approach 1:
The system changes the parameter of speech output timing by introducing variable delays based on detected speech patterns. When concurrent speech is detected, the system dynamically adjusts the delay parameter for subsequent speakers rather than using a fixed first-in-first-out queue, thereby reducing overlap without complex reordering logic
Solution Approach 2:
The system introduces an intermediary speech detection and control mechanism that monitors incoming speech signals and mediates the output timing of multiple participants. This intermediary layer analyzes speech patterns and applies appropriate delays to prevent overlap, simplifying the overall system architecture compared to full serial processing
2Object-affected harmful factors
If speech is delayed to prevent overlap, then speech clarity improves, but user frustration increases due to extended delays
Solution Approach 1:
The system applies partial delay action by only delaying speech when concurrent speech patterns are detected, rather than uniformly delaying all speech. The delay is applied selectively and minimally—just enough to prevent overlap—thereby maintaining speech clarity while reducing user frustration from excessive delays
Solution Approach 2:
The delay mechanism is dynamic rather than static, adjusting in real-time based on detected speech patterns. The system continuously monitors for concurrent speech and applies variable delays as needed, making the delay adaptive to actual communication conditions rather than imposing fixed delays that cause user frustration
3Speed
If frequent speech interruptions are allowed, then user responsiveness is maintained, but communication effectiveness decreases due to repeated speech
Solution Approach 1:
The system performs preliminary detection of speech patterns before outputting speech. By detecting potential concurrent speech in advance and applying preventive delays, the system avoids the need for frequent interruptions and repetitions, thereby maintaining both user responsiveness and communication effectiveness
Solution Approach 2:
The system uses feedback from speech pattern detection to continuously adjust speech output timing. By monitoring for concurrent speech and responding with appropriate delays, the system creates a feedback loop that prevents overlapping speech and reduces the need for interruptions and repetitions, improving overall communication effectiveness
Data Source
AI summary
Systems and methods are provided for handling concurrent speech in which temporally overlapping first speech data and second speech data is received from respective first and second participants of a session. A speech policy applied to the speech data specifies dropping the second speech when it interrupts the first speech within a first interval of the first speech data. The first interval is temporally bounded by the beginning of the first speech and a first predetermined amount of time after the beginning of the first speech. The speech policy specifies outputting the first speech data and then outputting the second speech data when the second speech data interrupts a second interval of the first speech data. The second interval of the first speech data is temporally bounded by the end of the first speech data and a second predetermined amount of time before the end of the first speech data.


