Audio Conferencing Utterance Segmentation and Playback
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing audio conferencing systems face issues such as interruptions, difficulty in understanding multiple speakers, and poor channel conditions, leading to distractions and loss of information, where participants may miss important content or struggle to contribute their thoughts effectively.
Innovation Solution
A system that detects the start and end of utterances from multiple conference participants, allowing for sequential playback of audio clips, enabling participants to listen to missed content without interruptions by switching between live audio and recorded playback, and storing audio clips for later access and tagging.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Speed
If real-time audio conferencing is used to connect participants, then communication speed is improved, but interruptions and information loss increase
Solution Approach 1:
The audio conferencing system segments the continuous audio stream into individual utterance clips for each participant. This segmentation allows the system to process, store, and replay utterances independently, preventing information loss during interruptions while maintaining real-time communication capabilities.
Solution Approach 2:
The system records and stores audio clips of each participant's utterances as they occur during the conference. By preparing these clips in advance, participants can later replay missed content without losing information, even if they were interrupted or distracted during the live session.
2Productivity
If multiple participants speak simultaneously in real-time conferencing, then conversation flow is improved, but understanding difficulty increases
Solution Approach 1:
The system separates simultaneous utterances from different participants into distinct audio clips. Each participant's speech is captured and stored as an individual segment, allowing listeners to replay and understand each utterance clearly without the confusion of overlapping voices, while still maintaining natural conversation flow.
3Speed
If live audio playback is used during conference, then real-time communication is improved, but distractions and missed content increase
Solution Approach 1:
The system records and stores audio clips of all participant utterances during the conference. This preliminary recording allows participants who become distracted or interrupted to later replay the exact content they missed without affecting real-time communication for other participants.
Solution Approach 2:
The system introduces an intermediary playback mechanism that bridges real-time communication and information retention. Audio clips serve as an intermediary between the live conference and participant understanding, allowing missed content to be recovered without disrupting the real-time flow for active participants.
4Loss of information
If sequential playback of audio clips is implemented, then information completeness is improved, but time consumption increases
Solution Approach 1:
The system allows participants to selectively replay only the specific audio clips they missed or need to review, rather than requiring them to listen to the entire conference sequentially. This partial action approach maintains information completeness while significantly reducing time consumption compared to full replay.
Data Source
Figure 1
Figure 2
Figure 3A
AI summary
Systems and methods are disclosed herein for improving audio conferencing services. One aspect relates to processing audio content of a conference. A first audio signal is received from a first conference participant, and a start and an end of a first utterance by the first conference participant are detected from the first audio signal. A second audio signal is received from a second conference participant, and a start and an end of a second utterance by the second conference participant is detected from the second audio signal. The second conference participant is provided with at least a portion of the first utterance, wherein at least one of start time, start point, and duration is determined based at least in part on the start, end, or both, of the second utterance.