Predicting Speaking Intent to Mitigate Teleconference Speech Collision
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Teleconferencing experiences significant challenges with speech collision due to latency in real-time communication, leading to awkwardness and inefficiency as participants often speak simultaneously, as visual cues used in in-person interactions are delayed or not fully facilitated.
Innovation Solution
Implementing a participant computing device with sensors (like accelerometers and gyroscopes) to predict speaking intent through machine-learned models, allowing for early indication of speaking intent to other participants, thereby reducing the likelihood of simultaneous speech.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If real-time audio transmission is used in teleconferencing, then communication efficiency is improved, but speech collision occurs due to latency
Solution Approach 1:
The system performs preliminary detection of speaking intent using sensors (accelerometer, gyroscope, microphone) before the participant actually speaks. The machine learning model analyzes sensor data in advance to predict speaking intent, allowing the system to prepare turn-indication signals before speech occurs, thereby preventing speech collision while maintaining efficient communication flow
2Reliability
If visual cues are used to indicate speaking intent, then speech collision is reduced, but latency in communication increases
Solution Approach 1:
Each participant's computing device performs self-service by locally processing sensor data through the machine learning model to determine speaking intent. The device independently generates and transmits turn-indication signals without requiring constant server mediation, reducing communication latency while maintaining reliable speech collision mitigation through distributed intelligence
3Measurement precision
If sensor-based intent detection is implemented, then speaking intent prediction accuracy is improved, but device complexity increases
Solution Approach 1:
The system merges multiple sensor types (accelerometer, gyroscope, microphone) into a unified detection framework where their data is combined and processed by a single machine learning model. This integration approach improves speaking intent detection accuracy by leveraging complementary sensor information while managing device complexity through centralized processing rather than separate independent systems
Data Source
AI summary
Sensor data is obtained from one or more sensors of a participant computing device. The participant computing device and one or more other participant computing devices are connected to a teleconference orchestrated by a teleconference computing system. Based at least in part on the sensor data, a participant associated with the participant computing device is determined to intend to speak to other participants of the teleconference. Information indicating that the participant intends to speak is provided to one or more of the teleconference computing system or at least one of the one or more other participant computing devices.


