Distributed Active Speaker Detection for Frontal Call Video Switching
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conference room video cameras have limitations in capturing high-quality, frontal views of users due to their placement, restricting the user experience in teleconferencing.
Innovation Solution
Employing active speaker detection on user devices to dynamically switch to a high-resolution frontal view from the user's device camera when they speak, using personalized enhancement models to isolate their voice and attenuate other noises.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If a conference room video camera is used to capture video of users, then video coverage of multiple users is achieved, but the video quality and frontal view are insufficient
Solution Approach 1:
The system divides the video capture function into multiple segments: the conference room camera provides general coverage, while individual user devices provide high-quality frontal views. This segmentation allows each camera to operate in its optimal position without requiring complex reconfiguration of the main camera system.
Solution Approach 2:
The solution adds a new dimension to video capture by utilizing the camera on each user's personal device. Instead of relying solely on the single conference room camera, the system incorporates secondary video feeds from multiple devices, creating a multi-dimensional video source architecture that provides both general and close-up views.
2Measurement precision
If the conference room camera provides video to all users, then bandwidth is consumed, but the video quality does not improve when a specific user is speaking
Solution Approach 1:
Instead of transmitting video from all users continuously, the system applies partial action by selectively transmitting video only when a user is detected as the active speaker. The bandwidth-efficient mode transmits video at lower quality or only upon demand, while high-quality video is transmitted selectively when needed, optimizing the balance between quality and bandwidth consumption.
Solution Approach 2:
The system implements feedback through active speaker detection, where audio signals are analyzed to determine who is speaking. This feedback mechanism triggers selective video transmission: when a user speaks, their video is prioritized for transmission to remote participants; when no one speaks or multiple users speak, the system defaults to bandwidth-efficient modes.
3Measurement precision
If personalized audio enhancement is applied to isolate a speaker's voice, then audio quality is improved, but processing complexity increases
Solution Approach 1:
The system performs preliminary action by pre-processing audio signals from all users through noise reduction and enhancement filters before active speaker detection. This preliminary processing simplifies subsequent speaker isolation by reducing background noise and enhancing speech signals in advance, making the personalized enhancement less computationally intensive when applied.
Solution Approach 2:
Each user device performs self-service audio processing by applying noise reduction and enhancement to its own microphone input locally. This distributes the processing load across multiple devices rather than concentrating it on a central server, reducing overall system complexity while maintaining high audio quality for the active speaker.
Data Source
AI summary
This document relates to active speaker detection using distributed devices. For example, the disclosed implementations can employ personal devices of one or more users to detect when those users are speaking during a call with other users. Then, a camera on the personal device can be employed to obtain a front-facing view of the user, which can be provided to other call participants. In some cases, a microphone and/or camera on the user's device are employed to detect when the user is actively speaking.


