Distributed Active Speaker Detection for Frontal Call Video Switching

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conference room video cameras have limitations in capturing high-quality, frontal views of users due to their placement, restricting the user experience in teleconferencing.

Innovation Solution

Employing active speaker detection on user devices to dynamically switch to a high-resolution frontal view from the user's device camera when they speak, using personalized enhancement models to isolate their voice and attenuate other noises.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If a conference room video camera is used to capture video of users, then video coverage of multiple users is achieved, but the video quality and frontal view are insufficient

Engineering Contradiction:
Improvevideo qualityVSAvoidcamera placement
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The system divides the video capture function into multiple segments: the conference room camera provides general coverage, while individual user devices provide high-quality frontal views. This segmentation allows each camera to operate in its optimal position without requiring complex reconfiguration of the main camera system.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The solution adds a new dimension to video capture by utilizing the camera on each user's personal device. Instead of relying solely on the single conference room camera, the system incorporates secondary video feeds from multiple devices, creating a multi-dimensional video source architecture that provides both general and close-up views.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

2Measurement precision

If the conference room camera provides video to all users, then bandwidth is consumed, but the video quality does not improve when a specific user is speaking

Engineering Contradiction:
Improveactive speaker video qualityVSAvoidbandwidth usage
Core Design Contradiction:
Measurement precisionVSLoss of energy

Solution Approach 1:

Instead of transmitting video from all users continuously, the system applies partial action by selectively transmitting video only when a user is detected as the active speaker. The bandwidth-efficient mode transmits video at lower quality or only upon demand, while high-quality video is transmitted selectively when needed, optimizing the balance between quality and bandwidth consumption.

Inventive Principle:
Principle #16Partial or excessive action

Solution Approach 2:

The system implements feedback through active speaker detection, where audio signals are analyzed to determine who is speaking. This feedback mechanism triggers selective video transmission: when a user speaks, their video is prioritized for transmission to remote participants; when no one speaks or multiple users speak, the system defaults to bandwidth-efficient modes.

Inventive Principle:
Principle #23Feedback

3Measurement precision

If personalized audio enhancement is applied to isolate a speaker's voice, then audio quality is improved, but processing complexity increases

Engineering Contradiction:
Improveaudio qualityVSAvoidprocessing complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The system performs preliminary action by pre-processing audio signals from all users through noise reduction and enhancement filters before active speaker detection. This preliminary processing simplifies subsequent speaker isolation by reducing background noise and enhancing speech signals in advance, making the personalized enhancement less computationally intensive when applied.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

Each user device performs self-service audio processing by applying noise reduction and enhancement to its own microphone input locally. This distributes the processing load across multiple devices rather than concentrating it on a central server, reducing overall system complexity while maintaining high audio quality for the active speaker.

Inventive Principle:
Principle #25Self-service

Data Source

PatentUS12494207B2Active speaker detection using distributed devices
Publication Date: 2025.12.09 MICROSOFT TECHNOLOGY LICENSING LLC
  • US12494207B2 patent drawing
  • US12494207B2 patent drawing
  • US12494207B2 patent drawing

AI summary

This document relates to active speaker detection using distributed devices. For example, the disclosed implementations can employ personal devices of one or more users to detect when those users are speaking during a call with other users. Then, a camera on the personal device can be employed to obtain a front-facing view of the user, which can be provided to other call participants. In some cases, a microphone and/or camera on the user's device are employed to detect when the user is actively speaking.