Videoconference Participant Identification via Audio-Visual Segmentation
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
In videoconferencing scenarios where multiple participants share a single device, existing technologies fail to accurately identify and represent each participant individually, leading to equity issues, security concerns, and degraded user experience due to unclear participant visibility.
Innovation Solution
Implementing analysis components that use facial detection, voice recognition, and motion detection to identify and separate device-sharing participants, providing individual connections and video streams, and determining active talkers to enhance visibility and security.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If multiple participants share a single device for videoconferencing, then device usage efficiency is improved, but participant identification accuracy deteriorates
Solution Approach 1:
The patent segments the single device's video and audio streams into multiple participant-specific streams by detecting individual faces, voices, and motion patterns. Each participant receives a dedicated video stream showing their own video feed, allowing accurate identification while maintaining shared device usage.
Solution Approach 2:
The system introduces analysis components as intermediaries that process the device's camera and microphone inputs to identify and separate individual participants. These components act as mediators between the shared device and multiple participants, enabling accurate participant differentiation without requiring separate devices.
2Device complexity
If a single name is used to identify all device-sharing participants, then device complexity is reduced, but security deteriorates
Solution Approach 1:
The system dynamically assigns participant identities by continuously analyzing video and audio inputs to detect which individual is currently speaking or active. This dynamic identification allows the simple device to securely distinguish between multiple participants based on real-time behavioral patterns rather than static device information.
3Device complexity
If all device-sharing participants are shown in a single video stream, then device complexity is reduced, but user experience deteriorates
Solution Approach 1:
The patent segments the single video stream into multiple participant-specific video streams, where each participant receives a stream showing their own video feed. This segmentation improves user experience by making each participant easily visible and recognizable, while the system manages the complexity through automated face and motion detection.
Data Source
AI summary
A plurality of device-sharing participants may be detected that are participating in a videoconference via a shared computing device. The detecting of the plurality of device-sharing participants may be performed based, at least in part, on at least one of an audio analysis of captured audio from one or more microphones or a video analysis of captured video from one or more cameras. A plurality of participant connections corresponding to the plurality of device-sharing participants may be joined to the videoconference. Each of the plurality of participant connections may be identified within the videoconference using a respective name. A plurality of video streams and a plurality of audio streams corresponding to the plurality of participant connections may be transmitted, and the plurality of video streams and the plurality of audio streams may be presented to at least one other conference participant.


