Multi-Camera Participant Correlation for Dynamic Meeting Framing
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Traditional video conferencing systems lack the ability to dynamically engage participants by considering social cues, speaker awareness, and spatial relationships, leading to a limited user experience, especially for those far from the camera.
Innovation Solution
A multi-camera system that uses AI to identify meeting participants, divide the room into zones, and selectively frame interactions, providing a dynamic viewing experience by alternating between speaker and listening shots, and offering spatial context.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Device complexity
If a single camera system is used to capture the meeting environment, then the device complexity is reduced, but the ability to feature non-speaking meeting participants and provide multiple viewing angles is limited
Solution Approach 1:
The patent divides the meeting environment into multiple zones (e.g., front zone near the speaker, side zones, back zone) and assigns different cameras to capture each zone. This segmentation allows the system to provide multiple viewing angles and feature non-speaking participants in different locations, resolving the contradiction between system simplicity and viewing versatility.
Solution Approach 2:
The patent implements a multi-camera system where each camera serves multiple functions: capturing speaker shots, capturing listening shots of non-speaking participants, and providing different viewing angles. This multi-functionality allows a single system to address diverse viewing needs without requiring separate dedicated systems for each function.
2Productivity
If traditional video conferencing systems display only speaking participants, then the focus remains on active communication, but the user experience lacks depth and interaction by excluding non-speaking participants
Solution Approach 1:
The patent dynamically switches between different camera outputs based on the meeting situation. When a participant is speaking, the system displays the speaker shot; when a participant is listening or reacting, the system displays the listening shot. This dynamic adaptation provides a rich user experience that reflects the actual interaction flow, resolving the contradiction between communication efficiency and experience depth.
Solution Approach 2:
The system uses feedback from multiple cameras to determine which participants are speaking and which are listening. By analyzing audio and visual feedback from all cameras, the system can accurately identify speaking participants and switch between shots accordingly, maintaining communication efficiency while including non-speaking participants in the display.
3Area of stationary object
If far end participants are located away from the camera, then the meeting environment can be captured broadly, but the facial expressions and engagement of these participants become difficult to convey
Solution Approach 1:
The patent positions multiple cameras at different locations around the meeting environment, including near the speaker and at far end locations. Each camera captures a specific zone, allowing the system to maintain both broad coverage and close-up views of facial expressions. This spatial segmentation resolves the contradiction between coverage area and expression detection precision.
Solution Approach 2:
The patent adds a temporal dimension to the video transmission by switching between different camera shots (speaker shots and listening shots) based on the interaction flow. This allows far end participants to be featured in close-up views when they are actively listening or reacting, conveying their facial expressions and engagement despite their physical distance from the primary camera.
4Device complexity
If the system displays a limited number of camera angles, then the device complexity is reduced, but the ability to create an engaging and dynamic viewing experience is limited
Solution Approach 1:
The patent implements dynamic shot selection that automatically switches between multiple camera angles based on the meeting situation. The system monitors who is speaking and who is listening, then dynamically transitions between speaker shots and listening shots. This dynamic approach creates an engaging viewing experience without requiring manual control, resolving the contradiction between system complexity and viewing dynamics.
Solution Approach 2:
The system uses automated detection and decision-making to select which camera output to display, eliminating the need for manual camera switching or complex user interfaces. The self-service mechanism automatically manages camera outputs based on detected speaking and listening participants, reducing the operational complexity while maintaining viewing experience diversity.
Data Source
AI summary
Consistent with disclosed embodiments, systems and methods for analyzing video output streams and generating a primary video stream may be provided. Embodiments may include automatically analyzing a first video output stream and a second video output stream, based on at least one identity indicator, to determine whether a first representation of a meeting participant and a second representation of a meeting participant correspond to a common meeting participant. Disclosed embodiments may involve evaluating the first representation and the second representation of the common meeting participant relative to one or more predetermined criteria. Embodiments may involve selecting, based on the evaluation, either the first video output stream or the second video output stream as a source of a framed representation of the common meeting participant to be output as a primary video stream. Furthermore, embodiments may include generating the primary video stream including the framed representation of the common meeting participant.


