Multi-Camera Participant Correlation for Dynamic Meeting Framing

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Traditional video conferencing systems lack the ability to dynamically engage participants by considering social cues, speaker awareness, and spatial relationships, leading to a limited user experience, especially for those far from the camera.

Innovation Solution

A multi-camera system that uses AI to identify meeting participants, divide the room into zones, and selectively frame interactions, providing a dynamic viewing experience by alternating between speaker and listening shots, and offering spatial context.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Device complexity

If a single camera system is used to capture the meeting environment, then the device complexity is reduced, but the ability to feature non-speaking meeting participants and provide multiple viewing angles is limited

Engineering Contradiction:
Improvecamera system complexityVSAvoidviewing angle variety
Core Design Contradiction:
Device complexityVSAdaptability or versatility

Solution Approach 1:

The patent divides the meeting environment into multiple zones (e.g., front zone near the speaker, side zones, back zone) and assigns different cameras to capture each zone. This segmentation allows the system to provide multiple viewing angles and feature non-speaking participants in different locations, resolving the contradiction between system simplicity and viewing versatility.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent implements a multi-camera system where each camera serves multiple functions: capturing speaker shots, capturing listening shots of non-speaking participants, and providing different viewing angles. This multi-functionality allows a single system to address diverse viewing needs without requiring separate dedicated systems for each function.

Inventive Principle:
Principle #6Universality (Multi-functionality)

2Productivity

If traditional video conferencing systems display only speaking participants, then the focus remains on active communication, but the user experience lacks depth and interaction by excluding non-speaking participants

Engineering Contradiction:
Improvecommunication efficiencyVSAvoiduser experience depth
Core Design Contradiction:
ProductivityVSAdaptability or versatility

Solution Approach 1:

The patent dynamically switches between different camera outputs based on the meeting situation. When a participant is speaking, the system displays the speaker shot; when a participant is listening or reacting, the system displays the listening shot. This dynamic adaptation provides a rich user experience that reflects the actual interaction flow, resolving the contradiction between communication efficiency and experience depth.

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

The system uses feedback from multiple cameras to determine which participants are speaking and which are listening. By analyzing audio and visual feedback from all cameras, the system can accurately identify speaking participants and switch between shots accordingly, maintaining communication efficiency while including non-speaking participants in the display.

Inventive Principle:
Principle #23Feedback

3Area of stationary object

If far end participants are located away from the camera, then the meeting environment can be captured broadly, but the facial expressions and engagement of these participants become difficult to convey

Engineering Contradiction:
Improvecoverage areaVSAvoidfacial expression detection
Core Design Contradiction:
Area of stationary objectVSMeasurement precision

Solution Approach 1:

The patent positions multiple cameras at different locations around the meeting environment, including near the speaker and at far end locations. Each camera captures a specific zone, allowing the system to maintain both broad coverage and close-up views of facial expressions. This spatial segmentation resolves the contradiction between coverage area and expression detection precision.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent adds a temporal dimension to the video transmission by switching between different camera shots (speaker shots and listening shots) based on the interaction flow. This allows far end participants to be featured in close-up views when they are actively listening or reacting, conveying their facial expressions and engagement despite their physical distance from the primary camera.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

4Device complexity

If the system displays a limited number of camera angles, then the device complexity is reduced, but the ability to create an engaging and dynamic viewing experience is limited

Engineering Contradiction:
Improvecamera output managementVSAvoidviewing experience dynamics
Core Design Contradiction:
Device complexityVSAdaptability or versatility

Solution Approach 1:

The patent implements dynamic shot selection that automatically switches between multiple camera angles based on the meeting situation. The system monitors who is speaking and who is listening, then dynamically transitions between speaker shots and listening shots. This dynamic approach creates an engaging viewing experience without requiring manual control, resolving the contradiction between system complexity and viewing dynamics.

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

The system uses automated detection and decision-making to select which camera output to display, eliminating the need for manual camera switching or complex user interfaces. The self-service mechanism automatically manages camera outputs based on detected speaking and listening participants, reducing the operational complexity while maintaining viewing experience diversity.

Inventive Principle:
Principle #25Self-service

Data Source

PatentUS20250336233A1Systems and methods for correlating individuals across outputs of a multi-camera system and framing interactions between meeting participants
Publication Date: 2025.10.30 HUDDLY AS
  • US20250336233A1 patent drawing
  • US20250336233A1 patent drawing
  • US20250336233A1 patent drawing

AI summary

Consistent with disclosed embodiments, systems and methods for analyzing video output streams and generating a primary video stream may be provided. Embodiments may include automatically analyzing a first video output stream and a second video output stream, based on at least one identity indicator, to determine whether a first representation of a meeting participant and a second representation of a meeting participant correspond to a common meeting participant. Disclosed embodiments may involve evaluating the first representation and the second representation of the common meeting participant relative to one or more predetermined criteria. Embodiments may involve selecting, based on the evaluation, either the first video output stream or the second video output stream as a source of a framed representation of the common meeting participant to be output as a primary video stream. Furthermore, embodiments may include generating the primary video stream including the framed representation of the common meeting participant.