Group and Conversational Framing for Speaker Tracking

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current video conference systems focus on framing the active speaker, which limits the contextual understanding for far-end participants by not showing reactions and body language of other participants, and often group people sitting close together as a single entity, leading to distracting view switches.

Innovation Solution

Implementing group and conversational framing techniques that detect participant proximity to group participants and adjust camera views to frame the active speaker with nearby participants, reducing unnecessary view switches and enhancing contextual understanding.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If the system frames only the active speaker in close-up views, then the speaker tracking function is improved, but the contextual understanding for far-end participants deteriorates because reactions and body language of other participants are not shown

Engineering Contradiction:
Improvespeaker tracking accuracyVSAvoidcontextual understanding
Core Design Contradiction:
Measurement precisionVSLoss of information

Solution Approach 1:

The system segments the video feed into multiple regions of interest, identifying both the active speaker and nearby participants as separate focal points. This allows the system to maintain close-up framing of the speaker while also capturing and transmitting reactions from other participants, thereby preserving contextual information without sacrificing speaker tracking accuracy.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The system transitions from a single-dimension close-up view to a multi-dimensional framing approach that includes both close-up speaker views and wider context views showing nearby participants. This dimensional expansion allows far-end participants to see both the speaker's expressions and the reactions of others simultaneously, resolving the contradiction between focused speaker tracking and contextual understanding.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

2Device complexity

If the system groups people sitting close together as a single entity, then the device complexity is reduced, but the viewing experience deteriorates due to distracting view switches

Engineering Contradiction:
Improveframing control complexityVSAvoidviewing experience
Core Design Contradiction:
Device complexityVSEase of operation

Solution Approach 1:

The system applies different framing qualities to different spatial zones within the conference room. Instead of treating all nearby participants as a single homogeneous group, it identifies individual participants or small sub-groups based on their specific positions and engagement levels. This local differentiation allows the system to maintain stable, targeted framing on relevant participants without causing distracting view switches, while keeping complexity manageable through automated spatial analysis.

Inventive Principle:
Principle #3Local quality

Data Source

PatentUS10708544B2Group and conversational framing for speaker tracking in a video conference system
Publication Date: 2020.07.07 CISCO TECHNOLOGY INC
  • US10708544B2 patent drawing
  • US10708544B2 patent drawing
  • US10708544B2 patent drawing

AI summary

In one embodiment, a method is provided to intelligently frame groups of participants in a meeting. This gives a more pleasing experience with fewer switches, better contextual understanding, and more natural framing, as would be seen in a video production made by a human director. Furthermore, in accordance with another embodiment, conversational framing techniques are provided. During speaker tracking, when two local participants are addressing each other, a method is provided to show a close-up framing showing both participants. By evaluating the direction participants are looking and a speaker history, it is determined if there is a local discussion going on, and an appropriate framing is selected to give far-end participants the most contextually rich experience.