Group and Conversational Framing for Speaker Tracking
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current video conference systems focus on framing the active speaker, which limits the contextual understanding for far-end participants by not showing reactions and body language of other participants, and often group people sitting close together as a single entity, leading to distracting view switches.
Innovation Solution
Implementing group and conversational framing techniques that detect participant proximity to group participants and adjust camera views to frame the active speaker with nearby participants, reducing unnecessary view switches and enhancing contextual understanding.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If the system frames only the active speaker in close-up views, then the speaker tracking function is improved, but the contextual understanding for far-end participants deteriorates because reactions and body language of other participants are not shown
Solution Approach 1:
The system segments the video feed into multiple regions of interest, identifying both the active speaker and nearby participants as separate focal points. This allows the system to maintain close-up framing of the speaker while also capturing and transmitting reactions from other participants, thereby preserving contextual information without sacrificing speaker tracking accuracy.
Solution Approach 2:
The system transitions from a single-dimension close-up view to a multi-dimensional framing approach that includes both close-up speaker views and wider context views showing nearby participants. This dimensional expansion allows far-end participants to see both the speaker's expressions and the reactions of others simultaneously, resolving the contradiction between focused speaker tracking and contextual understanding.
2Device complexity
If the system groups people sitting close together as a single entity, then the device complexity is reduced, but the viewing experience deteriorates due to distracting view switches
Solution Approach 1:
The system applies different framing qualities to different spatial zones within the conference room. Instead of treating all nearby participants as a single homogeneous group, it identifies individual participants or small sub-groups based on their specific positions and engagement levels. This local differentiation allows the system to maintain stable, targeted framing on relevant participants without causing distracting view switches, while keeping complexity manageable through automated spatial analysis.
Data Source
AI summary
In one embodiment, a method is provided to intelligently frame groups of participants in a meeting. This gives a more pleasing experience with fewer switches, better contextual understanding, and more natural framing, as would be seen in a video production made by a human director. Furthermore, in accordance with another embodiment, conversational framing techniques are provided. During speaker tracking, when two local participants are addressing each other, a method is provided to show a close-up framing showing both participants. By evaluating the direction participants are looking and a speaker history, it is determined if there is a local discussion going on, and an appropriate framing is selected to give far-end participants the most contextually rich experience.


