Conference Gallery View Intelligence System for Multi-Region Focus

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conferencing software typically uses a single camera view to capture all participants in a conference room, limiting individual focus and contributing to missed or misattributed conversations among participants both inside and outside the room.

Innovation Solution

Implementing a conference gallery view intelligence system that determines regions of interest within the conference room based on input video and audio streams, producing multiple output video streams for separate views in conferencing software to focus on specific participants based on criteria like presence and speaking time.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If a single camera view is used to capture all participants in a conference room, then all participants can be seen in one view, but individual focus is limited and conversations may be missed or misattributed

Engineering Contradiction:
Improveconversation detection accuracyVSAvoidvideo stream processing complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent divides the conference room into multiple regions of interest (ROIs) based on participant locations and conversation dynamics. Instead of treating the entire room as a single view, the system segments the video feed into multiple focused views, each capturing specific participants or conversation zones. This segmentation enables the system to track and attribute conversations more accurately while managing complexity through intelligent region definition.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent applies local quality by providing different video stream characteristics to different regions. Each region of interest receives customized video processing that focuses on specific participants or conversation areas, allowing for enhanced individual focus and conversation detection accuracy in each local area while maintaining overall system coherence.

Inventive Principle:
Principle #3Local quality

2Ease of operation

If multiple video streams are generated for different regions of interest, then individual participant focus is improved, but the system complexity increases

Engineering Contradiction:
Improveparticipant focus capabilityVSAvoidvideo stream generation complexity
Core Design Contradiction:
Ease of operationVSDevice complexity

Solution Approach 1:

The patent implements preliminary action by pre-defining regions of interest and establishing the framework for multiple video stream generation before the actual conference takes place. The system pre-processes video feeds to identify potential conversation zones and participant areas, creating a structured approach that simplifies real-time processing and reduces operational complexity during live events.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent applies dynamics by making the video stream generation adaptive and flexible. The system dynamically adjusts region definitions and video stream characteristics based on real-time conversation detection, participant movement, and engagement levels. This dynamic approach allows the system to maintain high participant focus capability while managing complexity through intelligent adaptation rather than rigid pre-programming.

Inventive Principle:
Principle #15Dynamics

3Reliability

If manual passing of information between communication services is required, then service integration is maintained, but user workload increases

Engineering Contradiction:
Improveservice integration reliabilityVSAvoidinformation transfer time
Core Design Contradiction:
ReliabilityVSLoss of time

Solution Approach 1:

The patent merges multiple communication services into a unified conference experience. By integrating video streaming, conversation detection, and information routing into a single coordinated system, the patent eliminates the need for manual information passing between separate services. The unified approach maintains service integration reliability while significantly reducing the time and effort required for information transfer.

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The patent introduces an intermediary system that automatically handles information routing between different communication services. This intermediary layer processes and forwards information between video streams, audio feeds, and communication platforms without requiring manual intervention. The intermediary maintains reliable service integration while eliminating the time-consuming manual passing of information.

Inventive Principle:
Principle #24Intermediary (Mediator)

Data Source

PatentUS20240364549A1Conference Gallery View
Publication Date: 2024.10.31 ZOOM COMMUNICATIONS INC
  • US20240364549A1 patent drawing
  • US20240364549A1 patent drawing
  • US20240364549A1 patent drawing

AI summary

A conference gallery view intelligence system determines at least two regions of interest within a conference room based on an input video stream received from a video capture device located within the conference room. An output video stream for rendering within conferencing software is produced for each of the at least two regions of interest. The output video stream for each of the at least two regions of interest is then transmitted to one or more client devices connected to the conferencing software.