Web Conference Head Detection and Contextual Data Overlay

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Remote participants in web conferencing events often struggle to identify local participants without introductions, leading to a disadvantage in contributing effectively to the conference.

Innovation Solution

An event computer detects local participants' heads and assigns contextual data, such as names and job titles, which is superimposed on the video feed and sent to remote participants, allowing them to identify local participants through virtual frames and recognition mechanisms.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Loss of information

If remote participants rely on verbal introductions to identify local participants, then the system maintains simplicity without additional technology, but remote participants who join late or miss introductions cannot identify local participants, reducing their ability to contribute effectively

Engineering Contradiction:
Improveidentification informationVSAvoidparticipant contribution
Core Design Contradiction:
Loss of informationVSEase of operation

Solution Approach 1:

The system performs head detection and assigns contextual data (names, titles, avatars) to local participants in advance, before remote participants need to identify them. This preliminary tagging ensures that identification information is always available, regardless of when a remote participant joins the conference.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent introduces an intermediary identification system that acts as a mediator between local and remote participants. Virtual frames with contextual data serve as an intermediary layer that translates visual information about local participants into identifiable information for remote participants, eliminating the need for verbal introductions.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Loss of information

If the system displays virtual frames with contextual data for all local participants, then remote participants can easily identify everyone in the meeting room, but the video feed becomes cluttered and harder to view

Engineering Contradiction:
Improveidentification informationVSAvoidvideo display area
Core Design Contradiction:
Loss of informationVSArea of stationary object

Solution Approach 1:

The system applies different visual qualities to different regions of the video feed. Virtual frames are displayed with varying levels of prominence depending on the participant's relevance, position, or engagement level. This allows identification information to be presented without uniformly cluttering the entire video display area.

Inventive Principle:
Principle #3Local quality

Solution Approach 2:

Instead of displaying comprehensive contextual data for all participants simultaneously, the system selectively displays virtual frames for only some participants at any given time. This partial action approach provides sufficient identification information without overwhelming the video display area with excessive visual elements.

Inventive Principle:
Principle #16Partial or excessive action

Data Source

PatentUS9191616B2Local participant identification in a web conferencing system
Publication Date: 2015.11.17 MICROSOFT TECHNOLOGY LICENSING LLC
  • US9191616B2 patent drawing
  • US9191616B2 patent drawing
  • US9191616B2 patent drawing

AI summary

An event computer receives video in which one or more local participants of a conferencing event are viewable. The event computer receives head detection information of the local participants and assigns contextual data to the head detection information for each of the local participants for which head detection information is received. The event computer then sends the video, the head detection information, and the contextual data to one or more remote participant computer systems by which one or more remote participants can view the local participants and their corresponding contextual data within the video.