Video Conference Face Detection Filtering Glass Windows

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Video conference systems incorrectly include individuals detected behind glass windows in the video layout due to inadequate head detection filtering, causing confusion for participants.

Innovation Solution

An endpoint device employs face detection and speaker tracking to create boundary data defining maximum distances and angles for valid face positions, rejecting faces detected outside these boundaries during a video conference.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Extent of automation

If automatic head detection is used to detect heads or faces of individuals, then the video framing can automatically include detected heads in the video layout, but individuals behind glass windows are incorrectly detected and included in the video layout

Engineering Contradiction:
Improveautomatic head detectionVSAvoidhead detection accuracy
Core Design Contradiction:
Extent of automationVSMeasurement precision

Solution Approach 1:

The patent introduces boundary data as an intermediary mechanism between the head detection system and the video framing system. This boundary data acts as a filter that mediates the inclusion of detected heads, allowing the system to automatically frame videos while preventing incorrect inclusions of individuals behind glass windows through spatial boundary constraints.

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The patent modifies the detection parameters by introducing boundary conditions (maximum distance and angle parameters) that constrain the detection results. By changing the parameter space to include spatial boundaries, the system can distinguish between valid in-room participants and invalid detections behind glass windows, improving measurement precision while maintaining automation.

Inventive Principle:
Principle #35Parameter changes

2Productivity

If face detection is used to detect faces in the video, then the system can identify participants, but it cannot distinguish between participants in the meeting room and individuals behind glass windows

Engineering Contradiction:
Improveparticipant identificationVSAvoidparticipant identification accuracy
Core Design Contradiction:
ProductivityVSReliability

Solution Approach 1:

The patent segments the detection space into valid in-room regions and invalid regions behind glass windows by introducing boundary data. This segmentation allows the system to identify all faces in the video (maintaining productivity) while filtering out faces that fall outside the defined spatial boundaries (improving reliability).

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent replaces purely mechanical vision-based detection with a hybrid system that incorporates acoustic information (speaker tracking) and spatial boundary constraints. This substitution enables the system to cross-validate face detections with acoustic data and spatial reasoning, improving identification accuracy without sacrificing productivity.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

3Area of stationary object

If the video layout includes all detected faces, then the video stream captures the entire meeting room view, but it creates confusion for participants at the far-end due to incorrect inclusions

Engineering Contradiction:
Improvevideo coverage areaVSAvoidconfusion caused by incorrect inclusions
Core Design Contradiction:
Area of stationary objectVSObject-affected harmful factors

Solution Approach 1:

The patent applies local quality by differentiating the treatment of detected faces based on their spatial location relative to the glass window boundaries. Faces within the valid boundary region are included in the video layout, while faces outside the boundary (behind glass windows) are excluded. This localized differentiation maintains comprehensive coverage of the meeting room while eliminating harmful inclusions that cause confusion.

Inventive Principle:
Principle #3Local quality

Data Source

PatentUS20240289984A1Method for rejecting head detections through windows in meeting rooms
Publication Date: 2024.08.29 CISCO TECHNOLOGY INC
  • US20240289984A1 patent drawing
  • US20240289984A1 patent drawing
  • US20240289984A1 patent drawing

AI summary

A method is performed by an endpoint device that includes a microphone array to detect audio and a camera to capture video. The method comprises: detecting faces in the video to produce detected faces; detecting talkers based on the audio to produce detected talkers; determining valid face positions based on the detected faces and the detected talkers; storing, as boundary data, face distances and face angles for the valid face positions as maximum distances for the face angles; detecting a face in the video to produce a detected face and a face position for the detected face; and including or not including the detected face in a video layout for transmission based on the face position and the boundary data.