Auto-Framing Video Conference Camera Tracking

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current video conferencing systems face challenges in accurately framing meeting participants due to difficulties in manually adjusting camera settings, with voice-tracking cameras often losing track of speakers or directing to reflections rather than actual sound sources, especially in reverberant environments.

Innovation Solution

The implementation of a video conferencing apparatus that uses a combination of stationary and adjustable cameras, employing fast and reliable people detection methods such as face, torso, and motion detection, along with audio localization, to automatically and continuously adjust the camera view to provide an optimal framing of all participants or the speaker, eliminating the need for manual initiation and improving accuracy even when participants turn away or change positions.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If voice-tracking cameras are used to direct toward speakers, then speaker tracking capability is improved, but the system may lose track of speakers or direct at reflections in reverberant environments

Engineering Contradiction:
Improvespeaker tracking accuracyVSAvoidtracking stability
Core Design Contradiction:
Measurement precisionVSReliability

Solution Approach 1:

The patent combines multiple detection methods (face detection, torso detection, motion detection) with audio localization to create a hybrid tracking system. This merging of visual and audio cues allows the system to maintain reliable speaker tracking even when audio alone is ambiguous due to reflections or when participants turn away from microphones.

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The system introduces visual detection results as an intermediary to verify and supplement audio-based speaker identification. When audio localization is uncertain (due to reverberation or speaker orientation), the visual detection of faces and torsos serves as a mediator to confirm the actual speaker location, preventing misdirection to reflection points.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Ease of operation

If manual camera adjustment is used, then framing control is achieved, but the operation becomes cumbersome and requires repeated adjustments

Engineering Contradiction:
Improvecamera control simplicityVSAvoidadjustment time
Core Design Contradiction:
Ease of operationVSLoss of time

Solution Approach 1:

The system implements automated framing that performs camera adjustments without user intervention. The detection and tracking modules continuously monitor participant positions and automatically control camera pan, tilt, and zoom to maintain optimal framing, eliminating the need for manual operations and repeated adjustments when participants move.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The system establishes a feedback loop where detection results from visual and audio sensors continuously inform camera control decisions. This real-time feedback mechanism allows the camera to automatically adapt to changing participant positions and speaker changes, providing responsive framing without manual input.

Inventive Principle:
Principle #23Feedback

3Extent of automation

If automated detection systems are implemented, then participant tracking is improved, but the system requires manual initiation and training mode switching

Engineering Contradiction:
Improvecamera control automationVSAvoidsystem initiation simplicity
Core Design Contradiction:
Extent of automationVSEase of operation

Solution Approach 1:

The system performs preliminary detection and tracking setup automatically upon initialization, eliminating the need for manual training mode activation. The detection modules begin operating immediately in the background, and the system transitions smoothly to active tracking without requiring user intervention to switch modes or initiate automated functions.

Inventive Principle:
Principle #10Preliminary action

4Area of stationary object

If wide-angle fixed camera view is used, then coverage of entire room is achieved, but participant size in display becomes too small

Engineering Contradiction:
Improvecamera field of viewVSAvoidparticipant visibility
Core Design Contradiction:
Area of stationary objectVSMeasurement precision

Solution Approach 1:

The system transforms the camera from a static wide-angle view to a dynamic view that automatically adjusts pan, tilt, and zoom based on detected participant positions and speaker location. This dynamic adjustment allows the camera to maintain appropriate framing and participant size in the display while still being able to cover the entire room when needed.

Inventive Principle:
Principle #15Dynamics

Data Source

PatentUS10574899B2People detection method for auto-framing and tracking in a video conference
Publication Date: 2020.02.25 HEWLETT PACKARD DEVELOPMENT COMPANY LP
  • US10574899B2 patent drawing
  • US10574899B2 patent drawing
  • US10574899B2 patent drawing

AI summary

A videoconference apparatus and method coordinates a stationary view obtained with a stationary camera to an adjustable view obtained with an adjustable camera. The stationary camera can be a web camera, while the adjustable camera can be a pan-tilt-zoom camera. As the stationary camera obtains video, participants are detected and localized by establishing a static perimeter around a participant in which no motion is detected. Thereafter, if no motion is detected in the perimeter, any personage objects such as head, face, or shoulders which are detected in the region bounded by the perimeter are determined to correspond to the participant.