Video Conference Camera View Selection via Audio-Visual Processing

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Video conferencing systems often struggle with optimal camera framing, leading to poor video quality and participant visibility due to manual adjustments and limitations in voice-tracking technology, especially when multiple cameras are used in daisy-chained configurations.

Innovation Solution

A video conferencing endpoint that automatically adjusts one or more cameras to provide an optimal view of all participants and the speaker, using a combination of audio and video processing to determine the best frontal view, incorporating steerable Pan-Tilt-Zoom cameras and Electronic Pan-Tilt-Zoom cameras to switch between wide and zoomed-in views dynamically.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Manufacturing precision

If manual camera adjustment is used, then video framing can be optimized, but operation complexity and time consumption increase

Engineering Contradiction:
Improvevideo framing qualityVSAvoidcamera adjustment operation
Core Design Contradiction:
Manufacturing precisionVSEase of operation

Solution Approach 1:

The system performs automatic camera framing by detecting participant positions and speaker locations, then autonomously adjusting pan-tilt-zoom camera parameters without requiring manual intervention. The camera controller receives commands based on detected spatial information and automatically positions cameras to capture optimal views.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

Manual mechanical camera adjustment is replaced with an automated control system that uses audio-visual processing to determine camera positioning. The system substitutes human operation with electronic control based on detected participant locations and speaker identification.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

2Extent of automation

If voice-tracking technology is used, then automatic speaker tracking is improved, but accuracy deteriorates in reverberant environments or when speakers turn away

Engineering Contradiction:
Improveautomatic speaker trackingVSAvoidspeaker location accuracy
Core Design Contradiction:
Extent of automationVSMeasurement precision

Solution Approach 1:

The system combines voice-tracking technology with visual detection methods to determine speaker location. By merging audio-based speaker identification with video-based participant position detection, the system achieves more accurate and reliable speaker tracking that works effectively in reverberant environments and when speakers turn away from microphones.

Inventive Principle:
Principle #5Merging (Combining)

3Area of stationary object

If multiple cameras are daisy-chained to extend range, then coverage area increases, but view selection complexity increases

Engineering Contradiction:
Improvecamera coverage areaVSAvoidcamera system configuration
Core Design Contradiction:
Area of stationary objectVSDevice complexity

Solution Approach 1:

The system uses feedback from audio-visual processing to automatically determine which camera provides the optimal view. The camera controller receives information about participant positions and speaker locations, then selects and adjusts the appropriate camera among daisy-chained devices to capture the best view of the current speaker or group.

Inventive Principle:
Principle #23Feedback

4Ease of operation

If fixed wide-angle camera view is used, then camera setup simplicity is maintained, but participant visibility and engagement deteriorate

Engineering Contradiction:
Improvecamera setup simplicityVSAvoidvideo quality and participant visibility
Core Design Contradiction:
Ease of operationVSManufacturing precision

Solution Approach 1:

The system transitions from a fixed static camera view to a dynamic adjustable view. Pan-tilt-zoom cameras are controlled to automatically adjust their positioning and framing based on detected participant locations and speaker identification, enabling the camera to dynamically optimize video quality without requiring manual setup adjustments.

Inventive Principle:
Principle #15Dynamics

Data Source

PatentEP3422705B1Optimal view selection method in a video conference
Publication Date: 2024.07.24 HEWLETT PACKARD DEVELOPMENT COMPANY LP
  • EP3422705B1 patent drawingFigure 1A~1B
  • EP3422705B1 patent drawingFigure 1C~2D
  • EP3422705B1 patent drawingFigure 3~4

AI summary

A system for ensuring that the best available view of a person's face is included in a video stream when the person's face is being captured by multiple cameras (50A, 50B) at multiple angles at a first endpoint (10). The system uses one or more microphone arrays (60A-B) to capture direct-reverberant ratio information corresponding to the views, and determines which view most closely matches a view of the person looking directly at the camera (50A, 50B), thereby improving the experience for viewers at a second endpoint (14).