Video Conference Camera View Selection via Audio-Visual Processing
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Video conferencing systems often struggle with optimal camera framing, leading to poor video quality and participant visibility due to manual adjustments and limitations in voice-tracking technology, especially when multiple cameras are used in daisy-chained configurations.
Innovation Solution
A video conferencing endpoint that automatically adjusts one or more cameras to provide an optimal view of all participants and the speaker, using a combination of audio and video processing to determine the best frontal view, incorporating steerable Pan-Tilt-Zoom cameras and Electronic Pan-Tilt-Zoom cameras to switch between wide and zoomed-in views dynamically.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Manufacturing precision
If manual camera adjustment is used, then video framing can be optimized, but operation complexity and time consumption increase
Solution Approach 1:
The system performs automatic camera framing by detecting participant positions and speaker locations, then autonomously adjusting pan-tilt-zoom camera parameters without requiring manual intervention. The camera controller receives commands based on detected spatial information and automatically positions cameras to capture optimal views.
Solution Approach 2:
Manual mechanical camera adjustment is replaced with an automated control system that uses audio-visual processing to determine camera positioning. The system substitutes human operation with electronic control based on detected participant locations and speaker identification.
2Extent of automation
If voice-tracking technology is used, then automatic speaker tracking is improved, but accuracy deteriorates in reverberant environments or when speakers turn away
Solution Approach 1:
The system combines voice-tracking technology with visual detection methods to determine speaker location. By merging audio-based speaker identification with video-based participant position detection, the system achieves more accurate and reliable speaker tracking that works effectively in reverberant environments and when speakers turn away from microphones.
3Area of stationary object
If multiple cameras are daisy-chained to extend range, then coverage area increases, but view selection complexity increases
Solution Approach 1:
The system uses feedback from audio-visual processing to automatically determine which camera provides the optimal view. The camera controller receives information about participant positions and speaker locations, then selects and adjusts the appropriate camera among daisy-chained devices to capture the best view of the current speaker or group.
4Ease of operation
If fixed wide-angle camera view is used, then camera setup simplicity is maintained, but participant visibility and engagement deteriorate
Solution Approach 1:
The system transitions from a fixed static camera view to a dynamic adjustable view. Pan-tilt-zoom cameras are controlled to automatically adjust their positioning and framing based on detected participant locations and speaker identification, enabling the camera to dynamically optimize video quality without requiring manual setup adjustments.
Data Source
Figure 1A~1B
Figure 1C~2D
Figure 3~4
AI summary
A system for ensuring that the best available view of a person's face is included in a video stream when the person's face is being captured by multiple cameras (50A, 50B) at multiple angles at a first endpoint (10). The system uses one or more microphone arrays (60A-B) to capture direct-reverberant ratio information corresponding to the views, and determines which view most closely matches a view of the person looking directly at the camera (50A, 50B), thereby improving the experience for viewers at a second endpoint (14).