Videoconferencing Endpoint Dual Camera Speaker Tracking
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Videoconferences often suffer from poor framing, where far-end participants struggle to see facial expressions and determine who is speaking due to small participant sizes on screen, and voice-tracking cameras face challenges in accurately locating speakers, especially in reverberant environments or when speakers turn away.
Innovation Solution
Implementing a system with at least two cameras at an endpoint that captures video in a controlled manner, using techniques such as motion detection, skin tone detection, and facial recognition to dynamically adjust the view based on who is speaking, and incorporating speech recognition to assist in camera control, allowing for automated switching between wide and tight views.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Area of stationary object
If a single camera captures a wide view to include all participants, then all participants are visible in the frame, but the participant size on screen becomes too small and facial expressions cannot be seen
Solution Approach 1:
The patent divides the video capture function into two separate cameras: a wide-area camera that captures all participants and a close-up camera that captures detailed facial expressions. This segmentation allows the system to provide both wide coverage and detailed views simultaneously, resolving the contradiction between field of view and detail visibility.
Solution Approach 2:
The patent adds a temporal dimension by dynamically switching between wide and close-up views based on speaker detection. When a participant speaks, the system transitions from the wide view to the close-up view of that participant, and switches back when they stop speaking. This temporal switching allows the system to provide appropriate view levels at different times.
2Loss of information
If manual camera adjustment is used to frame participants better, then viewing quality improves, but operation becomes cumbersome and requires repeated adjustments when participants change positions
Solution Approach 1:
The patent implements automated camera control that detects speaker positions and automatically switches between cameras without user intervention. The system monitors audio signals to identify active speakers and autonomously determines when to switch from wide to close-up views and which camera to use, eliminating the need for manual operation while maintaining optimal framing quality.
3Extent of automation
If voice-tracking camera is used to automatically direct toward speakers, then camera control becomes automated, but the camera may lose track of speakers when they turn away or point at reflection points in reverberant environments
Solution Approach 1:
The patent introduces video processing as an intermediary layer between audio detection and camera control. The system uses video data to detect participant positions and orientations, and uses this information to verify and correct audio-based speaker detection. This intermediary video verification prevents the camera from tracking reflection points or losing track of speakers who turn away, as the video data provides ground truth about actual speaker locations.
Data Source
AI summary
A videoconferencing apparatus automatically tracks speakers in a room and dynamically switches between a controlled, people-view camera and a fixed, room-view camera. When no one is speaking, the apparatus shows the room view to the far-end. When there is a dominant speaker in the room, the apparatus directs the people-view camera at the dominant speaker and switches from the room-view camera to the people-view camera. When there is a new speaker in the room, the apparatus switches to the room-view camera first, directs the people-view camera at the new speaker, and then switches to the people-view camera directed at the new speaker. When there are two near-end speakers engaged in a conversation, the apparatus tracks and zooms-in the people-view camera so that both speakers are in view.


