Videoconferencing Endpoint Dual Camera Speaker Tracking

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Videoconferences often suffer from poor framing, where far-end participants struggle to see facial expressions and determine who is speaking due to small participant sizes on screen, and voice-tracking cameras face challenges in accurately locating speakers, especially in reverberant environments or when speakers turn away.

Innovation Solution

Implementing a system with at least two cameras at an endpoint that captures video in a controlled manner, using techniques such as motion detection, skin tone detection, and facial recognition to dynamically adjust the view based on who is speaking, and incorporating speech recognition to assist in camera control, allowing for automated switching between wide and tight views.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Area of stationary object

If a single camera captures a wide view to include all participants, then all participants are visible in the frame, but the participant size on screen becomes too small and facial expressions cannot be seen

Engineering Contradiction:
Improvefield of viewVSAvoidfacial expression detail
Core Design Contradiction:
Area of stationary objectVSLoss of information

Solution Approach 1:

The patent divides the video capture function into two separate cameras: a wide-area camera that captures all participants and a close-up camera that captures detailed facial expressions. This segmentation allows the system to provide both wide coverage and detailed views simultaneously, resolving the contradiction between field of view and detail visibility.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent adds a temporal dimension by dynamically switching between wide and close-up views based on speaker detection. When a participant speaks, the system transitions from the wide view to the close-up view of that participant, and switches back when they stop speaking. This temporal switching allows the system to provide appropriate view levels at different times.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

2Loss of information

If manual camera adjustment is used to frame participants better, then viewing quality improves, but operation becomes cumbersome and requires repeated adjustments when participants change positions

Engineering Contradiction:
Improvevideo framing qualityVSAvoidcamera adjustment complexity
Core Design Contradiction:
Loss of informationVSEase of operation

Solution Approach 1:

The patent implements automated camera control that detects speaker positions and automatically switches between cameras without user intervention. The system monitors audio signals to identify active speakers and autonomously determines when to switch from wide to close-up views and which camera to use, eliminating the need for manual operation while maintaining optimal framing quality.

Inventive Principle:
Principle #25Self-service

3Extent of automation

If voice-tracking camera is used to automatically direct toward speakers, then camera control becomes automated, but the camera may lose track of speakers when they turn away or point at reflection points in reverberant environments

Engineering Contradiction:
Improvecamera tracking automationVSAvoidspeaker detection accuracy
Core Design Contradiction:
Extent of automationVSReliability

Solution Approach 1:

The patent introduces video processing as an intermediary layer between audio detection and camera control. The system uses video data to detect participant positions and orientations, and uses this information to verify and correct audio-based speaker detection. This intermediary video verification prevents the camera from tracking reflection points or losing track of speakers who turn away, as the video data provides ground truth about actual speaker locations.

Inventive Principle:
Principle #24Intermediary (Mediator)

Data Source

PatentUS9392221B2Videoconferencing endpoint having multiple voice-tracking cameras
Publication Date: 2016.07.12 HEWLETT PACKARD DEVELOPMENT COMPANY LP
  • US9392221B2 patent drawing
  • US9392221B2 patent drawing
  • US9392221B2 patent drawing

AI summary

A videoconferencing apparatus automatically tracks speakers in a room and dynamically switches between a controlled, people-view camera and a fixed, room-view camera. When no one is speaking, the apparatus shows the room view to the far-end. When there is a dominant speaker in the room, the apparatus directs the people-view camera at the dominant speaker and switches from the room-view camera to the people-view camera. When there is a new speaker in the room, the apparatus switches to the room-view camera first, directs the people-view camera at the new speaker, and then switches to the people-view camera directed at the new speaker. When there are two near-end speakers engaged in a conversation, the apparatus tracks and zooms-in the people-view camera so that both speakers are in view.