Audio-Driven Camera FOV Switching in Video Conference Endpoints

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Manual selection of cameras in video conference systems to capture talking participants is cumbersome and inconvenient, as it requires operators to manually switch between different cameras positioned in various areas of a room.

Innovation Solution

Implementing an automatic switching mechanism between camera field-of-views (FOVs) based on audio signals, using a microphone array to define audio search regions and determine if audio originates from these regions, thereby activating the appropriate camera FOV to capture video of the speaking participant.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Ease of operation

If manual camera selection is used, then operators can control which camera captures video, but the operation becomes cumbersome and inconvenient

Engineering Contradiction:
Improvecamera switching operationVSAvoidtime for manual camera selection
Core Design Contradiction:
Ease of operationVSLoss of time

Solution Approach 1:

The system automatically switches between cameras based on audio detection without requiring operator intervention. The endpoint device autonomously determines which participant is speaking and switches the active camera accordingly, making the system self-sufficient in performing the camera switching function.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The patent replaces the manual mechanical operation of switching cameras with an automated audio-based detection and control system. The system uses audio sensors to detect speaking participants and automatically triggers camera switching, eliminating the need for manual mechanical intervention.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

2Adaptability or versatility

If multiple cameras are positioned to capture different areas, then comprehensive video coverage is achieved, but manual switching between them becomes complex

Engineering Contradiction:
Improvevideo coverage capabilityVSAvoidcamera switching control
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The system continuously monitors audio signals from multiple microphones to detect which participant is currently speaking. This real-time audio feedback drives the automatic camera switching decision, creating a closed-loop control system that adapts to changing speaking participants without increasing operational complexity.

Inventive Principle:
Principle #23Feedback

Solution Approach 2:

The endpoint device integrates multiple functions including audio detection, speaker identification, and camera control into a single unified system. This multi-functional approach allows the system to handle comprehensive video coverage across multiple cameras while managing the switching logic centrally, reducing overall system complexity.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Applied Scientific Principles

This section explains which scientific principles are used to turn an abstract innovation direction into a practical engineering solution.

Function Achieved in This Case

Automatically switches between main and side camera FOVs to capture video of actively speaking participants, eliminating the need for manual camera control and enhancing the video conference experience by dynamically framing speakers and audience.

Implementation Method 1

defining main and side audio search regions angularly-separated from each other at a microphone array configured to transduce audio received from the audio search regions

Methodology Applied
Scientific EffectTransduction:

Data Source

PatentUS9693017B2Automatic switching between different cameras at a video conference endpoint based on audio
Publication Date: 2017.06.27 CISCO TECHNOLOGY INC
  • US9693017B2 patent drawing
  • US9693017B2 patent drawing
  • US9693017B2 patent drawing

AI summary

A video conference endpoint includes predefined main and side audio search regions angularly-separated from each other at a microphone array configured to transduce audio received from the search regions. The endpoint includes one or more cameras to capture video in a main field of view (FOV) that encompasses the main audio search region. The endpoint determines if audio originates from any of the main and side audio search regions based on the transduced audio and predetermined audio search criteria. If it is determined that audio originates from the side audio search region, the endpoint automatically switches from capturing video in the main FOV to one or more cameras to capture video in a side FOV that encompasses the side audio search region.