Multi-Camera Speaker Framing via Audio-Visual Correlation

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

In videoconferencing settings, speakers often do not face the camera when addressing others in the room, leading to side views being captured at the far end, and existing multi-camera configurations fail to provide satisfactory views when multiple individuals are present.

Innovation Solution

Implementing multiple cameras with microphone arrays for sound source localization, combined with neural network processing to identify and frame the speaker's face from various angles, ensuring the best facial view is selected and transmitted, with a default view provided if necessary.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If multiple cameras are deployed to capture speakers from different directions, then the view of the speaker is improved, but the system complexity increases

Engineering Contradiction:
Improvespeaker view qualityVSAvoidcamera system complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The system divides the camera selection task into independent modules: sound source localization module, face detection module, and camera selection module. Each module processes specific information (audio direction, facial features, view quality) separately before integrating results to select the optimal camera, reducing overall system complexity while maintaining high speaker view quality

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

A central processing unit acts as an intermediary that receives inputs from multiple cameras and microphones, processes the combined audio-video data, and selects the best camera view. This intermediary layer coordinates the complex multi-camera system by synthesizing information from various sources and making intelligent camera selection based on pre-computed features

Inventive Principle:
Principle #24Intermediary (Mediator)

2Ease of operation

If single camera is used, then the device complexity is low, but the speaker does not appear to be speaking to the far end when looking at others

Engineering Contradiction:
Improvespeaker natural behaviorVSAvoidspeaker view quality
Core Design Contradiction:
Ease of operationVSMeasurement precision

Solution Approach 1:

The system dynamically switches between different camera views based on real-time detection of speaker location and orientation. When a speaker turns to address another person, the system automatically transitions to capture that directional view, creating a dynamic adaptation to speaker behavior rather than relying on a fixed single camera position

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

The camera system serves multiple functions simultaneously: capturing speaker views, detecting speaker orientation through audio-visual correlation, identifying when speakers address different directions, and selecting appropriate views for transmission. This multi-functionality allows a single camera to effectively replace multiple fixed cameras by adapting to various speaking scenarios

Inventive Principle:
Principle #6Universality (Multi-functionality)

3Adaptability or versatility

If multiple cameras point in different directions to capture multiple speakers, then the adaptability improves, but it becomes difficult to identify which individual is the speaker

Engineering Contradiction:
Improvemulti-speaker coverageVSAvoidspeaker identification
Core Design Contradiction:
Adaptability or versatilityVSDifficulty of detecting and measuring

Solution Approach 1:

The system replaces mechanical camera switching with intelligent video processing. Instead of physically rotating cameras or manually switching feeds, the system uses audio-visual correlation algorithms and machine learning models to automatically identify speakers and select appropriate camera views, substituting mechanical operations with computational intelligence

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Solution Approach 2:

The system continuously monitors audio sources and video feeds to detect speaker changes. When a new speaker is detected or an existing speaker changes orientation, the system receives feedback about the current speaking state and adjusts camera selection accordingly, creating a closed-loop system that maintains accurate speaker identification throughout the conference

Inventive Principle:
Principle #23Feedback

Data Source

PatentUS11606510B2Intelligent multi-camera switching with machine learning
Publication Date: 2023.03.14 HEWLETT PACKARD DEVELOPMENT COMPANY LP
  • US11606510B2 patent drawing
  • US11606510B2 patent drawing
  • US11606510B2 patent drawing

AI summary

Multiple cameras in a conference room, each pointed in a different direction and including a microphone array to perform sound source localization (SSL). The SSL is used in combination with the video image to identify the speaker from among multiple individuals that appear in the video image. Neural network or machine learning processing is performed on the identified speaker to determine the quality of the front or facial view of the speaker. The best view of the speaker's face from the various cameras is selected to be provided to the far end. If no view is satisfactory, a default view is selected and that is provided to the far end. The use of the SSL allows selection of the proper individual from a group of individuals in the conference room, so that only the speaker's head is analyzed for the best facial view and then framed for transmission.