Auto-calibrating Speaker Tracking Systems for Conference Rooms

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing speaker tracking systems in video conference endpoints often require manual calibration and struggle to accurately frame active speakers when multiple systems are deployed at different angles, leading to suboptimal far-end experiences due to the lack of automatic spatial calibration.

Innovation Solution

The system automatically calibrates multiple speaker tracking systems by collecting data points from active speakers using cameras and microphone arrays, determining a reference coordinate system, and calculating the spatial locations of secondary systems relative to a master system, allowing for dynamic repositioning and improved framing without manual intervention.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If multiple speaker tracking systems are deployed at different angles to improve framing, then the far-end experience is improved, but manual calibration complexity increases

Engineering Contradiction:
Improveframing accuracyVSAvoidcalibration complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The speaker tracking systems automatically calibrate themselves by detecting common active speakers and computing relative positions without manual intervention. The system uses data from multiple tracking systems to determine a reference coordinate system and calculate spatial relationships autonomously, eliminating the need for manual calibration while maintaining accurate framing across multiple angled deployments.

Inventive Principle:
Principle #25Self-service

2Measurement precision

If multiple speaker tracking systems share and combine data to improve active speaker detection, then framing quality improves, but system complexity increases

Engineering Contradiction:
Improveactive speaker detection accuracyVSAvoiddata processing complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The system merges data from multiple speaker tracking systems by collecting data points from each system and combining them to determine a unified reference coordinate system. This integration allows the systems to share information about active speakers and compute relative positions, improving detection accuracy while managing complexity through systematic data fusion rather than ad-hoc processing.

Inventive Principle:
Principle #5Merging (Combining)

3Measurement precision

If manual calibration is performed to ensure accurate spatial positioning, then measurement precision is maintained, but time consumption increases

Engineering Contradiction:
Improvespatial positioning accuracyVSAvoidcalibration time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The system performs preliminary automatic calibration by detecting common active speakers and computing relative positions before actual video conferencing begins. This preliminary action establishes the reference coordinate system and spatial relationships in advance, eliminating the need for time-consuming manual calibration while ensuring measurement precision is maintained from the start of operations.

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS9986360B1Auto-calibration of relative positions of multiple speaker tracking systems
Publication Date: 2018.05.29 CISCO TECHNOLOGY INC
  • US9986360B1 patent drawing
  • US9986360B1 patent drawing
  • US9986360B1 patent drawing

AI summary

A system that automatically calibrates multiple speaker tracking systems with respect to one another based on detection of an active speaker at a collaboration endpoint is presented herein. The system collects a first data point set of an active speaker at the collaboration endpoint using at least a first camera and a first microphone array. The system then receives a plurality of second data point sets from one or more secondary speaker tracking systems located at the collaboration endpoint. Once enough data points have been collected, a reference coordinate system is determined using the first data point set and the one or more second data point sets. Finally, after a reference coordinate system has been determined, the system generates the locations of the one or more secondary speaker tracking systems with respect to the first speaker tracking system.