Speaker Framing via Face-Guided Sound Source Localization
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Videoconferencing systems face challenges in accurately focusing on speakers due to jitter in sound source localization (SSL) pan angle data caused by conference room acoustics and varying distances between speakers and microphone arrays, leading to unreliable framing.
Innovation Solution
Combining SSL pan angle data with image data processed by a trained machine learning system to identify facial features, using bins to correlate pixel ranges with SSL pan angles, and determining the final SSL pan angle by tallying bin entries for bounding boxes to improve the accuracy of speaker localization and framing.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Ease of operation
If sound source localization (SSL) is used to determine speaker direction, then camera focusing on speaker is enabled, but jitter in SSL pan angle data occurs due to conference room acoustics and distance variations
Solution Approach 1:
The patent combines SSL pan angle data with image data from face detection algorithms to determine the final camera pan angle. By merging audio-based SSL results with visual face detection results, the system compensates for jitter in SSL data through the more stable image-based face location information, thereby improving reliability while maintaining ease of operation
Solution Approach 2:
The patent introduces an intermediary processing step that correlates SSL pan angles with image data through a trained machine learning system. This intermediary system processes both audio and visual inputs separately and then integrates them to produce a stabilized final pan angle, acting as a mediator that filters out jitter from the raw SSL data
2Productivity
If SSL pan angle data is used directly for framing, then speaker localization is achieved, but framing accuracy deteriorates due to jitter from room acoustics and speaker-microphone distance variations
Solution Approach 1:
The patent merges SSL pan angle data with face detection image data to determine the final camera framing angle. By combining these two data sources through correlation and machine learning processing, the system maintains the speed advantage of SSL while improving framing accuracy through the complementary stability of visual face detection
Solution Approach 2:
The system implements feedback by continuously comparing SSL pan angle estimates with face detection results and using this comparison to adjust and stabilize the final framing angle. The machine learning system learns from the correlation between audio and visual data to provide more accurate and stable framing decisions
Data Source
AI summary
A videoconferencing system includes a camera acquiring image data and a microphone array acquiring audio data. Image data is used in conjunction with sound source localization (SSL) data to locate a talker depicted in the image data. SSL processes the audio data and determines SSL pan angle values indicative of an estimated direction of a sound. Columns of pixels in an image are associated with bins. A bin count is incremented for each SSL pan angle value of the audio data that falls within a given bin. A bounding box in the image data is determined that encompasses a face depicted in the image data. A range of pixels is determined for the bounding box, such as extending from a leftmost column to a rightmost column. The bin with the highest bin count that also overlaps a range of pixels for a bounding box is deemed to contain the talker.


