Audio-Video Talker Localization Using Motion Detection
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing videoconferencing solutions for localizing an active talker often rely solely on audio information, leading to reduced accuracy and increased hardware and computational requirements, especially when the speaker is facing away from the device.
Innovation Solution
A method that combines audio and motion information analysis, using a unique algorithm to weight lower frequencies and detect motion at candidate angles to accurately identify the active talker, allowing for high-definition display with fewer cameras and less computational resources.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If only audio information is used to localize an active talker, then hardware requirements are reduced, but localization accuracy decreases and the system becomes more cumbersome
Solution Approach 1:
The patent combines audio information processing with motion detection from video feed to localize the active talker. The audio module identifies candidate talker positions while the video module detects motion, and their results are integrated to determine the final talker location, achieving high accuracy without requiring multiple cameras
Solution Approach 2:
The video feed serves multiple functions: it provides motion detection data for talker localization and simultaneously serves as the display output for the conference participants. This multi-functionality reduces hardware requirements while maintaining localization accuracy
2Measurement precision
If multiple cameras are used to improve talker localization accuracy, then measurement precision increases, but device complexity and cost increase
Solution Approach 1:
The system merges audio-based candidate position identification with video-based motion detection to achieve accurate talker localization using only a single camera. The audio module provides directional information while the video module confirms motion at the predicted location, eliminating the need for multiple cameras
Solution Approach 2:
The audio processing module acts as an intermediary that predicts talker position from audio signals, allowing the single camera to focus on detecting motion at the predicted location rather than scanning the entire field of view. This intermediary step enables accurate localization with minimal video processing
3Reliability
If audio-only processing is used, then computational resources are reduced, but localization accuracy and reliability decrease
Solution Approach 1:
The audio processing module performs preliminary action by identifying candidate talker positions and time intervals before the video motion detection is executed. This preliminary filtering allows the system to focus computational resources on verifying motion only at the most likely candidate positions, improving reliability while controlling computational load
Solution Approach 2:
The system applies partial action by using audio processing to identify candidate positions and then applying video motion detection only to those specific candidates rather than analyzing the entire video feed. This selective approach achieves reliable localization with reduced computational requirements
Data Source
AI summary
A videoconferencing endpoint includes at least one processor a number of microphones and at least one camera. The endpoint can receive audio information and visual motion information during a teleconferencing session. The audio information includes one or more angles with respect to the microphone from a location of a teleconferencing session. The system evaluates the audio information is evaluated to determine at least one candidate angle corresponding to a possible location of an active talker. The candidate angle can be analyzed further with respect to the motion information to determine whether the candidate angle correctly corresponds to person who is speaking during the teleconferencing session. The person's face can then be framed within a frame view.


