Virtual Microphone Mapping for Noise-Resistant Speech Targeting
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current automated conference room camera systems struggle with accurately identifying valid speech sources due to reliance on visual cues and frequency-based filtering, leading to false triggers from noise sources like keyboard typing and crinkling bags, and are constrained by zone-based designs that require talkers to remain in specific zones, limiting flexibility and causing potential pickup discontinuities.
Innovation Solution
A virtual microphone map processing system combined with a machine learning-based speech/noise discriminator to accurately track acoustic energy locations and determine valid speech sources, allowing seamless camera switching and continuous tracking of desired speech events while ignoring noise.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If frequency-based filtering and energy-based detection are used to distinguish speech from noise, then the system can reduce false triggers from noise sources, but it cannot reliably distinguish between speech and noise sources that have similar frequency content and energy envelopes
Solution Approach 1:
The patent replaces traditional frequency-based filtering and energy-based detection mechanisms with a machine learning-based speech/noise discriminator. This substitution allows the system to accurately distinguish between speech and noise sources that have similar frequency content and energy envelopes, thereby improving speech detection accuracy without being constrained by the limitations of conventional filtering methods.
Solution Approach 2:
The patent transforms the approach by changing from static frequency and energy parameters to dynamic, context-aware features extracted by machine learning models. The system analyzes multiple parameters including spectral characteristics, temporal patterns, and contextual information to make intelligent distinctions between speech and noise, enabling reliable detection even when traditional parameters are insufficient.
2Stability of the object's composition
If zone-based camera switching is implemented, then the system can maintain structured camera coverage, but it causes pickup discontinuities when talkers move between zones and requires talkers to remain in specific zones
Solution Approach 1:
The patent implements dynamic camera switching based on real-time speech detection and talker location tracking. Instead of rigid zone-based switching, the system continuously adjusts camera views according to the current speaker's position and speech activity, enabling seamless camera coverage as talkers move freely throughout the room while maintaining stable and adaptive video feed.
3Extent of automation
If visual cues and speech energy signatures are used for talker identification, then the system can automatically switch cameras, but it cannot reliably identify talkers when visual cues are unclear or when noise sources are present
Solution Approach 1:
The patent replaces traditional visual cue-based and energy-based talker identification with machine learning-based speech detection and classification. This substitution enables the system to automatically and precisely identify talkers by analyzing acoustic characteristics, spectral features, and contextual information, thereby improving identification precision while maintaining full automation of camera switching.
Solution Approach 2:
The patent introduces machine learning models as intermediary components between the microphone array and camera switching system. These models process acoustic signals, extract meaningful features, and provide intelligent decisions about talker identification and camera selection, serving as a bridge that enhances both automation and precision.
Data Source
AI summary
The artificial intelligence (AI-VAD) engine utilizes virtual microphone (bubble) map processing for continuous acoustic energy location tracking throughout a room. A virtual microphone can be placed anywhere in the room and tracked based on acoustic energy location identification. A noise discriminator determines whether the acoustic energy is speech or a non-speech noise source. This allows talker location(s) to be accurately tracked in real-time while non-speech noise sources are ignored. The tracked talker location(s) can be used to accurately position cameras on the speaker for use with unified communication clients.


