Microphone Mute Notification via Acoustic Source Localization
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current video conferencing systems face issues with false or missed mute/unmute notifications due to interference from background noise, echo, and low signal-to-noise ratios, leading to degraded user experience.
Innovation Solution
A processor circuitry that receives audio and video signals from multiple microphones and cameras, detects silence and voice events, and classifies them as interference or speaker events using acoustic source localization, echo cancellation, and face detection, generating appropriate mute or unmute indications and notifications.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If traditional audio signal analysis is used to detect speaker activity, then the system can identify when a participant is speaking, but it produces false alarms due to background noise, echo, and low signal-to-noise ratios
Solution Approach 1:
The patent segments the audio signal processing into multiple independent analysis streams: acoustic echo cancellation, noise reduction, acoustic source localization, and voice activity detection. Each stream processes specific aspects of the audio signal separately, then their results are integrated to make the final speaker detection decision, reducing false alarms from any single analysis method
Solution Approach 2:
The patent introduces an intermediary classification mechanism that acts as a mediator between raw audio detection and final speaker identification. The system classifies detected audio events as either 'speaker events' or 'interference events' by analyzing multiple parameters simultaneously, filtering out false detections from background noise and echo before generating mute/unmute notifications
2Measurement precision
If multiple microphones and complex signal processing algorithms are used to improve speaker detection accuracy, then false alarms are reduced, but the computational complexity and processing requirements increase
Solution Approach 1:
The patent divides the complex signal processing task into separate functional modules: acoustic echo cancellation, noise reduction, acoustic source localization, and voice activity detection. Each module handles a specific aspect independently, allowing for optimized processing of each function rather than one monolithic complex algorithm
Solution Approach 2:
The patent applies partial processing by focusing computational resources on the most critical discrimination tasks. Rather than analyzing every parameter in full detail simultaneously, the system performs targeted analysis on key discriminators (acoustic location, echo cancellation, noise levels) to achieve sufficient accuracy without excessive computational overhead
3Measurement precision
If the system uses acoustic source localization and face detection to distinguish speakers from background noise, then notification accuracy is improved, but the processing time and computational load increase
Solution Approach 1:
The patent performs preliminary acoustic source localization and echo cancellation processing on incoming audio signals before voice activity detection is finalized. By pre-processing the audio data to establish baseline acoustic characteristics and cancel echo patterns in advance, the system reduces the computational burden during real-time speaker detection, minimizing processing delays
Solution Approach 2:
The patent merges multiple detection functions (acoustic source localization, face detection, voice activity detection) into a unified speaker verification system. By combining these functions and sharing computational resources and data across them, the system achieves high accuracy without the cumulative time penalty of completely separate processing pipelines
Data Source
Figure 1
Figure 2
Figure 3
AI summary
A processing system can include a processor that includes circuitry. The circuitry can be configured to: receive far-end and near-end audio signals; detect silence events and voice activities from the audio signals; determine whether an audio event in the audio signals is an interference event or a speaker event based on the detected silence events and voice activities, and further based on localized acoustic source data and faces or motion detected from an image; and generate a mute or unmute indication based on whether the audio event is the interference event or the speaker event. The system can include a near-end microphone array to output the near-end audio signals, one or more far-end microphones to output the far-end audio signals, and one or more cameras to capture the image of the environment.