Speaker Activity Detection Using Segmented Event Detectors
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
In speech communication systems, especially in automotive environments, existing technologies face challenges in accurately detecting speaker activity and selecting the optimal microphone, particularly in the presence of interfering acoustic events like wind noise and local distortions, leading to misdetection and poor speech signal processing.
Innovation Solution
The system employs energy-based speaker activity detection using power spectra and signal-to-noise ratio (SNR) analysis to differentiate between speaker activity and acoustic events, allowing for robust joint speaker activity and event detection, and dynamically selects the best microphone for processing based on SNR and event detection results.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Speed
If energy-based speaker activity detection is used to quickly identify speaking passengers, then response speed is improved, but misdetection occurs during interfering acoustic events like wind noise and local distortions
Solution Approach 1:
The detection system is segmented into multiple independent event detectors, each specialized for detecting specific types of acoustic events (wind noise, local distortions, double-talk). This allows the system to quickly identify different event types without requiring a single complex detector, maintaining fast response while improving accuracy through specialized detection for each event category.
Solution Approach 2:
The system changes detection parameters dynamically by switching between different detection modes and thresholds based on the current acoustic environment. When interfering events are detected, the system adjusts its detection parameters to distinguish between actual speaker activity and environmental disturbances, thereby maintaining high detection accuracy across varying conditions.
2Reliability
If the system switches to the next best microphone when distortion is detected, then robustness is improved, but false switching occurs when inactive microphones get distorted
Solution Approach 1:
The system performs preliminary detection of acoustic events before making microphone selection decisions. By detecting and identifying interfering events in advance, the system can distinguish between distortion caused by environmental factors and actual speaker activity, preventing false switching to alternative microphones when the current microphone is still providing valid speech signals.
Solution Approach 2:
The system uses feedback from multiple event detectors to continuously monitor microphone signal quality and adjust selection decisions. The feedback mechanism allows the system to learn from detection results and make more accurate microphone selection decisions, reducing false switching while maintaining robustness against actual signal degradation.
3Reliability
If multiple microphones are provided per speaker for redundancy, then robustness against distortion is improved, but processing complexity increases
Solution Approach 1:
The processing system is segmented into parallel event detection modules, each handling specific types of acoustic events independently. This segmentation allows the system to process multiple microphone signals simultaneously with dedicated detectors for different event types, managing complexity through modular design while maintaining robustness through comprehensive monitoring.
Solution Approach 2:
The event detectors are designed with multi-functionality to handle various types of acoustic events (wind noise, local distortions, double-talk) using unified detection algorithms. This universality reduces overall processing complexity by avoiding the need for separate specialized processors for each event type, while still providing comprehensive detection across all microphone channels.
Data Source
AI summary
Method and apparatus to determine a speaker activity detection measure from energy-based characteristics of signals from a plurality of speaker-dedicated microphones, detect acoustic events using power spectra for the microphone signals, and determine a robust speaker activity detection measure from the speaker activity measure and the detected acoustic events.


