Speaker Activity Detection Using Segmented Event Detectors

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

In speech communication systems, especially in automotive environments, existing technologies face challenges in accurately detecting speaker activity and selecting the optimal microphone, particularly in the presence of interfering acoustic events like wind noise and local distortions, leading to misdetection and poor speech signal processing.

Innovation Solution

The system employs energy-based speaker activity detection using power spectra and signal-to-noise ratio (SNR) analysis to differentiate between speaker activity and acoustic events, allowing for robust joint speaker activity and event detection, and dynamically selects the best microphone for processing based on SNR and event detection results.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Speed

If energy-based speaker activity detection is used to quickly identify speaking passengers, then response speed is improved, but misdetection occurs during interfering acoustic events like wind noise and local distortions

Engineering Contradiction:
Improveresponse speedVSAvoiddetection accuracy
Core Design Contradiction:
SpeedVSReliability

Solution Approach 1:

The detection system is segmented into multiple independent event detectors, each specialized for detecting specific types of acoustic events (wind noise, local distortions, double-talk). This allows the system to quickly identify different event types without requiring a single complex detector, maintaining fast response while improving accuracy through specialized detection for each event category.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The system changes detection parameters dynamically by switching between different detection modes and thresholds based on the current acoustic environment. When interfering events are detected, the system adjusts its detection parameters to distinguish between actual speaker activity and environmental disturbances, thereby maintaining high detection accuracy across varying conditions.

Inventive Principle:
Principle #35Parameter changes

2Reliability

If the system switches to the next best microphone when distortion is detected, then robustness is improved, but false switching occurs when inactive microphones get distorted

Engineering Contradiction:
Improvemicrophone selection robustnessVSAvoidfalse switching
Core Design Contradiction:
ReliabilityVSLoss of information

Solution Approach 1:

The system performs preliminary detection of acoustic events before making microphone selection decisions. By detecting and identifying interfering events in advance, the system can distinguish between distortion caused by environmental factors and actual speaker activity, preventing false switching to alternative microphones when the current microphone is still providing valid speech signals.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The system uses feedback from multiple event detectors to continuously monitor microphone signal quality and adjust selection decisions. The feedback mechanism allows the system to learn from detection results and make more accurate microphone selection decisions, reducing false switching while maintaining robustness against actual signal degradation.

Inventive Principle:
Principle #23Feedback

3Reliability

If multiple microphones are provided per speaker for redundancy, then robustness against distortion is improved, but processing complexity increases

Engineering Contradiction:
ImproverobustnessVSAvoidprocessing complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The processing system is segmented into parallel event detection modules, each handling specific types of acoustic events independently. This segmentation allows the system to process multiple microphone signals simultaneously with dedicated detectors for different event types, managing complexity through modular design while maintaining robustness through comprehensive monitoring.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The event detectors are designed with multi-functionality to handle various types of acoustic events (wind noise, local distortions, double-talk) using unified detection algorithms. This universality reduces overall processing complexity by avoiding the need for separate specialized processors for each event type, while still providing comprehensive detection across all microphone channels.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Data Source

PatentUS9767826B2Methods and apparatus for robust speaker activity detection
Publication Date: 2017.09.19 CERENCE OPERATING CO
  • US9767826B2 patent drawing
  • US9767826B2 patent drawing
  • US9767826B2 patent drawing

AI summary

Method and apparatus to determine a speaker activity detection measure from energy-based characteristics of signals from a plurality of speaker-dedicated microphones, detect acoustic events using power spectra for the microphone signals, and determine a robust speaker activity detection measure from the speaker activity measure and the detected acoustic events.