Sub-Band Speech Onset Detection for Low-Power Noise Robustness
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing voice-controlled devices struggle to accurately detect the onset of speech in noisy environments with low power consumption and complexity, leading to inefficient power usage and high false positive rates.
Innovation Solution
A low-power and low-complexity speech onset detector (SOD) using a fractional-band filter structure and spectral subtraction technique to derive sub-band energy profiles, which includes a cascade of low-pass filters and spectral subtraction to distinguish speech from noise.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If the SOD remains in an active state to detect speech onset at any time, then the detection accuracy is improved, but the power consumption increases
Solution Approach 1:
The SOD dynamically transitions between sleep mode and active state based on detection needs. The system activates the SOD only when speech onset detection is required, rather than maintaining a static active state, thus reducing power consumption while preserving detection capability.
Solution Approach 2:
The SOD operates in periodic cycles of sleep and active states. By scheduling detection periods and using idle states between detections, the system reduces overall power consumption while maintaining the ability to detect speech onset when needed.
2Measurement precision
If the SOD uses complex algorithms to improve detection accuracy in noisy environments, then the detection precision is improved, but the device complexity increases
Solution Approach 1:
The detection process is segmented into multiple stages: a low-complexity energy-based detector first identifies potential speech onsets, followed by a more complex neural network classifier that processes only these candidate segments. This segmentation allows accurate detection while keeping the always-on portion simple.
Solution Approach 2:
An energy-based detection mechanism serves as an intermediary between the simple always-on sensor and the complex neural network classifier. This intermediary filters and pre-processes signals, reducing the computational burden on the complex algorithm while maintaining overall detection accuracy.
3Reliability
If the SOD activates the wake word detection algorithm frequently to ensure no speech is missed, then the detection reliability is improved, but the power consumption increases
Solution Approach 1:
The system applies different quality levels of detection to different time periods and signal conditions. High-reliability neural network classification is applied only to candidate segments identified by the energy detector, while other periods use simpler or no detection, thus maintaining reliability where needed while reducing overall power consumption.
Data Source
AI summary
Techniques are disclosed for a low-power and low-complexity speech onset detector (SOD) that uses a fractional-band filter structure and spectral subtraction technique to derive sub-band energy profiles to detect the onset of speech in the presence of noise. The SOD derives the sub-band energy profiles by filtering and down-sampling a full-band input audio signal using the fractional-bandwidth filter structure, which may be a low-pass filter with a cut-off frequency that is a fraction of the full bandwidth of the input signal. The SOD flexibly estimates the average noise energy across frames and the current frame speech energy in each sub-band to track noise and speech energy levels across the frames for each of the sub-bands to determine one or more band thresholds used to detect active speech. The sub-band energy profiles leverage any separation in frequency between noise and speech to detect the onset of speech in a target signal.


