Non-spatial Speech Detection Using Covariance Matrix Algorithm
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing microphone systems in vehicular environments face challenges in accurately detecting speech in noisy conditions due to multiple noise sources and acoustic reflections, which are not effectively addressed by previous algorithms that rely heavily on phase/angular data and fail to utilize amplitude information effectively.
Innovation Solution
A non-spatial speech detection system utilizing a plurality of microphones with both fixed and adaptive beamformers, processing outputs through a covariance matrix-based algorithm to identify speech from noise, leveraging the determinant of a Gram matrix for improved speech detection in high-noise environments.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If traditional speech detection algorithms are used in vehicular environments, then the system structure remains simple, but speech detection accuracy deteriorates due to noise interference and acoustic reflections
Solution Approach 1:
The patent divides the speech detection task into multiple stages: initial speech/noise classification using spectral features, followed by refinement using spectral subtraction and inverse filtering. This segmented approach allows each stage to address specific aspects of noise reduction, improving overall detection accuracy in vehicular environments
Solution Approach 2:
The patent transitions from traditional spatial-based speech detection to a spectral-domain approach by analyzing frequency spectra, spectral ratios, and spectral slopes. This dimensional shift to spectral features enables effective speech detection without relying on spatial separation, overcoming limitations in noisy vehicular acoustics
2Reliability
If algorithms relying heavily on phase/angular data are used, then spatial information is utilized, but performance deteriorates in high-noise environments where phase information is unreliable
Solution Approach 1:
The patent changes the detection parameters from phase-based spatial features to spectral-based features including spectral ratios, spectral slopes, and zero-crossing rates. This parameter transformation makes the detection robust to phase corruption in high-noise environments while maintaining reliability through alternative acoustic cues
3Measurement precision
If amplitude information is not utilized effectively, then the system remains simple, but speech detection performance deteriorates in low signal-to-noise ratio conditions
Solution Approach 1:
The patent implements an iterative refinement process where initial speech detection results feed into spectral subtraction, which then feeds into inverse filtering, with each stage refining the estimate based on feedback from previous stages. This feedback loop progressively improves speech detection precision by leveraging amplitude information at multiple processing levels
Data Source
AI summary
A non-spatial speech detection system includes a plurality of microphones whose output is supplied to a fixed beamformer. An adaptive beamformer is used for receiving the output of the plurality of microphones and one or more processors are used for processing an output from the fixed beamformer and identifying speech from noise though the use of an algorithm utilizing a covariance matrix.


