Non-spatial Speech Detection Using Covariance Matrix Algorithm

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing microphone systems in vehicular environments face challenges in accurately detecting speech in noisy conditions due to multiple noise sources and acoustic reflections, which are not effectively addressed by previous algorithms that rely heavily on phase/angular data and fail to utilize amplitude information effectively.

Innovation Solution

A non-spatial speech detection system utilizing a plurality of microphones with both fixed and adaptive beamformers, processing outputs through a covariance matrix-based algorithm to identify speech from noise, leveraging the determinant of a Gram matrix for improved speech detection in high-noise environments.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If traditional speech detection algorithms are used in vehicular environments, then the system structure remains simple, but speech detection accuracy deteriorates due to noise interference and acoustic reflections

Engineering Contradiction:
Improvespeech detection accuracyVSAvoidnoise interference
Core Design Contradiction:
Measurement precisionVSObject-affected harmful factors

Solution Approach 1:

The patent divides the speech detection task into multiple stages: initial speech/noise classification using spectral features, followed by refinement using spectral subtraction and inverse filtering. This segmented approach allows each stage to address specific aspects of noise reduction, improving overall detection accuracy in vehicular environments

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent transitions from traditional spatial-based speech detection to a spectral-domain approach by analyzing frequency spectra, spectral ratios, and spectral slopes. This dimensional shift to spectral features enables effective speech detection without relying on spatial separation, overcoming limitations in noisy vehicular acoustics

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

2Reliability

If algorithms relying heavily on phase/angular data are used, then spatial information is utilized, but performance deteriorates in high-noise environments where phase information is unreliable

Engineering Contradiction:
Improvespeech identification reliabilityVSAvoidspeech detection in high noise
Core Design Contradiction:
ReliabilityVSDifficulty of detecting and measuring

Solution Approach 1:

The patent changes the detection parameters from phase-based spatial features to spectral-based features including spectral ratios, spectral slopes, and zero-crossing rates. This parameter transformation makes the detection robust to phase corruption in high-noise environments while maintaining reliability through alternative acoustic cues

Inventive Principle:
Principle #35Parameter changes

3Measurement precision

If amplitude information is not utilized effectively, then the system remains simple, but speech detection performance deteriorates in low signal-to-noise ratio conditions

Engineering Contradiction:
Improvespeech detection precisionVSAvoidamplitude information utilization
Core Design Contradiction:
Measurement precisionVSLoss of information

Solution Approach 1:

The patent implements an iterative refinement process where initial speech detection results feed into spectral subtraction, which then feeds into inverse filtering, with each stage refining the estimate based on feedback from previous stages. This feedback loop progressively improves speech detection precision by leveraging amplitude information at multiple processing levels

Inventive Principle:
Principle #23Feedback

Data Source

PatentUS8935164B2Non-spatial speech detection system and method of using same
Publication Date: 2015.01.13 HL KLEMOVE CORP
  • US8935164B2 patent drawing
  • US8935164B2 patent drawing
  • US8935164B2 patent drawing

AI summary

A non-spatial speech detection system includes a plurality of microphones whose output is supplied to a fixed beamformer. An adaptive beamformer is used for receiving the output of the plurality of microphones and one or more processors are used for processing an output from the fixed beamformer and identifying speech from noise though the use of an algorithm utilizing a covariance matrix.