Speech Detection System Using Noise Estimates
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Speech processing systems face challenges in identifying speech segments effectively in noisy environments, leading to reduced speech intelligibility and recognition errors, especially in non-stationary noise conditions.
Innovation Solution
A system that converts time-varying input signals into digital-domain signals, applies a window function to filter out unwanted frequencies, and uses a background voice detector and noise estimator to compare the strength of speech segments against noise levels, dynamically adjusting thresholds to improve speech detection accuracy.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If speech processing systems operate in noisy environments, then speech recognition capability is maintained, but speech intelligibility decreases and recognition errors increase
Solution Approach 1:
The patent introduces background voice estimates and noise estimates as intermediary parameters that mediate between the raw audio signal and the speech recognition decision. The background voice detector estimates background speech strength, while the noise estimator quantifies noise levels. These intermediary estimates allow the system to distinguish between relevant speech and irrelevant background noise, improving speech intelligibility while maintaining recognition capability in noisy environments.
Solution Approach 2:
The system uses feedback from background voice detectors and noise estimators to dynamically adjust speech detection thresholds and processing parameters. By continuously monitoring background voice strength and noise levels, the system can adapt its speech recognition sensitivity in real-time, preventing misidentification of background noise as speech while maintaining the ability to detect relevant speech segments.
2Measurement precision
If the system detects background voice strength and noise levels, then speech detection accuracy improves, but system complexity increases
Solution Approach 1:
The patent segments the audio processing system into distinct functional modules: a background voice detector that estimates background speech strength, a noise estimator that quantifies noise levels, and a speech detector that integrates these estimates. This segmentation allows each module to perform a specific function independently, making the overall complex task of accurate speech detection in noisy environments more manageable and implementable.
Solution Approach 2:
The system performs preliminary actions by estimating background voice strength and noise levels before making final speech detection decisions. The background voice detector and noise estimator operate in advance to characterize the acoustic environment, providing preparatory information that guides subsequent speech segment detection and reduces false alarms.
3Reliability
If the system uses multiple frequency bins and dynamic threshold adjustment, then speech recognition performance improves, but processing time increases
Solution Approach 1:
The patent applies partial action by processing only the necessary frequency bins rather than analyzing all possible frequencies. The noise estimator and background voice detector focus on specific frequency ranges where speech and noise are most distinguishable, reducing the computational burden while maintaining recognition performance. This selective processing approach achieves adequate performance without requiring exhaustive analysis of all frequency components.
Data Source
AI summary
A system detects a speech segment that may include unvoiced, fully voiced, or mixed voice content. The system includes a digital converter that converts a time-varying input signal into a digital-domain signal. A window function passes signals within a programmed aural frequency range while substantially blocking signals above and below the programmed aural frequency range when multiplied by an output of the digital converter. A frequency converter converts the signals passing within the programmed aural frequency range into a plurality of frequency bins. A background voice detector estimates the strength of a background speech segment relative to the noise of selected portions of the aural spectrum. A noise estimator estimates a maximum distribution of noise to an average of an acoustic noise power of some of the plurality of frequency bins. A voice detector compares the strength of a desired speech segment to a criterion based on an output of the background voice detector and an output of the noise estimator.


