Speech Detection System Using Noise Estimates

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Speech processing systems face challenges in identifying speech segments effectively in noisy environments, leading to reduced speech intelligibility and recognition errors, especially in non-stationary noise conditions.

Innovation Solution

A system that converts time-varying input signals into digital-domain signals, applies a window function to filter out unwanted frequencies, and uses a background voice detector and noise estimator to compare the strength of speech segments against noise levels, dynamically adjusting thresholds to improve speech detection accuracy.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If speech processing systems operate in noisy environments, then speech recognition capability is maintained, but speech intelligibility decreases and recognition errors increase

Engineering Contradiction:
Improvespeech recognition capabilityVSAvoidspeech intelligibility
Core Design Contradiction:
ReliabilityVSMeasurement precision

Solution Approach 1:

The patent introduces background voice estimates and noise estimates as intermediary parameters that mediate between the raw audio signal and the speech recognition decision. The background voice detector estimates background speech strength, while the noise estimator quantifies noise levels. These intermediary estimates allow the system to distinguish between relevant speech and irrelevant background noise, improving speech intelligibility while maintaining recognition capability in noisy environments.

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The system uses feedback from background voice detectors and noise estimators to dynamically adjust speech detection thresholds and processing parameters. By continuously monitoring background voice strength and noise levels, the system can adapt its speech recognition sensitivity in real-time, preventing misidentification of background noise as speech while maintaining the ability to detect relevant speech segments.

Inventive Principle:
Principle #23Feedback

2Measurement precision

If the system detects background voice strength and noise levels, then speech detection accuracy improves, but system complexity increases

Engineering Contradiction:
Improvespeech detection accuracyVSAvoidsystem complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent segments the audio processing system into distinct functional modules: a background voice detector that estimates background speech strength, a noise estimator that quantifies noise levels, and a speech detector that integrates these estimates. This segmentation allows each module to perform a specific function independently, making the overall complex task of accurate speech detection in noisy environments more manageable and implementable.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The system performs preliminary actions by estimating background voice strength and noise levels before making final speech detection decisions. The background voice detector and noise estimator operate in advance to characterize the acoustic environment, providing preparatory information that guides subsequent speech segment detection and reduces false alarms.

Inventive Principle:
Principle #10Preliminary action

3Reliability

If the system uses multiple frequency bins and dynamic threshold adjustment, then speech recognition performance improves, but processing time increases

Engineering Contradiction:
Improvespeech recognition performanceVSAvoidprocessing time
Core Design Contradiction:
ReliabilityVSLoss of time

Solution Approach 1:

The patent applies partial action by processing only the necessary frequency bins rather than analyzing all possible frequencies. The noise estimator and background voice detector focus on specific frequency ranges where speech and noise are most distinguishable, reducing the computational burden while maintaining recognition performance. This selective processing approach achieves adequate performance without requiring exhaustive analysis of all frequency components.

Inventive Principle:
Principle #16Partial or excessive action

Data Source

PatentUS8311819B2System for detecting speech with background voice estimates and noise estimates
Publication Date: 2012.11.13 BLACKBERRY LTD
  • US8311819B2 patent drawing
  • US8311819B2 patent drawing
  • US8311819B2 patent drawing

AI summary

A system detects a speech segment that may include unvoiced, fully voiced, or mixed voice content. The system includes a digital converter that converts a time-varying input signal into a digital-domain signal. A window function passes signals within a programmed aural frequency range while substantially blocking signals above and below the programmed aural frequency range when multiplied by an output of the digital converter. A frequency converter converts the signals passing within the programmed aural frequency range into a plurality of frequency bins. A background voice detector estimates the strength of a background speech segment relative to the noise of selected portions of the aural spectrum. A noise estimator estimates a maximum distribution of noise to an average of an acoustic noise power of some of the plurality of frequency bins. A voice detector compares the strength of a desired speech segment to a criterion based on an output of the background voice detector and an output of the noise estimator.