Erroneous Detection Determination Device for Speech Recognition
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing speech recognition systems face challenges in accurately distinguishing speech from noise, particularly in noisy environments with non-stationary noise, leading to erroneous detection and reduced recognition accuracy.
Innovation Solution
The system employs a microphone array to acquire audio signals, calculates a speech arrival rate by determining the proportion of sound from a specific direction, and uses this rate to differentiate between speech and noise, thereby reducing erroneous detection.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If speech recognition is performed in noisy environments, then the system can operate in real-world conditions, but erroneous detection increases and recognition accuracy decreases
Solution Approach 1:
The audio signal is segmented into multiple frequency bands using band-pass filters. By dividing the frequency spectrum into distinct segments (e.g., low, mid, high bands), the system can independently analyze speech characteristics in each band and determine whether speech is present more reliably, reducing erroneous detection in noisy environments.
Solution Approach 2:
The system introduces intermediate processing steps between raw audio input and final recognition output. These include acoustic feature extraction, spectrum envelope calculation, and speech presence determination in multiple frequency bands. These intermediary processes filter out noise and extract meaningful speech characteristics before recognition.
2Reliability
If acoustic feature quantity comparison with stored noise signals is used, then noise resistance improves, but the system cannot handle non-stationary noise effectively
Solution Approach 1:
The system dynamically adapts to changing noise conditions by continuously analyzing acoustic features in real-time and adjusting speech presence determination accordingly. Unlike static comparison with pre-stored noise templates, the system can handle non-stationary noise by evaluating spectrum envelopes and speech characteristics dynamically across multiple frequency bands.
Solution Approach 2:
The system changes analysis parameters adaptively by examining multiple frequency bands with different characteristics. When noise conditions change, the system can weight or emphasize different frequency bands based on where speech characteristics are most distinguishable from noise, effectively handling non-stationary noise environments.
3Object-affected harmful factors
If spectrum envelope removal is applied, then sharp peaks in non-stationary noise are suppressed, but stationary noise with gentle peaks may be affected
Solution Approach 1:
The system applies different processing characteristics to different frequency bands. By analyzing speech presence independently in low, mid, and high frequency bands, the system can locally adapt to different noise types in different spectral regions, suppressing non-stationary noise where it occurs while preserving stationary speech components.
Solution Approach 2:
The system applies spectrum envelope removal selectively rather than uniformly across all frequencies. By determining speech presence in multiple frequency bands and combining results, the system can apply noise suppression where needed while maintaining speech recognition accuracy, avoiding over-suppression of legitimate speech signals.
Data Source
AI summary
An erroneous detection determination device includes: a signal acquisition unit configured to acquire, from each of microphones, a plurality of audio signals relating to ambient sound including sound from a sound source in a certain direction; a result acquisition unit configured to acquire a recognition result including voice activity information indicating the inclusion of a voice activity relating to at least one of the audio signals; a calculation unit configured to calculate, for each of audio signals on the basis of the signals in respective unit times and the certain direction, a speech arrival rate representing the proportion of the sound from the certain direction to the ambient sound in each of the unit times; and an error detection unit configured to determine, on the basis of the recognition result and the speech arrival rate, whether or not the voice activity information is the result of erroneous detection.


