Non-Speech Section Detection via Spectrum Bias Analysis
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing speech recognition devices face challenges in accurately detecting non-speech sections, particularly in environments with strong non-stationary noise, leading to erroneous speech recognition due to incorrect identification of noise as speech.
Innovation Solution
A non-speech section detecting device that generates frames from sound data and calculates the bias of the spectrum, using thresholds to identify consecutive frames without voice data, thereby detecting non-speech sections and improving speech recognition accuracy.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If power-based speech section detection is used, then speech sections can be detected, but non-speech sections containing noise with strong non-stationarity are erroneously detected as speech sections
Solution Approach 1:
The patent changes the detection parameter from power-based to spectrum bias-based. By calculating the bias of the spectrum obtained by converting sound data into frequency components, the system can distinguish speech from non-speech more accurately. Speech sections have different spectrum bias characteristics compared to non-speech sections, allowing the system to avoid erroneous detection of non-stationary noise as speech.
Solution Approach 2:
The patent introduces dynamic threshold adjustment based on consecutive frame analysis. Instead of using a fixed threshold, the system counts the number of consecutive frames with similar spectrum bias values and dynamically adjusts the detection threshold. This dynamic approach allows the system to adapt to varying noise conditions and avoid false detection of transient noise as speech.
2Reliability
If correction coefficient is calculated from maximum speech power and speech recognition result, then future criterion value is corrected, but detection accuracy degrades in environments with strong non-stationary noise
Solution Approach 1:
The patent fundamentally changes the parameter used for criterion value correction from power-based to spectrum bias-based. By using spectrum bias, which captures the distribution characteristics of frequency components, the system can correct criterion values in a way that is robust to non-stationary noise. The spectrum bias reflects the inherent structure of speech signals rather than just their power, making the correction more reliable in noisy environments.
Solution Approach 2:
The patent implements feedback through consecutive frame analysis. The system continuously monitors spectrum bias values across multiple frames and uses this feedback to adjust detection decisions. By counting consecutive frames with similar bias characteristics, the system provides feedback that helps distinguish true speech sections from noise, improving both reliability and precision in criterion value application.
Data Source
AI summary
A non-speech section detecting device generating a plurality of frames having a given time length on the basis of sound data obtained by sampling sound, and detecting a non-speech section having a frame not containing voice data based on speech uttered by a person, the device including: a calculating part calculating a bias of a spectrum obtained by converting sound data of each frame into components on a frequency axis; a judging part judging whether the bias is greater than or equal to a given threshold or alternatively smaller than or equal to a given threshold; a counting part counting the number of consecutive frames judged as having a bias greater than or equal to the threshold or alternatively smaller than or equal to the threshold; a count judging part judging whether the obtained number of consecutive frames is greater than or equal to a given value.


