Speech Segment Determination Using Spectral Entropy
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing methods struggle to accurately determine speech segments in input signals, especially when non-stationary noise is present, as they rely on signal power or spectral entropy alone, which can lead to inaccurate differentiation between speech and noise.
Innovation Solution
A device and method that divides input signals into frames, calculates and increases the power spectrum, and then determines speech segments based on spectral entropy, distinguishing between speech and colored noise by altering the power spectrum values to enhance entropy differences.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Ease of manufacture
If signal power is used to determine speech segments, then the method is simple to implement, but accuracy deteriorates when signal level varies or non-stationary noise is present
Solution Approach 1:
The patent transforms the speech detection approach from using raw signal power to using spectral entropy calculated from the power spectrum. This parameter transformation enables accurate speech segment determination even when signal levels vary or non-stationary noise is present, as spectral entropy is normalized and insensitive to absolute power levels.
2Measurement precision
If spectral entropy is used to determine speech segments, then accuracy improves for varying signal levels, but real-time determination becomes difficult when non-stationary noise is present
Solution Approach 1:
The patent applies dynamic noise spectrum estimation that adapts to non-stationary noise conditions. The noise spectrum is continuously updated based on recent signal characteristics, allowing the spectral entropy calculation to remain effective in real-time even as noise properties change over time.
3Productivity
If traditional speech detection methods are used, then processing is fast, but differentiation between speech and colored noise becomes inaccurate
Solution Approach 1:
The patent introduces spectral entropy as an intermediary measure between the raw power spectrum and speech detection decision. This intermediary transformation enhances the distinguishability between speech and colored noise by converting power spectrum values into an entropy metric that highlights spectral distribution differences rather than absolute power differences.
Data Source
AI summary
A speech segment determination device includes a frame division portion, a power spectrum calculation portion, a power spectrum operation portion, a spectral entropy calculation portion and a determination portion. The frame division portion divides an input signal in units of frames. The power spectrum calculation portion calculates, using an analysis length, a power spectrum of the input signal for each of the frames that have been divided. The power spectrum operation portion adds a value of the calculated power spectrum to a value of power spectrum in each of frequency bins. The spectral entropy calculation portion calculates spectral entropy using the power spectrum whose value has been increased. The determination portion determines, based on a value of the spectral entropy, whether the input signal is a signal in a speech segment.


