Voice Activity Detection Using Entropy-Energy Square Root
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing voice activity detection methods face challenges in accurately determining voiced frames due to threshold setting issues affected by recording environments and the inability of spectral entropy-energy product to effectively combine time and frequency domain characteristics, leading to low endpoint detection accuracy.
Innovation Solution
A voice activity detection method that calculates the spectral entropy-energy square root of speech signals, combining time and frequency domain features by determining energy and spectral entropy, and using this value to classify frames as voiced or unvoiced based on preset thresholds, thereby improving detection accuracy.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If traditional voice activity detection methods use threshold setting based on recording environment, then the detection process is simple, but the endpoint detection accuracy deteriorates due to environmental effects
Solution Approach 1:
The patent transforms the traditional linear energy feature into a logarithmic scale by calculating the logarithm of energy, and combines it with spectral entropy to create a new composite feature (log-energy spectral entropy). This parameter transformation makes the feature more robust to environmental variations and improves detection accuracy without significantly increasing process complexity
Solution Approach 2:
The patent creates a composite feature by combining log-energy and spectral entropy components. This composite feature integrates both amplitude information (from energy) and spectral distribution information (from entropy), providing a more comprehensive representation of speech characteristics that is less sensitive to environmental noise
2Measurement precision
If spectral entropy-energy product is used to combine features, then time and frequency domain characteristics are integrated, but the combination effectiveness deteriorates because it cannot adequately reflect voiced frame characteristics
Solution Approach 1:
The patent applies logarithmic transformation to the energy component, changing it from linear to logarithmic scale. This transformation compresses the dynamic range and emphasizes relative energy changes, making the feature more sensitive to voiced frame characteristics while reducing the dominance of high-energy segments
Solution Approach 2:
The patent replaces the traditional multiplicative combination (spectral entropy-energy product) with an additive combination of logarithmic energy and spectral entropy. This substitution changes the interaction mechanism between features, allowing for more balanced contribution of both components and preventing information loss
Data Source
AI summary
This application discloses a voice activity detection method. The method includes receiving speech data, the speech data including a multi-frame speech signal; determining energy and spectral entropy of a frame of speech signal; calculating a square root of the energy of the speech signal and/or calculating a square root of the spectral entropy of the frame of the speech signal; determining a spectral entropy-energy square root of the frame of the speech signal based on at least one of the square root of the energy and the square root of the spectral entropy; and determining that the frame of the speech signal is an unvoiced frame if the spectral entropy-energy square root of the speech signal is less than a first threshold, or that it is a voiced frame if the spectral entropy-energy square root of the speech signal is greater than or equal to the first threshold.


