Voice Activity Detection Using Entropy-Energy Square Root

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing voice activity detection methods face challenges in accurately determining voiced frames due to threshold setting issues affected by recording environments and the inability of spectral entropy-energy product to effectively combine time and frequency domain characteristics, leading to low endpoint detection accuracy.

Innovation Solution

A voice activity detection method that calculates the spectral entropy-energy square root of speech signals, combining time and frequency domain features by determining energy and spectral entropy, and using this value to classify frames as voiced or unvoiced based on preset thresholds, thereby improving detection accuracy.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If traditional voice activity detection methods use threshold setting based on recording environment, then the detection process is simple, but the endpoint detection accuracy deteriorates due to environmental effects

Engineering Contradiction:
Improveendpoint detection accuracyVSAvoiddetection process complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent transforms the traditional linear energy feature into a logarithmic scale by calculating the logarithm of energy, and combines it with spectral entropy to create a new composite feature (log-energy spectral entropy). This parameter transformation makes the feature more robust to environmental variations and improves detection accuracy without significantly increasing process complexity

Inventive Principle:
Principle #35Parameter changes

Solution Approach 2:

The patent creates a composite feature by combining log-energy and spectral entropy components. This composite feature integrates both amplitude information (from energy) and spectral distribution information (from entropy), providing a more comprehensive representation of speech characteristics that is less sensitive to environmental noise

Inventive Principle:
Principle #40Composite materials

2Measurement precision

If spectral entropy-energy product is used to combine features, then time and frequency domain characteristics are integrated, but the combination effectiveness deteriorates because it cannot adequately reflect voiced frame characteristics

Engineering Contradiction:
Improvevoiced frame detection accuracyVSAvoidinformation loss in feature combination
Core Design Contradiction:
Measurement precisionVSLoss of information

Solution Approach 1:

The patent applies logarithmic transformation to the energy component, changing it from linear to logarithmic scale. This transformation compresses the dynamic range and emphasizes relative energy changes, making the feature more sensitive to voiced frame characteristics while reducing the dominance of high-energy segments

Inventive Principle:
Principle #35Parameter changes

Solution Approach 2:

The patent replaces the traditional multiplicative combination (spectral entropy-energy product) with an additive combination of logarithmic energy and spectral entropy. This substitution changes the interaction mechanism between features, allowing for more balanced contribution of both components and preventing information loss

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Data Source

PatentUS11138992B2Voice activity detection based on entropy-energy feature
Publication Date: 2021.10.05 TENCENT TECHNOLOGY (SHENZHEN) CO LTD
  • US11138992B2 patent drawing
  • US11138992B2 patent drawing
  • US11138992B2 patent drawing

AI summary

This application discloses a voice activity detection method. The method includes receiving speech data, the speech data including a multi-frame speech signal; determining energy and spectral entropy of a frame of speech signal; calculating a square root of the energy of the speech signal and/or calculating a square root of the spectral entropy of the frame of the speech signal; determining a spectral entropy-energy square root of the frame of the speech signal based on at least one of the square root of the energy and the square root of the spectral entropy; and determining that the frame of the speech signal is an unvoiced frame if the spectral entropy-energy square root of the speech signal is less than a first threshold, or that it is a voiced frame if the spectral entropy-energy square root of the speech signal is greater than or equal to the first threshold.