Speech Segment Determination Using Spectral Entropy

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing methods struggle to accurately determine speech segments in input signals, especially when non-stationary noise is present, as they rely on signal power or spectral entropy alone, which can lead to inaccurate differentiation between speech and noise.

Innovation Solution

A device and method that divides input signals into frames, calculates and increases the power spectrum, and then determines speech segments based on spectral entropy, distinguishing between speech and colored noise by altering the power spectrum values to enhance entropy differences.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Ease of manufacture

If signal power is used to determine speech segments, then the method is simple to implement, but accuracy deteriorates when signal level varies or non-stationary noise is present

Engineering Contradiction:
Improveimplementation simplicityVSAvoidspeech segment determination accuracy
Core Design Contradiction:
Ease of manufactureVSMeasurement precision

Solution Approach 1:

The patent transforms the speech detection approach from using raw signal power to using spectral entropy calculated from the power spectrum. This parameter transformation enables accurate speech segment determination even when signal levels vary or non-stationary noise is present, as spectral entropy is normalized and insensitive to absolute power levels.

Inventive Principle:
Principle #35Parameter changes

2Measurement precision

If spectral entropy is used to determine speech segments, then accuracy improves for varying signal levels, but real-time determination becomes difficult when non-stationary noise is present

Engineering Contradiction:
Improvespeech segment determination accuracyVSAvoidreal-time processing capability
Core Design Contradiction:
Measurement precisionVSProductivity

Solution Approach 1:

The patent applies dynamic noise spectrum estimation that adapts to non-stationary noise conditions. The noise spectrum is continuously updated based on recent signal characteristics, allowing the spectral entropy calculation to remain effective in real-time even as noise properties change over time.

Inventive Principle:
Principle #15Dynamics

3Productivity

If traditional speech detection methods are used, then processing is fast, but differentiation between speech and colored noise becomes inaccurate

Engineering Contradiction:
Improveprocessing speedVSAvoidspeech vs noise differentiation accuracy
Core Design Contradiction:
ProductivityVSMeasurement precision

Solution Approach 1:

The patent introduces spectral entropy as an intermediary measure between the raw power spectrum and speech detection decision. This intermediary transformation enhances the distinguishability between speech and colored noise by converting power spectrum values into an entropy metric that highlights spectral distribution differences rather than absolute power differences.

Inventive Principle:
Principle #24Intermediary (Mediator)

Data Source

PatentUS9123351B2Speech segment determination device, and storage medium
Publication Date: 2015.09.01 OKI ELECTRIC INDUSTRY CO LTD
  • US9123351B2 patent drawing
  • US9123351B2 patent drawing
  • US9123351B2 patent drawing

AI summary

A speech segment determination device includes a frame division portion, a power spectrum calculation portion, a power spectrum operation portion, a spectral entropy calculation portion and a determination portion. The frame division portion divides an input signal in units of frames. The power spectrum calculation portion calculates, using an analysis length, a power spectrum of the input signal for each of the frames that have been divided. The power spectrum operation portion adds a value of the calculated power spectrum to a value of power spectrum in each of frequency bins. The spectral entropy calculation portion calculates spectral entropy using the power spectrum whose value has been increased. The determination portion determines, based on a value of the spectral entropy, whether the input signal is a signal in a speech segment.