Voice Signal Detection Using Multi-Resolution Framing

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing voice signal detection methods have low accuracy in identifying abrupt exceptions, such as interruptions, starts, and stops, which can affect voice quality assessment and are not effectively reflected in existing quality models.

Innovation Solution

A method and apparatus that frame continuous voice samples into timeframes, analyze energy relationships to detect potential abrupt exceptions, and process tone features using fast-Fourier transforms to determine real abrupt exceptions by examining power density spectra and sound pressure levels.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If existing voice quality assessment models are used, then assessment can be performed, but accuracy in detecting abrupt exceptions is low

Engineering Contradiction:
Improvedetection accuracyVSAvoidassessment reliability
Core Design Contradiction:
Measurement precisionVSReliability

Solution Approach 1:

The patent segments the voice signal into multiple timeframes (first timeframes and second timeframes) with different frame lengths. The first timeframes are used for initial energy-based detection of potential abrupt exceptions, while the second timeframes are used for more detailed tone feature analysis. This segmentation allows the system to process the signal at multiple resolutions, improving both detection accuracy and reliability by combining coarse and fine-grained analysis results.

Inventive Principle:
Principle #1Segmentation

2Measurement precision

If tone feature analysis is added, then detection accuracy improves, but device complexity increases

Engineering Contradiction:
Improvedetection accuracyVSAvoidprocessing complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent performs preliminary action by first detecting potential abrupt exceptions using energy analysis on first timeframes before conducting the more complex tone feature analysis on second timeframes. This preliminary detection filters out obvious non-abrupt segments, allowing the complex FFT-based tone feature extraction to be applied only where needed, thus improving detection accuracy while managing processing complexity through selective application of computationally intensive operations.

Inventive Principle:
Principle #10Preliminary action

Applied Scientific Principles

This section explains which scientific principles are used to turn an abstract innovation direction into a practical engineering solution.

Function Achieved in This Case

Improves the accuracy of detecting abrupt voice signal exceptions by analyzing tone features, effectively identifying real interruptions, starts, and stops, enhancing voice quality assessment.

Implementation Method 1

performing a fast-Fourier transform on each one of the second timeframes to acquire a power density spectrum

Methodology Applied
Scientific EffectFast-Fourier transform:

Data Source

PatentEP2927906B1Method and apparatus for detecting voice signal
Publication Date: 2016.10.05 HUAWEI TECH CO LTD
  • EP2927906B1 patent drawingFigure 1A~2A
  • EP2927906B1 patent drawingFigure 2B
  • EP2927906B1 patent drawingFigure 3

AI summary

The present invention discloses a method and an apparatus for detecting a voice signal. The method includes: performing, in a unit of first timeframe frame length, framing on a continuous voice sample to obtain a plurality of first timeframes, detecting energy of each of the first timeframes, and determining a target first timeframe including a potential abrupt exception of a voice signal by analyzing a relationship between the energy of the plurality of first timeframes; performing, in a unit of second timeframe frame length, framing on the continuous voice sample to obtain a plurality of second timeframes, where each second timeframe frame length is an integral multiple of the first timeframe frame length, and a second timeframe including the target first timeframe is a target second timeframe; and processing each of the second timeframes to acquire a tone feature, and determining, by analyzing a tone feature of at least one of the second timeframes including at least one target second timeframe, whether the potential abrupt exception of a voice signal included in the target first timeframe included in the target second timeframe is a real abrupt exception of a voice signal. The technical solution can improve accuracy in detecting an abrupt exception of a voice signal.