Adaptive Speech Segmentation via Instantaneous Phase

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional speech signal segmentation techniques, such as fixed frame size and rate (FFSR), are ineffective in recognizing consonants due to their fixed frame lengths and inability to adapt to signal properties, leading to poor noise robustness and recognition accuracy.

Innovation Solution

A method that segments speech signals into frames of varying lengths based on the low frequency component's instantaneous phase, using Hilbert transform to extract phase information and divide it into phase-sections, allowing adaptive frame determination and improved noise robustness.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Ease of manufacture

If fixed frame size and rate (FFSR) technique is used to segment speech signals, then the processing is simple and consistent, but the recognition accuracy for consonants deteriorates due to inability to adapt to signal properties

Engineering Contradiction:
Improveprocessing simplicityVSAvoidrecognition accuracy
Core Design Contradiction:
Ease of manufactureVSMeasurement precision

Solution Approach 1:

The patent applies dynamics by transitioning from fixed frame size to variable frame size segmentation. The frame size is dynamically adjusted based on the instantaneous phase of the low frequency component, allowing the segmentation to adapt to the temporal characteristics of different speech signals. This resolves the contradiction by making the processing flexible enough to maintain accuracy for both vowels and consonants while preserving relative simplicity through automated phase-based determination.

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

The patent changes the parameter of frame size from fixed to variable based on the instantaneous phase parameter. By using the instantaneous phase of the low frequency component as a basis for determining frame boundaries, the system adapts the segmentation parameters to match the signal properties, thereby improving recognition accuracy without significantly complicating the processing pipeline.

Inventive Principle:
Principle #35Parameter changes

2Stability of the object's composition

If fixed frame size is used for all speech signals, then the processing is consistent, but the noise robustness deteriorates because features unique to speech signals are not well represented

Engineering Contradiction:
Improveprocessing consistencyVSAvoidnoise robustness
Core Design Contradiction:
Stability of the object's compositionVSReliability

Solution Approach 1:

The patent introduces dynamic frame size adjustment based on instantaneous phase to improve noise robustness. By adapting the frame size to the signal's temporal characteristics, the segmentation better preserves speech-specific features even in noisy conditions, while the automated phase-based approach maintains processing consistency through a systematic adaptive method.

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

The patent applies local quality by allowing different frame sizes for different segments of the speech signal based on their instantaneous phase characteristics. This enables the processing to be locally optimized for each segment's properties, improving noise robustness by preserving unique speech features in each local region while maintaining overall processing consistency through the unified phase-based framework.

Inventive Principle:
Principle #3Local quality

3Quantity of substance

If conventional MFCC technique is used to extract features, then all frequency components are captured, but the recognition accuracy deteriorates because noise frequency components are included in the feature vectors

Engineering Contradiction:
Improvefrequency information completenessVSAvoidrecognition accuracy
Core Design Contradiction:
Quantity of substanceVSMeasurement precision

Solution Approach 1:

The patent extracts only the low frequency component and its instantaneous phase information, discarding the full spectrum analysis approach of conventional MFCC. By taking out only the essential phase information from the low frequency component, the system eliminates noise frequency components while preserving the critical temporal structure needed for accurate speech recognition.

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The patent uses the instantaneous phase of the low frequency component as an intermediary to represent the speech signal's temporal characteristics. This intermediary approach allows the system to capture essential speech features without directly processing all frequency components, thereby excluding noise while maintaining recognition accuracy through the phase-based representation.

Inventive Principle:
Principle #24Intermediary (Mediator)

Applied Scientific Principles

This section explains which scientific principles are used to turn an abstract innovation direction into a practical engineering solution.

Function Achieved in This Case

This approach enhances speech recognition accuracy by effectively capturing phoneme transitions and reducing overlapping time intervals, providing more robust feature extraction and improved noise resistance compared to traditional methods.

Implementation Method 1

By performing Hilbert transform on the low frequency signal (similar to LFP) of the speech signal, it is possible to extract instantaneous phase information having a value of −π to π.

Methodology Applied
Scientific EffectHilbert transform:

Data Source

PatentUS10008198B2Nested segmentation method for speech recognition based on sound processing of brain
Publication Date: 2018.06.26 KOREA ADVANCED INST OF SCI & TECH
  • US10008198B2 patent drawing
  • US10008198B2 patent drawing
  • US10008198B2 patent drawing

AI summary

A method of segmenting input speech signal into plurality of frames for speech recognition is disclosed. The method includes extracting a low frequency signal from the speech signal, and segmenting the speech signal into a plurality of time-intervals according to a plurality of instantaneous phase-sections of the low frequency signal.