Speech Recognition Syllabic Nucleus Detection via Phase Features

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional speech recognition systems face challenges in accurately detecting syllabic nuclei and determining speaking rates, especially when syllables are connected or sonorants are present, leading to reduced performance in spontaneous speech recognition.

Innovation Solution

A method and apparatus utilizing a deep neural network to detect syllabic nuclei by extracting both magnitude and phase features from voice signals transformed into the frequency domain, determining speaking rates, and adjusting cepstrum length or overlap factors for improved recognition performance.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If conventional energy and periodicity peak detection is used to determine syllabic nuclei, then syllabic nuclei with clear energy peaks can be detected, but syllabic nuclei connected by sonorants or adjacent to each other cannot be detected accurately

Engineering Contradiction:
Improvesyllabic nucleus detection accuracyVSAvoidhandling of connected syllables and sonorants
Core Design Contradiction:
Measurement precisionVSAdaptability or versatility

Solution Approach 1:

The patent introduces an intermediary acoustic feature (energy decay rate) to bridge the gap between syllabic nuclei that are connected by sonorants or adjacent to each other. This intermediary feature allows the system to detect syllabic boundaries even when traditional peak detection fails, by measuring how quickly energy decays between potential syllable boundaries.

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The patent changes the detection parameter from simple energy peak detection to energy decay rate measurement. By transforming the detection criterion from looking for maximum energy points to measuring the rate of energy change, the system can identify syllabic nuclei even when they are connected by sonorants or occur adjacent to each other without clear energy peaks.

Inventive Principle:
Principle #35Parameter changes

2Adaptability or versatility

If speaking rate variation is not compensated, then the acoustic model can process various speaking rates, but recognition performance deteriorates for spontaneous speech with varying speaking rates

Engineering Contradiction:
Improvehandling of various speaking ratesVSAvoidrecognition performance
Core Design Contradiction:
Adaptability or versatilityVSReliability

Solution Approach 1:

The patent implements dynamic adjustment of acoustic model parameters based on detected speaking rate. The system continuously monitors syllabic nuclei density and adjusts the acoustic model's time scaling parameters in real-time, allowing the model to adapt its processing speed to match the speaker's natural rhythm while maintaining recognition accuracy.

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

The patent changes the acoustic model's processing parameters (time scaling factors) based on the detected speaking rate. By dynamically adjusting these parameters to match the speaker's tempo, the system maintains high recognition performance across varying speaking rates without requiring multiple specialized models.

Inventive Principle:
Principle #35Parameter changes

3Device complexity

If only magnitude features are used as input to the deep neural network, then the network structure is simpler, but syllabic nuclei detection performance is reduced

Engineering Contradiction:
Improvenetwork input feature structureVSAvoidsyllabic nucleus detection accuracy
Core Design Contradiction:
Device complexityVSMeasurement precision

Solution Approach 1:

The patent combines multiple types of acoustic features (magnitude features and phase features) to form a composite feature vector input to the deep neural network. This composite approach leverages the complementary information in both feature types, with magnitude features providing energy information and phase features providing temporal structure information, resulting in improved detection accuracy.

Inventive Principle:
Principle #40Composite materials

Data Source

PatentUS10388275B2Method and apparatus for improving spontaneous speech recognition performance
Publication Date: 2019.08.20 ELECTRONICS & TELECOMM RES INST
  • US10388275B2 patent drawing
  • US10388275B2 patent drawing
  • US10388275B2 patent drawing

AI summary

The present invention relates to a method and apparatus for improving spontaneous speech recognition performance. The present invention is directed to providing a method and apparatus for improving spontaneous speech recognition performance by extracting a phase feature as well as a magnitude feature of a voice signal transformed to the frequency domain, detecting a syllabic nucleus on the basis of a deep neural network using a multi-frame output, determining a speaking rate by dividing the number of syllabic nuclei by a voice section interval detected by a voice detector, calculating a length variation or an overlap factor according to the speaking rate, and performing cepstrum length normalization or time scale modification with a voice length appropriate for an acoustic model.