Speech Recognition Syllabic Nucleus Detection via Phase Features
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional speech recognition systems face challenges in accurately detecting syllabic nuclei and determining speaking rates, especially when syllables are connected or sonorants are present, leading to reduced performance in spontaneous speech recognition.
Innovation Solution
A method and apparatus utilizing a deep neural network to detect syllabic nuclei by extracting both magnitude and phase features from voice signals transformed into the frequency domain, determining speaking rates, and adjusting cepstrum length or overlap factors for improved recognition performance.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If conventional energy and periodicity peak detection is used to determine syllabic nuclei, then syllabic nuclei with clear energy peaks can be detected, but syllabic nuclei connected by sonorants or adjacent to each other cannot be detected accurately
Solution Approach 1:
The patent introduces an intermediary acoustic feature (energy decay rate) to bridge the gap between syllabic nuclei that are connected by sonorants or adjacent to each other. This intermediary feature allows the system to detect syllabic boundaries even when traditional peak detection fails, by measuring how quickly energy decays between potential syllable boundaries.
Solution Approach 2:
The patent changes the detection parameter from simple energy peak detection to energy decay rate measurement. By transforming the detection criterion from looking for maximum energy points to measuring the rate of energy change, the system can identify syllabic nuclei even when they are connected by sonorants or occur adjacent to each other without clear energy peaks.
2Adaptability or versatility
If speaking rate variation is not compensated, then the acoustic model can process various speaking rates, but recognition performance deteriorates for spontaneous speech with varying speaking rates
Solution Approach 1:
The patent implements dynamic adjustment of acoustic model parameters based on detected speaking rate. The system continuously monitors syllabic nuclei density and adjusts the acoustic model's time scaling parameters in real-time, allowing the model to adapt its processing speed to match the speaker's natural rhythm while maintaining recognition accuracy.
Solution Approach 2:
The patent changes the acoustic model's processing parameters (time scaling factors) based on the detected speaking rate. By dynamically adjusting these parameters to match the speaker's tempo, the system maintains high recognition performance across varying speaking rates without requiring multiple specialized models.
3Device complexity
If only magnitude features are used as input to the deep neural network, then the network structure is simpler, but syllabic nuclei detection performance is reduced
Solution Approach 1:
The patent combines multiple types of acoustic features (magnitude features and phase features) to form a composite feature vector input to the deep neural network. This composite approach leverages the complementary information in both feature types, with magnitude features providing energy information and phase features providing temporal structure information, resulting in improved detection accuracy.
Data Source
AI summary
The present invention relates to a method and apparatus for improving spontaneous speech recognition performance. The present invention is directed to providing a method and apparatus for improving spontaneous speech recognition performance by extracting a phase feature as well as a magnitude feature of a voice signal transformed to the frequency domain, detecting a syllabic nucleus on the basis of a deep neural network using a multi-frame output, determining a speaking rate by dividing the number of syllabic nuclei by a voice section interval detected by a voice detector, calculating a length variation or an overlap factor according to the speaking rate, and performing cepstrum length normalization or time scale modification with a voice length appropriate for an acoustic model.


