Pitch Feature Normalization for Unvoiced Frames in ASR
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Automatic speech recognition systems for tonal languages like Mandarin often ignore tone and pitch information, leading to decreased recognition performance, especially in small vocabulary tasks, and struggle to distinguish between homophonic words, as they assign random values to unvoiced frames which cause abnormal likelihood and decrease accuracy.
Innovation Solution
An apparatus and method that evaluates and normalizes the global distribution of random values for unvoiced frames based on the distribution of pitch features from voiced frames, adjusting these values to match the standard deviation and mean of pitch features, and combines these with non-pitch features and voice-level parameters for improved recognition.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If random values are assigned to unvoiced frames as pitch features, then a continuous feature stream can be output, but abnormal likelihood occurs in decoding and recognition performance decreases
Solution Approach 1:
The patent changes the parameter distribution of random values by normalizing them to match the statistical properties (mean and standard deviation) of pitch features from voiced frames. This transforms the random values from an arbitrary distribution to one that aligns with the expected pitch feature distribution, resolving the contradiction between maintaining continuity and ensuring recognition accuracy.
2Reliability
If pitch features are extracted and combined with conventional acoustic features, then recognition performance for tonal languages improves, but the problem of assigning feature values in unvoiced frames becomes more critical
Solution Approach 1:
The system uses the statistical properties (mean and standard deviation) derived from voiced frame pitch features to automatically normalize the random values for unvoiced frames. This self-service approach eliminates the need for manual intervention or complex external algorithms to assign appropriate values, thereby improving recognition performance without proportionally increasing system complexity.
3Measurement precision
If the global distribution of random values for unvoiced frames is normalized based on pitch features of voiced frames, then statistical discrimination between voiced and unvoiced frames improves, but additional processing steps are required
Solution Approach 1:
The patent performs preliminary normalization of random values using the statistical properties (mean and standard deviation) of pitch features from voiced frames before they are used in recognition. This preliminary action ensures that the random values for unvoiced frames already have the correct distribution characteristics, improving statistical discrimination without requiring complex real-time processing during decoding.
Data Source
AI summary
According to one embodiment, an apparatus for applying pitch features in automatic speech recognition is provided. The apparatus includes a distribution evaluation module, normalization module, and random value adjusting module. The distribution evaluation module evaluates the global distribution of pitch features of voiced frames in speech signals, and the global distribution of random values for unvoiced frames in speech signals. The normalization module normalizes the global distribution of random values for unvoiced frames based on the global distribution of pitch features of voiced frames. The random value adjusting module adjusts random values for unvoiced frames based on the normalized global distribution, so that the adjusted random values can be assigned to unvoiced frames in speech signals as pitch features of the unvoiced frames.


