Pitch Feature Normalization for Unvoiced Frames in ASR

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Automatic speech recognition systems for tonal languages like Mandarin often ignore tone and pitch information, leading to decreased recognition performance, especially in small vocabulary tasks, and struggle to distinguish between homophonic words, as they assign random values to unvoiced frames which cause abnormal likelihood and decrease accuracy.

Innovation Solution

An apparatus and method that evaluates and normalizes the global distribution of random values for unvoiced frames based on the distribution of pitch features from voiced frames, adjusting these values to match the standard deviation and mean of pitch features, and combines these with non-pitch features and voice-level parameters for improved recognition.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If random values are assigned to unvoiced frames as pitch features, then a continuous feature stream can be output, but abnormal likelihood occurs in decoding and recognition performance decreases

Engineering Contradiction:
Improvecontinuous feature stream outputVSAvoidrecognition performance
Core Design Contradiction:
ProductivityVSReliability

Solution Approach 1:

The patent changes the parameter distribution of random values by normalizing them to match the statistical properties (mean and standard deviation) of pitch features from voiced frames. This transforms the random values from an arbitrary distribution to one that aligns with the expected pitch feature distribution, resolving the contradiction between maintaining continuity and ensuring recognition accuracy.

Inventive Principle:
Principle #35Parameter changes

2Reliability

If pitch features are extracted and combined with conventional acoustic features, then recognition performance for tonal languages improves, but the problem of assigning feature values in unvoiced frames becomes more critical

Engineering Contradiction:
Improverecognition performance for tonal languagesVSAvoidfeature assignment complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The system uses the statistical properties (mean and standard deviation) derived from voiced frame pitch features to automatically normalize the random values for unvoiced frames. This self-service approach eliminates the need for manual intervention or complex external algorithms to assign appropriate values, thereby improving recognition performance without proportionally increasing system complexity.

Inventive Principle:
Principle #25Self-service

3Measurement precision

If the global distribution of random values for unvoiced frames is normalized based on pitch features of voiced frames, then statistical discrimination between voiced and unvoiced frames improves, but additional processing steps are required

Engineering Contradiction:
Improvestatistical discrimination precisionVSAvoidprocessing steps
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent performs preliminary normalization of random values using the statistical properties (mean and standard deviation) of pitch features from voiced frames before they are used in recognition. This preliminary action ensures that the random values for unvoiced frames already have the correct distribution characteristics, improving statistical discrimination without requiring complex real-time processing during decoding.

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS9076436B2Apparatus and method for applying pitch features in automatic speech recognition
Publication Date: 2015.07.07 TOSHIBA DIGITAL SOLUTIONS CORP
  • US9076436B2 patent drawing
  • US9076436B2 patent drawing
  • US9076436B2 patent drawing

AI summary

According to one embodiment, an apparatus for applying pitch features in automatic speech recognition is provided. The apparatus includes a distribution evaluation module, normalization module, and random value adjusting module. The distribution evaluation module evaluates the global distribution of pitch features of voiced frames in speech signals, and the global distribution of random values for unvoiced frames in speech signals. The normalization module normalizes the global distribution of random values for unvoiced frames based on the global distribution of pitch features of voiced frames. The random value adjusting module adjusts random values for unvoiced frames based on the normalized global distribution, so that the adjusted random values can be assigned to unvoiced frames in speech signals as pitch features of the unvoiced frames.