Speech Frame Identification via Weighted Suitability Measures
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional speech recognition systems face challenges in accurately identifying speech frames in both quiet and noisy environments due to limitations in noise interference mitigation and inter-speaker variability, leading to degraded recognition accuracy.
Innovation Solution
A computer-implemented system that computes a weighted suitability measure for speech frames using spectral flatness measure, energy normalized variance, entropy, signal-to-noise ratio, and similarity measure to identify significant speech frames, enabling better training and scoring models for speech recognition.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If conventional speech recognition systems use MFCC features, then recognition accuracy is improved in quiet environments, but recognition accuracy deteriorates in noisy environments
Solution Approach 1:
The patent changes the feature representation parameters from conventional MFCC to a hybrid approach combining spectral flatness measure (SFM), energy normalized variance (ENV), and entropy. These alternative parameters are more robust to noise, allowing the system to maintain high recognition accuracy across both quiet and noisy environments by using parameters that capture different aspects of speech signals less susceptible to noise contamination.
Solution Approach 2:
The patent creates a composite feature representation by combining multiple spectral characteristics (SFM, ENV, and entropy) into a unified feature set. This composite approach leverages the strengths of each individual measure: SFM for noise robustness, ENV for voicing detection, and entropy for speech/non-speech discrimination, resulting in a feature representation that adapts to various environmental conditions.
2Object-affected harmful factors
If noise-robust speech recognition software is used, then noise interference mitigation is improved, but recognition accuracy deteriorates in quiet conditions
Solution Approach 1:
The patent employs multiple spectral parameters (SFM, ENV, entropy) that have different sensitivities to noise. By monitoring these parameters simultaneously and selecting or weighting them based on environmental conditions, the system achieves noise robustness when needed while maintaining high accuracy in quiet conditions, avoiding the trade-off inherent in single-approach noise-robust systems.
3Speed
If speech frames are processed using conventional methods, then processing speed is maintained, but separation of inter-speaker variability from channel distortion deteriorates
Solution Approach 1:
The patent segments the speech signal into frames and processes each frame independently using the spectral parameters. This frame-based segmentation allows for efficient processing while the use of multiple parameters (SFM, ENV, entropy) within each frame provides sufficient information to separate inter-speaker variability from channel distortion without requiring complex cross-frame analysis, thus maintaining processing speed.
Data Source
AI summary
The present disclosure envisages a computer implemented system for identifying significant speech frames within speech signals for facilitating speech recognition. The system receives an input speech signal having a plurality of feature vectors which is passed through a spectrum analyzer. The spectrum analyzer divides the input speech signal into a plurality of speech frames and computes a spectral magnitude of each of the speech frames. There is provided a suitability engine which is enabled to compute a suitability measure for each of the speech frames corresponding to spectral flatness measure (SFM), energy normalized variance (ENV), entropy, signal-to-noise ratio (SNR) and similarity measure. The suitability engine further computes a weighted suitability measure for each of the speech frames.


