Speech Frame Identification via Weighted Suitability Measures

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional speech recognition systems face challenges in accurately identifying speech frames in both quiet and noisy environments due to limitations in noise interference mitigation and inter-speaker variability, leading to degraded recognition accuracy.

Innovation Solution

A computer-implemented system that computes a weighted suitability measure for speech frames using spectral flatness measure, energy normalized variance, entropy, signal-to-noise ratio, and similarity measure to identify significant speech frames, enabling better training and scoring models for speech recognition.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If conventional speech recognition systems use MFCC features, then recognition accuracy is improved in quiet environments, but recognition accuracy deteriorates in noisy environments

Engineering Contradiction:
Improvespeech recognition accuracyVSAvoidenvironmental adaptability
Core Design Contradiction:
Measurement precisionVSAdaptability or versatility

Solution Approach 1:

The patent changes the feature representation parameters from conventional MFCC to a hybrid approach combining spectral flatness measure (SFM), energy normalized variance (ENV), and entropy. These alternative parameters are more robust to noise, allowing the system to maintain high recognition accuracy across both quiet and noisy environments by using parameters that capture different aspects of speech signals less susceptible to noise contamination.

Inventive Principle:
Principle #35Parameter changes

Solution Approach 2:

The patent creates a composite feature representation by combining multiple spectral characteristics (SFM, ENV, and entropy) into a unified feature set. This composite approach leverages the strengths of each individual measure: SFM for noise robustness, ENV for voicing detection, and entropy for speech/non-speech discrimination, resulting in a feature representation that adapts to various environmental conditions.

Inventive Principle:
Principle #40Composite materials

2Object-affected harmful factors

If noise-robust speech recognition software is used, then noise interference mitigation is improved, but recognition accuracy deteriorates in quiet conditions

Engineering Contradiction:
Improvenoise interference mitigationVSAvoidrecognition accuracy
Core Design Contradiction:
Object-affected harmful factorsVSMeasurement precision

Solution Approach 1:

The patent employs multiple spectral parameters (SFM, ENV, entropy) that have different sensitivities to noise. By monitoring these parameters simultaneously and selecting or weighting them based on environmental conditions, the system achieves noise robustness when needed while maintaining high accuracy in quiet conditions, avoiding the trade-off inherent in single-approach noise-robust systems.

Inventive Principle:
Principle #35Parameter changes

3Speed

If speech frames are processed using conventional methods, then processing speed is maintained, but separation of inter-speaker variability from channel distortion deteriorates

Engineering Contradiction:
Improveprocessing speedVSAvoidinter-speaker variability separation
Core Design Contradiction:
SpeedVSMeasurement precision

Solution Approach 1:

The patent segments the speech signal into frames and processes each frame independently using the spectral parameters. This frame-based segmentation allows for efficient processing while the use of multiple parameters (SFM, ENV, entropy) within each frame provides sufficient information to separate inter-speaker variability from channel distortion without requiring complex cross-frame analysis, thus maintaining processing speed.

Inventive Principle:
Principle #1Segmentation

Data Source

PatentUS9659578B2Computer implemented system and method for identifying significant speech frames within speech signals
Publication Date: 2017.05.23 TATA CONSULTANCY SERVICES LTD
  • US9659578B2 patent drawing
  • US9659578B2 patent drawing
  • US9659578B2 patent drawing

AI summary

The present disclosure envisages a computer implemented system for identifying significant speech frames within speech signals for facilitating speech recognition. The system receives an input speech signal having a plurality of feature vectors which is passed through a spectrum analyzer. The spectrum analyzer divides the input speech signal into a plurality of speech frames and computes a spectral magnitude of each of the speech frames. There is provided a suitability engine which is enabled to compute a suitability measure for each of the speech frames corresponding to spectral flatness measure (SFM), energy normalized variance (ENV), entropy, signal-to-noise ratio (SNR) and similarity measure. The suitability engine further computes a weighted suitability measure for each of the speech frames.