Phonetic Distance Measurement Using Speech Recognition Error Rates

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current methods for quantifying phonetic distance in speech recognition engines are limited by their reliance on physiological mechanisms and fail to accurately account for unique characteristics of specific speech recognition engines and speakers, particularly in handling insertion and deletion errors.

Innovation Solution

A system and method that compares recognized speech files with reference files to determine error occurrences and rates, calculating phonetic distances based on these errors, and normalizes them using a mapping function with coefficients to create a phonetic distance matrix that can be used for grammar selection and language training.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If phonetic distance is estimated based on physiological mechanisms, then the measurement can be obtained using conventional methods, but the measurement fails to account for unique characteristics of specific speech recognition engines and speakers

Engineering Contradiction:
Improveadaptability to specific speech recognition engines and speakersVSAvoidphonetic distance measurement accuracy
Core Design Contradiction:
Adaptability or versatilityVSMeasurement precision

Solution Approach 1:

The patent changes the measurement parameters from physiological-based estimates to actual error rates observed in speech recognition operations. By measuring substitution, insertion, and deletion errors specific to each engine-speaker combination, the system adapts phonetic distance measurements to actual performance characteristics rather than relying on conventional physiological models.

Inventive Principle:
Principle #35Parameter changes

Solution Approach 2:

The system uses feedback from actual speech recognition error occurrences to refine phonetic distance measurements. By continuously monitoring substitution, insertion, and deletion errors in recognized speech files and comparing them against reference files, the system adjusts phonetic distance values to reflect real-world performance of specific engines and speakers.

Inventive Principle:
Principle #23Feedback

2Reliability

If phonetic distance is measured using conventional physiological-based methods, then the process is simple, but it cannot handle insertion and deletion errors effectively

Engineering Contradiction:
Improveerror handling capabilityVSAvoidmeasurement system complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The patent segments the error analysis into three distinct categories: substitution errors, insertion errors, and deletion errors. Each error type is measured and weighted separately to calculate phonetic distance, allowing the system to handle different error mechanisms independently rather than using a single conventional metric.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The system introduces an intermediary comparison module that analyzes recognized speech files against reference files to identify and categorize errors. This intermediary process bridges the gap between raw speech recognition output and phonetic distance measurement, enabling reliable handling of insertion and deletion errors through systematic error detection and classification.

Inventive Principle:
Principle #24Intermediary (Mediator)

3Measurement precision

If phonetic distance measurements are normalized using mapping functions with coefficients, then the total separation between measured distances and existing matrices is minimized, but the normalization process adds computational complexity

Engineering Contradiction:
Improvephonetic distance normalization accuracyVSAvoidnormalization process complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent applies parameter changes through mapping functions that transform raw error rate measurements into normalized phonetic distance values. By using three normalization coefficients (a, b, c) in the mapping function d(i,j) = a + b/(e(i,j) - c), the system adjusts the scale and distribution of phonetic distances to minimize total separation from existing phonetic distance matrices while maintaining measurement accuracy.

Inventive Principle:
Principle #35Parameter changes

Data Source

PatentUS9659559B2Phonetic distance measurement system and related methods
Publication Date: 2017.05.23 ADACEL SYST
  • US9659559B2 patent drawing
  • US9659559B2 patent drawing
  • US9659559B2 patent drawing

AI summary

Phonetic distances are empirically measured as a function of speech recognition engine recognition error rates. The error rates are determined by comparing a recognized speech file with a reference file. The phonetic distances can be normalized to earlier measurements. The phonetic distances/error rates can also be used to improve speech recognition engine grammar selection, as an aid in language training and evaluation, and in other applications.