Speech Recognition Confidence Scoring via Phoneme Acoustic Maps

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current speech recognition systems are hardware dependent, speaker dependent, and lack flexibility in creating and modifying concepts and grammars, requiring time-consuming modifications and recompilation, and do not allow dynamic creation of concepts with multiple phrases or simultaneous decodes using different grammar and voice samples.

Innovation Solution

A speech recognition system API that is hardware independent, speaker independent, allows dynamic creation and modification of concepts and grammars, uses flexible phrase formats, and includes a voice channel model or grammar set model for multiple simultaneous decodes, enabling the determination of a confidence score for speech recognition engine decoding through a method involving phoneme acoustic score maps and weighted averages.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Stability of the object's composition

If speech recognition systems use hardware components or fixed software modules, then system stability is improved, but adaptability and ease of modification deteriorate

Engineering Contradiction:
Improvesystem stabilityVSAvoidadaptability
Core Design Contradiction:
Stability of the object's compositionVSAdaptability or versatility

Solution Approach 1:

The system dynamically creates and modifies concepts and grammars at runtime without recompilation. The speech recognition engine loads grammars and voice samples into memory, allowing concepts to be added, removed, or modified dynamically. This enables the system to adapt to different hardware platforms and user needs while maintaining stable core functionality.

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

The API is designed to be hardware independent and speaker independent, making it universally applicable across different platforms. The system can handle multiple voice samples and grammar sets simultaneously, providing multi-functional capability that works across diverse hardware configurations without requiring platform-specific modifications.

Inventive Principle:
Principle #6Universality (Multi-functionality)

2Reliability

If speech recognition systems require recompilation and reloading for modifications, then system reliability is improved, but productivity and ease of operation deteriorate

Engineering Contradiction:
Improvesystem reliabilityVSAvoidmodification speed
Core Design Contradiction:
ReliabilityVSProductivity

Solution Approach 1:

The system allows dynamic creation and modification of concepts and grammars during runtime. Changes to the speech recognition configuration are loaded into memory immediately without requiring system recompilation or reloading, enabling rapid iteration and modification while maintaining system reliability through controlled memory management.

Inventive Principle:
Principle #15Dynamics

3Measurement precision

If speech recognition systems use fixed grammar and voice sample combinations, then measurement precision is improved, but adaptability deteriorates

Engineering Contradiction:
Improverecognition accuracyVSAvoidflexibility
Core Design Contradiction:
Measurement precisionVSAdaptability or versatility

Solution Approach 1:

The system dynamically selects and combines different grammar sets and voice samples based on the specific recognition task. Multiple grammar and voice sample combinations can be loaded simultaneously, allowing the system to optimize recognition accuracy for different scenarios while maintaining flexibility to switch between configurations as needed.

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

The voice channel model and grammar set model enable the system to handle multiple simultaneous decodes using different combinations of grammar and voice samples. This universal approach allows the same system to achieve high precision for various speech patterns, accents, and domains without sacrificing adaptability.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Data Source

PatentUS7324940B1Speech recognition concept confidence measurement
Publication Date: 2008.01.29 AI SOFTWARE LLC
  • US7324940B1 patent drawing
  • US7324940B1 patent drawing
  • US7324940B1 patent drawing

AI summary

Systems and methods for determining a confidence score associated with a decoding output of a speech recognition engine. In one embodiment, a method of determining the confidence score comprises arranging time frame and acoustic score data into an array, determining a phoneme sequence in the array that yields the highest sum of acoustic scores under certain constraints, e.g., minimum number of time frames and order of phonemes in a phoneme string. A relative score is derived by applying a functional relationship between the acoustic score and different sums comprising acoustic scores from the array. The confidence score, in some embodiments, depends at least in part on the relative score and a measure of ambiguity associated with similar sounding phrases being included in different concepts of a specified grammar.