Speech Recognition Confidence Scoring via Phoneme Acoustic Maps
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current speech recognition systems are hardware dependent, speaker dependent, and lack flexibility in creating and modifying concepts and grammars, requiring time-consuming modifications and recompilation, and do not allow dynamic creation of concepts with multiple phrases or simultaneous decodes using different grammar and voice samples.
Innovation Solution
A speech recognition system API that is hardware independent, speaker independent, allows dynamic creation and modification of concepts and grammars, uses flexible phrase formats, and includes a voice channel model or grammar set model for multiple simultaneous decodes, enabling the determination of a confidence score for speech recognition engine decoding through a method involving phoneme acoustic score maps and weighted averages.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Stability of the object's composition
If speech recognition systems use hardware components or fixed software modules, then system stability is improved, but adaptability and ease of modification deteriorate
Solution Approach 1:
The system dynamically creates and modifies concepts and grammars at runtime without recompilation. The speech recognition engine loads grammars and voice samples into memory, allowing concepts to be added, removed, or modified dynamically. This enables the system to adapt to different hardware platforms and user needs while maintaining stable core functionality.
Solution Approach 2:
The API is designed to be hardware independent and speaker independent, making it universally applicable across different platforms. The system can handle multiple voice samples and grammar sets simultaneously, providing multi-functional capability that works across diverse hardware configurations without requiring platform-specific modifications.
2Reliability
If speech recognition systems require recompilation and reloading for modifications, then system reliability is improved, but productivity and ease of operation deteriorate
Solution Approach 1:
The system allows dynamic creation and modification of concepts and grammars during runtime. Changes to the speech recognition configuration are loaded into memory immediately without requiring system recompilation or reloading, enabling rapid iteration and modification while maintaining system reliability through controlled memory management.
3Measurement precision
If speech recognition systems use fixed grammar and voice sample combinations, then measurement precision is improved, but adaptability deteriorates
Solution Approach 1:
The system dynamically selects and combines different grammar sets and voice samples based on the specific recognition task. Multiple grammar and voice sample combinations can be loaded simultaneously, allowing the system to optimize recognition accuracy for different scenarios while maintaining flexibility to switch between configurations as needed.
Solution Approach 2:
The voice channel model and grammar set model enable the system to handle multiple simultaneous decodes using different combinations of grammar and voice samples. This universal approach allows the same system to achieve high precision for various speech patterns, accents, and domains without sacrificing adaptability.
Data Source
AI summary
Systems and methods for determining a confidence score associated with a decoding output of a speech recognition engine. In one embodiment, a method of determining the confidence score comprises arranging time frame and acoustic score data into an array, determining a phoneme sequence in the array that yields the highest sum of acoustic scores under certain constraints, e.g., minimum number of time frames and order of phonemes in a phoneme string. A relative score is derived by applying a functional relationship between the acoustic score and different sums comprising acoustic scores from the array. The confidence score, in some embodiments, depends at least in part on the relative score and a measure of ambiguity associated with similar sounding phrases being included in different concepts of a specified grammar.


