Dynamic Vocabulary Loading for Speech Recognition Accuracy
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current speech recognition systems face challenges in accurately transcribing specific terms and context-specific inputs due to the use of general vocabularies that result in high error rates, as they fail to leverage repeated dictations in specific form fields or by specific users, lacking context-specific vocabularies and modifications.
Innovation Solution
The method involves loading specified vocabularies associated with specific contexts or users into computer storage, evaluating user voice input against these criteria, and selecting appropriate data values for form fields, with a fallback to a base vocabulary if the input does not meet the evaluation criterion, using a context identification module to dynamically select and prioritize sub-databases for improved accuracy.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If general vocabularies are used in speech recognition systems, then the system can handle a wide range of inputs, but transcription accuracy decreases for specific terms and context-specific inputs
Solution Approach 1:
The patent segments the vocabulary into multiple hierarchical levels: a base vocabulary containing general terms and multiple specified vocabularies containing context-specific terms. The system dynamically selects and combines appropriate vocabularies based on the recognition context, allowing it to handle both general and specific inputs effectively while maintaining high transcription accuracy for context-specific terms.
2Measurement precision
If context-specific vocabularies are implemented, then transcription accuracy improves, but system complexity increases
Solution Approach 1:
The patent implements dynamic vocabulary selection where the system automatically adjusts which vocabularies to use based on the recognition context. The context identification module dynamically determines the appropriate context and selects corresponding specified vocabularies, eliminating the need for manual configuration and reducing operational complexity despite having multiple vocabularies.
Solution Approach 2:
The patent introduces a context identification module as an intermediary between the speech input and the vocabulary selection process. This module analyzes the input speech and determines the appropriate context, then selects the corresponding specified vocabulary from multiple available vocabularies. This intermediary simplifies the overall system architecture by centralizing the decision-making logic for vocabulary selection.
3Measurement precision
If multiple specified vocabularies are loaded into memory, then transcription accuracy for specific contexts improves, but memory usage increases
Solution Approach 1:
The patent performs preliminary organization of vocabularies into hierarchical groups (base vocabulary and multiple specified vocabularies associated with different contexts). This pre-organization allows the system to load only the necessary specified vocabularies into memory based on the current recognition context, rather than loading all possible vocabularies simultaneously, thus optimizing memory usage while maintaining high transcription accuracy.
Data Source
AI summary
The present invention involves the dynamic loading and unloading of relatively small text-string vocabularies within a speech recognition system. In one embodiment, sub-databases of high likelihood text strings are created and prioritized such that those text strings are made available within definable portions of computer-transcribed dictations as a first-pass vocabulary for text matches. Failing a match within the first-pass vocabulary, the voice recognition software attempts to match the speech input to text strings within a more general vocabulary. In another embodiment, the first-pass text string vocabularies are organized and prioritized and loaded in relation to specific fields within an electronic form, specific users of the system and/or other general context-based, interrelationships of the data that provide a higher probability of text string matches then those otherwise provided by commercially available speech recognition systems and their general vocabulary databases.


