Dynamic Section Grammar Loading for Speech Recognition
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current speech recognition systems lack the ability to dynamically switch between domains within a single document, especially in unstructured environments like traditional telephony dictation, and are limited by the need for well-defined data fields and vocabulary constraints.
Innovation Solution
A system and method for loading and unloading dynamic grammars and language models that identify and adapt to section-based structures in documents, allowing for seamless domain switching without requiring predefined data fields or vocabulary constraints, using dynamic section identification and incremental training based on user dictation patterns.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If domain-specific language models are used to improve speech recognition accuracy in specific domains, then accuracy is improved, but the system cannot switch between different domains within a single document
Solution Approach 1:
The system dynamically loads and unloads language models and grammars based on the current section being processed. Instead of using a static domain-specific model throughout, the system adapts by switching between different language models corresponding to different sections (e.g., narrative section model, structured section model), enabling both accuracy and domain switching capability
Solution Approach 2:
The document is divided into sections with different structural characteristics (narrative vs. structured). Each section type has its own specialized language model. The system segments the processing task by identifying section boundaries and applying the appropriate language model to each section, resolving the contradiction between specialized accuracy and overall adaptability
2Productivity
If structured report organization with well-defined data fields is used to improve recognition efficiency, then efficiency is improved, but the system cannot handle unstructured environments like traditional telephony dictation
Solution Approach 1:
The system is designed to handle both structured and unstructured documents through a universal framework. It uses section identification algorithms that work regardless of document structure, and can switch between narrative language models and structured section models. This multi-functionality allows the same system to efficiently process both well-defined data fields and unstructured telephony dictation
3Measurement precision
If vocabulary constraints are applied to improve recognition accuracy, then accuracy is improved, but the system becomes limited by predefined vocabulary
Solution Approach 1:
The vocabulary constraints are applied dynamically rather than statically. The system identifies the current section type and applies vocabulary constraints appropriate to that section. This allows the system to maintain high accuracy within each section type while remaining flexible across different section types and domains
Data Source
AI summary
The invention involves the loading and unloading of dynamic section grammars and language models in a speech recognition system. The values of the sections of the structured document are either determined in advance from a collection of documents of the same domain, document type, and speaker; or collected incrementally from documents of the same domain, document type, and speaker; or added incrementally to an already existing set of values. Speech recognition in the context of the given field is constrained to the contents of these dynamic values. If speech recognition fails or produces a poor match within this grammar or section language model, speech recognition against a larger, more general vocabulary that is not constrained to the given section is performed.


