Multipass ASR Demultiplexer for Low-Memory Speech Recognition
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing automatic speech recognition (ASR) systems face significant error rates due to large vocabularies, requiring substantial memory and processing resources, which is a challenge for portable applications needing high accuracy, especially in urgent conditions like emergency calls or healthcare assessments where memory footprint is limited.
Innovation Solution
A multipass processing system using a reduced grammar-based ASR that processes audio files through a stateless, real-time system with a demultiplexer selecting the best matching grammar-based ASR based on confidence scores, allowing for accurate and scalable speech recognition without monitoring the state of ASRs, enabling processing of multiple words, phrases, or sentences without prior completion.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If a large vocabulary ASR system is used to improve recognition coverage, then the vocabulary coverage is improved, but the memory footprint and processing resources increase substantially
Solution Approach 1:
The patent divides the ASR system into multiple passes with specialized grammars for different domains (e.g., healthcare, emergency services). Each pass uses a focused vocabulary appropriate to its domain, avoiding the need to load all vocabularies simultaneously. This segmentation allows the system to maintain high adaptability within each domain while keeping overall memory footprint manageable.
Solution Approach 2:
The system dynamically selects and switches between different grammar-based ASRs based on the spoken utterance content. Rather than using a static large-vocabulary system, the ASR adapts its vocabulary scope in real-time by selecting the appropriate specialized grammar for the current context, optimizing the balance between coverage and memory usage.
2Adaptability or versatility
If a large vocabulary ASR system is used to improve recognition coverage, then the vocabulary coverage is improved, but the processing resources increase substantially
Solution Approach 1:
Processing resources are segmented across multiple specialized ASR engines, each optimized for a specific domain. This avoids the computational overhead of a single large-vocabulary system while maintaining comprehensive coverage through coordinated use of multiple smaller engines.
Solution Approach 2:
The system performs preliminary classification to determine which specialized grammar is appropriate for the current utterance before engaging the full ASR processing. This preliminary action reduces the processing burden by directing only relevant utterances to the appropriate specialized engine, avoiding unnecessary computation.
3Measurement precision
If traditional ASR systems are used to improve recognition accuracy, then the vocabulary size is increased, but the error rates increase due to frequent mismatches
Solution Approach 1:
Each specialized grammar is optimized for its specific domain with locally appropriate vocabulary and patterns. This local optimization ensures high recognition accuracy and reliability within each domain, avoiding the mismatches that occur when a general large-vocabulary system attempts to handle specialized terminology.
4Measurement precision
If a stateful ASR system is used to monitor ASR state for accuracy, then the recognition accuracy is improved, but the memory footprint increases
Solution Approach 1:
The demultiplexer uses confidence scores generated by each grammar-based ASR to automatically determine the best match and route the utterance accordingly. This self-service mechanism eliminates the need for external state monitoring while maintaining high recognition accuracy through confidence-based selection.
Data Source
Figure 1
Figure 2
Figure 3
AI summary
A multipass processing system includes a first grammar-based speech recognition system that compares a spoken utterance to a sub-grammar. The sub-grammar includes keywords or key phrases from active grammars that each uniquely identifies one of many application engines. The first grammar-based speech recognition system generates a first grammar-based speech recognition result and a first grammar-based confidence score. A demultiplexer receives the spoken utterance through an input. The demultiplexer transmits the spoken utterance to one of many other grammar-based speech recognition systems based on the first grammar-based speech recognition-result.