Dynamic Arc Weights in Speech Recognition FST Models
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional automatic speech recognition (ASR) systems face challenges in achieving accurate results when users utter uncommon words or when the context changes during a multi-turn dialog, as large general language models may not provide the most accurate results and smaller context-specific models incur overhead and poor performance if incorrectly loaded.
Innovation Solution
Implementing a dynamic finite state transducer (FST) language model that adjusts arc weights based on contextual factors such as user identity, time, and previous interactions, allowing a shared large model to be customized for the current context without the need for training or loading a specific model, using dynamic weight maps to replace arc weights for improved decoding accuracy.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If a large general language model is used, then coverage for most words is provided, but recognition accuracy deteriorates when users utter uncommon words or when context changes
Solution Approach 1:
The patent implements dynamic arc weights in the FST language model that can be adjusted at runtime based on contextual factors such as user identity, time, and previous interactions. This allows the system to transition from static general-purpose weights to dynamic context-specific weights without changing the underlying model structure, resolving the contradiction between general coverage and context-specific accuracy.
Solution Approach 2:
The system changes the parameters (arc weights) of the language model based on contextual conditions. By maintaining a base FST model with general weights and applying dynamic weight adjustments from context-specific data structures, the system adapts its behavior to different contexts while preserving the underlying model architecture.
2Measurement precision
If smaller context-specific models are used, then recognition accuracy improves for specific contexts, but system overhead increases and performance deteriorates if incorrect models are loaded
Solution Approach 1:
The patent creates a universal base FST language model that can serve multiple contexts by dynamically adjusting its arc weights. Instead of maintaining separate models for different contexts, the single model adapts to various contexts through weight modifications, reducing system overhead while maintaining context-specific accuracy.
Solution Approach 2:
The system creates lightweight copies of weight data structures (context-specific FST weight data structures) rather than copying entire language models. This allows rapid switching between contexts by replacing weight data structures while keeping the base model intact, minimizing overhead.
3Measurement precision
If context-specific models are maintained, then accuracy for uncommon words improves, but model loading time and computational overhead increase
Solution Approach 1:
The patent segments the language model into a base FST structure and separate context-specific weight data structures. This segmentation allows the system to load only the lightweight weight data structures for a given context rather than loading entire context-specific models, significantly reducing loading time while maintaining accuracy.
Solution Approach 2:
The system pre-processes and stores context-specific weight data structures in advance, organized by contextual factors. When a context is encountered, the corresponding pre-computed weights are quickly retrieved and applied to the base model, eliminating the need for on-the-fly model training or loading.
4Productivity
If a shared large model is used, then system resources are efficiently utilized, but the model cannot be customized for current context without training
Solution Approach 1:
The patent introduces dynamic arc weights to the shared large FST model, enabling it to adapt to different contexts at runtime. The base model remains shared and efficient, while context-specific weight data structures provide customization without requiring model training or reloading.
Solution Approach 2:
The system introduces context-specific weight data structures as an intermediary layer between the shared base model and the decoding process. These weights act as a mediator that translates general model capabilities into context-specific behavior without modifying the base model itself.
Data Source
AI summary
Features are disclosed for performing speech recognition on utterances using dynamic weights with speech recognition models. An automatic speech recognition system may use a general speech recognition model, such a large finite state transducer-based language model, to generate speech recognition results for various utterances. The general speech recognition model may include sub-models or other portions that are customized for particular tasks, such as speech recognition on utterances regarding particular topics. Individual weights within the general speech recognition model can be dynamically replaced based on the context in which an utterance is made or received, thereby providing a further degree of customization without requiring additional speech recognition models to generated, maintained, or loaded.


