Dynamic Arc Weights in Speech Recognition FST Models

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional automatic speech recognition (ASR) systems face challenges in achieving accurate results when users utter uncommon words or when the context changes during a multi-turn dialog, as large general language models may not provide the most accurate results and smaller context-specific models incur overhead and poor performance if incorrectly loaded.

Innovation Solution

Implementing a dynamic finite state transducer (FST) language model that adjusts arc weights based on contextual factors such as user identity, time, and previous interactions, allowing a shared large model to be customized for the current context without the need for training or loading a specific model, using dynamic weight maps to replace arc weights for improved decoding accuracy.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If a large general language model is used, then coverage for most words is provided, but recognition accuracy deteriorates when users utter uncommon words or when context changes

Engineering Contradiction:
ImprovecoverageVSAvoidrecognition accuracy
Core Design Contradiction:
Adaptability or versatilityVSMeasurement precision

Solution Approach 1:

The patent implements dynamic arc weights in the FST language model that can be adjusted at runtime based on contextual factors such as user identity, time, and previous interactions. This allows the system to transition from static general-purpose weights to dynamic context-specific weights without changing the underlying model structure, resolving the contradiction between general coverage and context-specific accuracy.

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

The system changes the parameters (arc weights) of the language model based on contextual conditions. By maintaining a base FST model with general weights and applying dynamic weight adjustments from context-specific data structures, the system adapts its behavior to different contexts while preserving the underlying model architecture.

Inventive Principle:
Principle #35Parameter changes

2Measurement precision

If smaller context-specific models are used, then recognition accuracy improves for specific contexts, but system overhead increases and performance deteriorates if incorrect models are loaded

Engineering Contradiction:
Improverecognition accuracyVSAvoidsystem overhead
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent creates a universal base FST language model that can serve multiple contexts by dynamically adjusting its arc weights. Instead of maintaining separate models for different contexts, the single model adapts to various contexts through weight modifications, reducing system overhead while maintaining context-specific accuracy.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The system creates lightweight copies of weight data structures (context-specific FST weight data structures) rather than copying entire language models. This allows rapid switching between contexts by replacing weight data structures while keeping the base model intact, minimizing overhead.

Inventive Principle:
Principle #26Copying

3Measurement precision

If context-specific models are maintained, then accuracy for uncommon words improves, but model loading time and computational overhead increase

Engineering Contradiction:
Improverecognition accuracyVSAvoidmodel loading time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The patent segments the language model into a base FST structure and separate context-specific weight data structures. This segmentation allows the system to load only the lightweight weight data structures for a given context rather than loading entire context-specific models, significantly reducing loading time while maintaining accuracy.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The system pre-processes and stores context-specific weight data structures in advance, organized by contextual factors. When a context is encountered, the corresponding pre-computed weights are quickly retrieved and applied to the base model, eliminating the need for on-the-fly model training or loading.

Inventive Principle:
Principle #10Preliminary action

4Productivity

If a shared large model is used, then system resources are efficiently utilized, but the model cannot be customized for current context without training

Engineering Contradiction:
Improveresource efficiencyVSAvoidcontext customization
Core Design Contradiction:
ProductivityVSAdaptability or versatility

Solution Approach 1:

The patent introduces dynamic arc weights to the shared large FST model, enabling it to adapt to different contexts at runtime. The base model remains shared and efficient, while context-specific weight data structures provide customization without requiring model training or reloading.

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

The system introduces context-specific weight data structures as an intermediary layer between the shared base model and the decoding process. These weights act as a mediator that translates general model capabilities into context-specific behavior without modifying the base model itself.

Inventive Principle:
Principle #24Intermediary (Mediator)

Data Source

PatentUS10140981B1Dynamic arc weights in speech recognition models
Publication Date: 2018.11.27 AMAZON TECH INC
  • US10140981B1 patent drawing
  • US10140981B1 patent drawing
  • US10140981B1 patent drawing

AI summary

Features are disclosed for performing speech recognition on utterances using dynamic weights with speech recognition models. An automatic speech recognition system may use a general speech recognition model, such a large finite state transducer-based language model, to generate speech recognition results for various utterances. The general speech recognition model may include sub-models or other portions that are customized for particular tasks, such as speech recognition on utterances regarding particular topics. Individual weights within the general speech recognition model can be dynamically replaced based on the context in which an utterance is made or received, thereby providing a further degree of customization without requiring additional speech recognition models to generated, maintained, or loaded.