Dynamic Language Model Switching in Speech Recognition

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional speech recognition systems are inflexible, requiring users to terminate existing sessions and log back in to access different language models and resources, which is tedious and time-consuming, especially when switching between diverse tasks that require tailored language models.

Innovation Solution

A distributed speech recognition system that allows dynamic switching of language models and resources using voice commands, enabling seamless resource management without the need to terminate existing logons, by utilizing a text recognizer to identify triggers that initiate the loading of new language models and resources.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If conventional speech recognition systems use a single language model per user account, then the system is simple to manage, but the user cannot switch between different language models without terminating the session and logging back in, which is time-consuming

Engineering Contradiction:
Improveability to switch between language modelsVSAvoidtime to switch between language models
Core Design Contradiction:
Adaptability or versatilityVSLoss of time

Solution Approach 1:

The system enables dynamic switching of language models during an active speech recognition session. Instead of being locked to a single language model per user account, the system allows real-time model switching through voice commands, transforming the static configuration into a dynamic, adaptable system that responds to user needs without session termination.

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

The system uses voice-activated triggers that automatically initiate language model switching without requiring manual intervention. When a user speaks a trigger phrase, the system autonomously identifies the requested language model and switches to it, eliminating the need for users to manually navigate through menus or re-login to change models.

Inventive Principle:
Principle #25Self-service

2Adaptability or versatility

If the system allows multiple language models per user account, then the adaptability improves, but the complexity of managing and distinguishing between different language models increases

Engineering Contradiction:
Improveaccess to multiple language modelsVSAvoidcomplexity of language model switching mechanism
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The system introduces voice commands as an intermediary layer between the user and the language model switching mechanism. Instead of requiring users to directly interact with complex model management interfaces, the voice trigger acts as a simple mediator that translates spoken phrases into automated model switching actions, reducing the perceived complexity.

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The system replaces manual mechanical interactions (clicking, navigating menus, re-login procedures) with acoustic field-based voice commands. This substitution transforms the complex mechanical process of model switching into a simple acoustic interaction, making the system easier to use despite supporting multiple language models.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

3Measurement precision

If the system distinguishes between dictation audio and command audio, then the accuracy of speech recognition improves, but the difficulty of detecting and measuring command triggers increases

Engineering Contradiction:
Improveaccuracy of speech recognitionVSAvoiddifficulty of detecting command audio
Core Design Contradiction:
Measurement precisionVSDifficulty of detecting and measuring

Solution Approach 1:

The system segments the audio stream into distinct phases: dictation phases and command phases. By detecting trigger phrases and using them as delimiters, the system divides the continuous audio into manageable segments, allowing accurate recognition of both dictation content and command intentions without confusion between the two types of audio.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The system implements feedback by monitoring the recognized text for trigger phrases and automatically initiating language model switches. This feedback loop allows the system to distinguish between regular dictation and command audio by detecting specific patterns in the recognized text, thereby maintaining high accuracy while simplifying the detection of command triggers.

Inventive Principle:
Principle #23Feedback

Data Source

PatentUS9812130B1Apparatus and methods for dynamically changing a language model based on recognized text
Publication Date: 2017.11.07 NVOQ INC
  • US9812130B1 patent drawing
  • US9812130B1 patent drawing
  • US9812130B1 patent drawing

AI summary

The technology of the present application provides a method and apparatus to manage speech resources. The method includes using a text recognizer to detect a change in a speech application that requires the use of different resources. On detection of the change, the method loads the different resources without the user needing to exit the currently executing speech application.