Standardized Speech Recognition Infrastructure Model Selection

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current automatic speech recognition (ASR) systems are incompatible and lack access to improved recognition models, as each system uses its own proprietary models trained from human or machine transcriptions, limiting the sharing and reuse of speech recognition models across applications and environments.

Innovation Solution

The development of a standardized speech recognition infrastructure that allows for the selection and adaptation of speech recognition models, including supervised, unsupervised, and generic models, which can be retrieved and reused across devices, enabling compatibility and adaptation for improved speech recognition.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If each ASR system uses its own proprietary recognition models, then the system can achieve customized and accurate recognition for specific applications and speakers, but other applications cannot access or benefit from those models, and compatibility between different ASR systems is lost

Engineering Contradiction:
Improverecognition accuracyVSAvoidmodel sharing capability
Core Design Contradiction:
Measurement precisionVSAdaptability or versatility

Solution Approach 1:

The speech recognition model is segmented into multiple components: a base model containing general speech recognition capabilities and task-specific components or adapters that can be selectively applied. This allows the base model to be shared across different applications while maintaining customization through task-specific components, resolving the contradiction between model sharing and application-specific accuracy.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent creates a universal base speech recognition model that can serve multiple applications and domains. This base model is designed to be adaptable through parameter adjustments and task-specific component attachments, enabling one model to fulfill multiple functions across different ASR systems while maintaining compatibility and allowing model sharing.

Inventive Principle:
Principle #6Universality (Multi-functionality)

2Adaptability or versatility

If a standardized speech recognition infrastructure is implemented to enable model sharing and compatibility, then access to improved recognition models across applications is enabled, but the complexity of managing multiple model types (supervised, unsupervised, generic) increases

Engineering Contradiction:
Improvemodel compatibilityVSAvoidmodel management complexity
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The system performs preliminary actions by pre-processing and categorizing speech data into different types (supervised, unsupervised, generic) before model training. This preliminary organization of data and models simplifies subsequent model selection and management, reducing the complexity of handling multiple model types in the standardized infrastructure.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent introduces an intermediary layer or interface that manages the standardized speech recognition infrastructure. This intermediary handles model selection, retrieval, and adaptation automatically, shielding users from the complexity of managing multiple model types while enabling seamless model sharing and compatibility across applications.

Inventive Principle:
Principle #24Intermediary (Mediator)

3Productivity

If speech models are trained from large amounts of untranscribed speech to augment transcribed speech, then model training efficiency is improved, but the quality and accuracy of the trained models may be reduced compared to fully supervised training

Engineering Contradiction:
Improvemodel training efficiencyVSAvoidmodel accuracy
Core Design Contradiction:
ProductivityVSMeasurement precision

Solution Approach 1:

The patent applies different training qualities to different parts of the model training process. Supervised training with high-quality transcribed speech is used for critical components where accuracy is paramount, while unsupervised training on untranscribed speech is used for augmenting less critical aspects. This local differentiation of training quality maintains overall model accuracy while improving training efficiency through scalable use of untranscribed data.

Inventive Principle:
Principle #3Local quality

Data Source

PatentUS9336773B2System and method for standardized speech recognition infrastructure
Publication Date: 2016.05.10 INTERACTIONS LLC (US)
  • US9336773B2 patent drawing
  • US9336773B2 patent drawing
  • US9336773B2 patent drawing

AI summary

Disclosed herein are systems, methods, and computer-readable storage media for selecting a speech recognition model in a standardized speech recognition infrastructure. The system receives speech from a user, and if a user-specific supervised speech model associated with the user is available, retrieves the supervised speech model. If the user-specific supervised speech model is unavailable and if an unsupervised speech model is available, the system retrieves the unsupervised speech model. If the user-specific supervised speech model and the unsupervised speech model are unavailable, the system retrieves a generic speech model associated with the user. Next the system recognizes the received speech from the user with the retrieved model. In one embodiment, the system trains a speech recognition model in a standardized speech recognition infrastructure. In another embodiment, the system handshakes with a remote application in a standardized speech recognition infrastructure.