ASR Compute Graph Batching for Dynamic AI Model Swapping

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Modern ASR systems face challenges in efficiently handling a large number of diverse AI models due to inflexibility in hardware configuration, leading to inefficiencies and increased latency in processing thousands or millions of custom AI models.

Innovation Solution

The system employs dynamic model swapping and compute graph generation to efficiently load and unload AI models on hardware modules, utilizing parallel processing capabilities to handle multiple requests simultaneously and prioritize based on priority data.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If traditional ASR systems use a handful of AI models with fixed hardware configuration, then hardware management is simple, but the system cannot handle the exponential increase in customized AI models efficiently

Engineering Contradiction:
Improveability to handle diverse AI modelsVSAvoidhardware configuration complexity
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The patent implements dynamic model swapping mechanisms that allow AI models to be loaded and unloaded from hardware modules on-demand. The system dynamically determines which models to swap based on request patterns, maintaining adaptability to diverse AI models while managing hardware complexity through automated model lifecycle management

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

The patent creates a universal hardware module architecture that can execute multiple different AI models across various domains (finance, healthcare, legal, etc.). The standardized hardware modules can be configured with different models through a common interface, enabling one hardware system to serve multiple specialized functions

Inventive Principle:
Principle #6Universality (Multi-functionality)

2Adaptability or versatility

If the system loads multiple customized AI models on hardware modules, then it can service diverse ASR requests, but the reconfiguration time increases

Engineering Contradiction:
Improvemodel diversity supportVSAvoidreconfiguration time
Core Design Contradiction:
Adaptability or versatilityVSLoss of time

Solution Approach 1:

The system performs preliminary actions by pre-loading AI models into hardware modules before they are needed for processing requests. The model swapping mechanism is prepared in advance, allowing transitions between models to occur rapidly when requests arrive, thereby reducing reconfiguration time while maintaining support for diverse models

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent ensures continuous processing by maintaining multiple AI models ready in hardware modules simultaneously. While one model is being used, other models remain loaded and ready for immediate execution, eliminating idle reconfiguration time and maintaining continuous useful action across diverse ASR request types

Inventive Principle:
Principle #20Continuity of useful action

3Productivity

If the system processes a large volume of ASR requests with multiple AI models, then service coverage increases, but hardware overhead and latency increase

Engineering Contradiction:
Improverequest processing throughputVSAvoidhardware resource consumption
Core Design Contradiction:
ProductivityVSLoss of energy

Solution Approach 1:

The patent merges multiple AI model processing capabilities into shared hardware modules that can execute different models sequentially or in parallel. By consolidating model execution resources rather than dedicating separate hardware to each model, the system reduces overall hardware overhead and energy consumption while maintaining the ability to process large volumes of diverse ASR requests

Inventive Principle:
Principle #5Merging (Combining)

Data Source

PatentUS20250372096A1Hardware efficient automatic speech recognition
Publication Date: 2025.12.04 DEEPGRAM INC
  • US20250372096A1 patent drawing
  • US20250372096A1 patent drawing
  • US20250372096A1 patent drawing

AI summary

Modern automatic speech recognition (ASR) systems can utilize artificial intelligence (AI) models to service ASR requests. The number and scale of AI models used in a modern ASR system can be substantial. The process of configuring and reconfiguring hardware to execute various AI models corresponding to a substantial number of ASR requests can be time consuming and inefficient. Among other features, the described technology utilizes batching of ASR requests, splitting of the ASR requests, and/or parallel processing to efficiently use hardware tasked with executing AI models corresponding to ASR requests. In one embodiment, the compute graphs of ASR tasks are used to batch the ASR requests. The corresponding AI models of each batch can be loaded into hardware, and batches can be processed in parallel. In some embodiments, the ASR requests are split, batched, and processed in parallel.