Predictive Speech Model Fetching for Embedded Devices
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Mobile devices with limited storage capacity struggle to store a wide range of speech processing models required for interactive speech technologies, leading to storage constraints and increased reliance on network-based servers, which result in high latency and inefficiencies.
Innovation Solution
Implementing a system that automatically manages and fetches speech processing models based on predictive factors such as geographic location, app content, usage patterns, and user preferences, allowing local devices to determine which models are needed and prioritize storage, thereby reducing the need for network-based processing.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Loss of time
If speech processing models are stored on local mobile devices, then response latency is reduced and network dependency is decreased, but storage space is consumed and device storage capacity is limited
Solution Approach 1:
The system performs preliminary actions by predicting which speech processing models will be needed based on various factors (geographic location, app content, usage patterns) and fetching them to local storage in advance, before they are actually required. This ensures low-latency local processing while only consuming storage space for models that are likely to be used soon.
Solution Approach 2:
The system dynamically manages the local model repository by continuously monitoring predictive factors and adjusting which models are stored locally. Models are fetched when prediction metrics indicate they will be needed, and removed when they are no longer expected to be used, creating a dynamic balance between storage consumption and response latency.
2Adaptability or versatility
If a wide range of speech processing models are stored locally, then model availability and functionality are improved, but storage competition with other device content increases
Solution Approach 1:
Instead of storing all possible speech processing models locally (excessive action), the system stores only a partial set of models that are predicted to be needed based on current context and usage patterns. This partial storage approach provides sufficient model availability for actual use cases while avoiding storage competition with other device content.
Solution Approach 2:
The system changes the parameter of model selection from static (pre-defined set of models) to dynamic (context-dependent model selection). By using predictive factors such as geographic location, app content, and usage patterns, the system adapts which models are stored locally, optimizing the balance between model availability and storage space.
3Productivity
If speech processing models are fetched and stored locally, then embedded speech technology performance is improved, but device storage capacity is limited
Solution Approach 1:
The system fetches and stores speech processing models locally in advance based on predictive analysis, enabling high-performance embedded speech processing without requiring all models to be permanently stored. Only the necessary subset of models occupies storage capacity, while performance is improved through local processing of predicted use cases.
Data Source
AI summary
Disclosed herein are systems, methods, and computer-readable storage devices for fetching speech processing models based on context changes in advance of speech requests using the speech processing models. An example local device configured to practice the method, having a local speech processor, and having access to remote speech models, detects a change in context. The change in context can be based on geographical location, language translation, speech in a different language, user language settings, installing or removing an app, and so forth. The local device can determine a speech processing model that is likely to be needed based on the change in context, and that is not stored on the local device. Independently of an explicit request to process speech, the local device can retrieve, from a remote server, the speech processing model for use on the mobile device.


