Domain-Specific ASR Engine Selection by Communication Interval
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing automatic speech recognition (ASR) systems in devices with limited processing resources face challenges due to computational complexities and power constraints, making it difficult to efficiently process speech recognition tasks.
Innovation Solution
A method that predicts the topic of communication between entities and selectively activates domain-specific ASR engines based on elapsed time since the previous communication, utilizing a smaller resource footprint than general ASR engines, allowing for efficient recognition of specific vocabulary and context.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If cloud-based ASR processing is used, then speech recognition accuracy is improved, but device power consumption and processing resource usage increase
Solution Approach 1:
The patent segments the ASR system into multiple domain-specific engines (e.g., music ASR, news ASR, weather ASR) that can be selectively activated based on the communication context. This allows the device to use only the necessary computational resources for the specific domain rather than running a full cloud-based system or a general-purpose ASR engine, thereby reducing power consumption while maintaining accuracy for the targeted domain.
Solution Approach 2:
The system dynamically selects and activates the appropriate ASR engine based on real-time analysis of communication metadata (such as time of day, duration, participants, and content). This dynamic adaptation enables the device to optimize power consumption by activating lightweight domain-specific engines for routine communications while reserving cloud-based processing for complex or unfamiliar domains.
2Use of energy by moving object
If domain-specific ASR engines with smaller resource footprint are used, then power consumption is reduced, but the ability to handle diverse communication topics decreases
Solution Approach 1:
The patent implements a universal ASR engine selection framework that can adapt to multiple communication domains through metadata analysis. The system maintains a library of domain-specific engines but uses a smart selection mechanism that can also fall back to general-purpose processing or activate multiple specialized engines in combination, ensuring versatility across diverse communication topics while maintaining low power consumption through selective activation.
Solution Approach 2:
The system performs preliminary analysis of communication metadata (time, duration, participants, content type) before activating the ASR engine. This preliminary action allows the system to predict the likely domain and pre-load or activate the appropriate domain-specific ASR engine, ensuring readiness for diverse topics while minimizing power consumption by avoiding activation of unnecessary engines.
3Speed
If real-time ASR processing is performed on-device, then response time is improved, but computational complexity and processing requirements increase
Solution Approach 1:
The patent applies local quality by implementing domain-specific ASR engines optimized for particular communication types (e.g., music, news, weather). Each engine is tailored to the specific vocabulary and patterns of its domain, reducing the computational complexity required compared to a general-purpose engine. The system selects the most appropriate localized engine based on metadata analysis, achieving fast response times with reduced computational burden.
Data Source
AI summary
A method and data processing device for detecting a communication between a first and second entity. The method includes identifying whether a previous communication between the first and second entity has been detected. In response to identifying that the previous communication between the first and second entity has been detected, the method determines an elapsed time since detection of the previous communication. The method predicts a topic of the communication, in part based on the determined elapsed time. The topic corresponds to a specific domain from among a plurality of available domains for automatic speech recognition (ASR) processing. The method triggers selection and activation of a first domain specific (DS) ASR engine from among a plurality of available DS ASR engines to utilize a smaller resource footprint than a general ASR engine and facilitate recognition of specific vocabulary and context, in part, based on the elapsed time since the previous communication.


