Dynamic Speech Resource Allocation for Personalization
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing speech recognition applications rely on generic models that are not well-suited for individual speakers, leading to inefficient resource usage and high costs for personalization, as complete personalization for every speaker is often unnecessary and resource-intensive.
Innovation Solution
A system that recognizes speech using a set of allocated resources, records metrics such as confidence scores and dialog behavior, and modifies these resources dynamically to provide personalized speech recognition, allocating more resources as needed for speakers with unique dialects or accents, while minimizing unnecessary personalization for others.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If complete personalization is provided for every speaker, then speech recognition accuracy for individual speakers is improved, but resource consumption (processing power, storage, bandwidth) becomes prohibitive
Solution Approach 1:
The patent applies local quality by providing personalized speech recognition models only to specific speakers who need them, rather than uniformly to all speakers. The system identifies speakers with unusual dialects or accents and allocates personalization resources selectively to these local cases, while using generic models for the majority of speakers.
Solution Approach 2:
The system dynamically adjusts the level of personalization based on speaker characteristics and performance metrics. It monitors speech recognition confidence scores and dialog behavior in real-time, modifying resource allocation and model complexity adaptively to match actual speaker needs rather than using static personalization for all users.
2Measurement precision
If personal speech recognition models are generated for all users through training phase, then individual speaker recognition accuracy is improved, but processing time and complexity increase significantly
Solution Approach 1:
The patent implements partial personalization by providing complete training and personalization only when necessary, rather than for all users. The system performs full personalization selectively for speakers with unusual characteristics who benefit most from it, while accepting partial performance with generic models for other speakers, thus reducing overall training time and complexity.
Solution Approach 2:
The system allows speakers to implicitly request personalization through their speech patterns during normal interaction. By monitoring confidence scores and dialog behavior, the system automatically identifies when a speaker would benefit from personalization and initiates the training process only for those cases, rather than requiring upfront training for all users.
3Productivity
If generic speech model is used for all speakers, then resource efficiency is improved, but speech recognition accuracy for speakers with unusual dialects or accents deteriorates
Solution Approach 1:
The patent introduces an intermediary monitoring system that tracks speech recognition confidence scores and dialog behavior metrics. This intermediary layer detects when generic model performance degrades for specific speakers and triggers appropriate personalization actions, bridging the gap between efficient generic processing and accurate personalized recognition.
Solution Approach 2:
The system implements feedback loops that continuously monitor speech recognition performance metrics including confidence scores, requests for repeats, and negative responses to confirmations. This feedback information is used to dynamically adjust resource allocation and trigger personalization for speakers experiencing poor performance with the generic model, thereby maintaining both efficiency and accuracy.
Data Source
AI summary
Disclosed herein are systems, computer-implemented methods, and tangible computer-readable storage media for speaker recognition personalization. The method recognizes speech received from a speaker interacting with a speech interface using a set of allocated resources, the set of allocated resources including bandwidth, processor time, memory, and storage. The method records metrics associated with the recognized speech, and after recording the metrics, modifies at least one of the allocated resources in the set of allocated resources commensurate with the recorded metrics. The method recognizes additional speech from the speaker using the modified set of allocated resources. Metrics can include a speech recognition confidence score, processing speed, dialog behavior, requests for repeats, negative responses to confirmations, and task completions. The method can further store a speaker personalization profile having information for the modified set of allocated resources and recognize speech associated with the speaker based on the speaker personalization profile.


