Dynamic Speech Resource Allocation for Personalization

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing speech recognition applications rely on generic models that are not well-suited for individual speakers, leading to inefficient resource usage and high costs for personalization, as complete personalization for every speaker is often unnecessary and resource-intensive.

Innovation Solution

A system that recognizes speech using a set of allocated resources, records metrics such as confidence scores and dialog behavior, and modifies these resources dynamically to provide personalized speech recognition, allocating more resources as needed for speakers with unique dialects or accents, while minimizing unnecessary personalization for others.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If complete personalization is provided for every speaker, then speech recognition accuracy for individual speakers is improved, but resource consumption (processing power, storage, bandwidth) becomes prohibitive

Engineering Contradiction:
Improvespeech recognition accuracyVSAvoidresource consumption
Core Design Contradiction:
Measurement precisionVSUse of energy by moving object

Solution Approach 1:

The patent applies local quality by providing personalized speech recognition models only to specific speakers who need them, rather than uniformly to all speakers. The system identifies speakers with unusual dialects or accents and allocates personalization resources selectively to these local cases, while using generic models for the majority of speakers.

Inventive Principle:
Principle #3Local quality

Solution Approach 2:

The system dynamically adjusts the level of personalization based on speaker characteristics and performance metrics. It monitors speech recognition confidence scores and dialog behavior in real-time, modifying resource allocation and model complexity adaptively to match actual speaker needs rather than using static personalization for all users.

Inventive Principle:
Principle #15Dynamics

2Measurement precision

If personal speech recognition models are generated for all users through training phase, then individual speaker recognition accuracy is improved, but processing time and complexity increase significantly

Engineering Contradiction:
Improveindividual speaker recognition accuracyVSAvoidtraining time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The patent implements partial personalization by providing complete training and personalization only when necessary, rather than for all users. The system performs full personalization selectively for speakers with unusual characteristics who benefit most from it, while accepting partial performance with generic models for other speakers, thus reducing overall training time and complexity.

Inventive Principle:
Principle #16Partial or excessive action

Solution Approach 2:

The system allows speakers to implicitly request personalization through their speech patterns during normal interaction. By monitoring confidence scores and dialog behavior, the system automatically identifies when a speaker would benefit from personalization and initiates the training process only for those cases, rather than requiring upfront training for all users.

Inventive Principle:
Principle #25Self-service

3Productivity

If generic speech model is used for all speakers, then resource efficiency is improved, but speech recognition accuracy for speakers with unusual dialects or accents deteriorates

Engineering Contradiction:
Improveresource efficiencyVSAvoidspeech recognition accuracy for unusual speakers
Core Design Contradiction:
ProductivityVSMeasurement precision

Solution Approach 1:

The patent introduces an intermediary monitoring system that tracks speech recognition confidence scores and dialog behavior metrics. This intermediary layer detects when generic model performance degrades for specific speakers and triggers appropriate personalization actions, bridging the gap between efficient generic processing and accurate personalized recognition.

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The system implements feedback loops that continuously monitor speech recognition performance metrics including confidence scores, requests for repeats, and negative responses to confirmations. This feedback information is used to dynamically adjust resource allocation and trigger personalization for speakers experiencing poor performance with the generic model, thereby maintaining both efficiency and accuracy.

Inventive Principle:
Principle #23Feedback

Data Source

PatentUS11620988B2System and method for speech personalization by need
Publication Date: 2023.04.04 MICROSOFT TECHNOLOGY LICENSING LLC
  • US11620988B2 patent drawing
  • US11620988B2 patent drawing
  • US11620988B2 patent drawing

AI summary

Disclosed herein are systems, computer-implemented methods, and tangible computer-readable storage media for speaker recognition personalization. The method recognizes speech received from a speaker interacting with a speech interface using a set of allocated resources, the set of allocated resources including bandwidth, processor time, memory, and storage. The method records metrics associated with the recognized speech, and after recording the metrics, modifies at least one of the allocated resources in the set of allocated resources commensurate with the recorded metrics. The method recognizes additional speech from the speaker using the modified set of allocated resources. Metrics can include a speech recognition confidence score, processing speed, dialog behavior, requests for repeats, negative responses to confirmations, and task completions. The method can further store a speaker personalization profile having information for the modified set of allocated resources and recognize speech associated with the speaker based on the speaker personalization profile.