Speech Recognition Personalization Through Context-Specific Word Boosting

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing automatic speech recognition systems struggle to accurately reproduce personalized content such as company names, company jargon, industry-specific terms, and slang terms, especially in end-to-end models, due to ambiguity and variability in pronunciation.

Innovation Solution

A method that detects unrecognized words in a communication session, associates them with contextual data, and applies context-specific boosting to improve the accuracy of speech recognition models by identifying a subset of unrecognized words based on contextual criteria.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Extent of automation

If end-to-end speech recognition models are used, then processing speed and automation are improved, but accuracy in recognizing personalized content deteriorates

Engineering Contradiction:
Improveautomation of speech recognitionVSAvoidaccuracy in recognizing personalized content
Core Design Contradiction:
Extent of automationVSMeasurement precision

Solution Approach 1:

The system segments the speech recognition process into multiple components: an end-to-end model for general transcription and a separate context-specific boosting component for personalized content. This segmentation allows each component to specialize - the end-to-end model handles overall processing while the boosting component targets specific weaknesses in recognizing company names, jargon, and industry terms.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The system applies local quality by implementing context-specific boosting that enhances recognition accuracy for particular types of personalized content (company names, jargon, industry-specific terms) without affecting the entire speech recognition system. This targeted approach improves local accuracy where needed while maintaining overall system efficiency.

Inventive Principle:
Principle #3Local quality

2Measurement precision

If context-specific boosting is applied to improve recognition accuracy, then measurement precision is improved, but device complexity increases

Engineering Contradiction:
Improveaccuracy in transcribing personalized contentVSAvoidcomplexity of speech recognition system
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The system performs preliminary action by pre-identifying and boosting context-specific words before the actual speech recognition process. Context-specific boosting is applied in advance based on contextual data from the audio, allowing the end-to-end model to focus on general transcription while personalized content receives pre-enhanced attention.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The system introduces an intermediary mechanism - context-specific boosting - that acts as a mediator between the audio input and the end-to-end model. This intermediary layer processes contextual data and applies selective enhancement to recognized words, bridging the gap between general speech recognition and personalized content accuracy without requiring complete system redesign.

Inventive Principle:
Principle #24Intermediary (Mediator)

Data Source

PatentUS20250259627A1Automatic personalization for speech recognition systems
Publication Date: 2025.08.14 CISCO TECHNOLOGY INC
  • US20250259627A1 patent drawing
  • US20250259627A1 patent drawing
  • US20250259627A1 patent drawing

AI summary

According to one or more embodiments of the disclosure, automatic personalization for speech recognition systems is provided by a method that includes detecting, by a device, unrecognized words within an automated transcript of audio from a communication session and associating, by the device, the unrecognized words with corresponding contextual data. The method further includes identifying, by the device, a subset of the unrecognized words for boosting and applying, by the device, a context-specific boosting to the subset of the unrecognized words within an automated speech recognition model when criteria, identified based on contextual data associated with the subset of the unrecognized words, are met within the audio from the communication session.