Dynamic ASR Custom Vocabulary for Chatbot Transcription Accuracy
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing chatbot systems face challenges in accurately recognizing uncommon phrases and domain-specific words during voice interactions, leading to suboptimal intent detection and transcription accuracy.
Innovation Solution
The implementation of a custom vocabulary feature that allows for the upload and usage of domain-specific words and phrases, with varying scopes of recognition based on runtime hints, user metadata, intent/slot type, and other data sources, to enhance transcription accuracy in chatbot interactions.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If a standard vocabulary is used for speech recognition, then the system is simple and easy to operate, but transcription accuracy for domain-specific words and uncommon phrases deteriorates
Solution Approach 1:
The system performs preliminary actions by pre-processing and uploading custom vocabulary lists before speech recognition operations. The custom vocabulary is prepared in advance with domain-specific words, uncommon phrases, and context information, then integrated into the ASR system to improve transcription accuracy for specific domains without requiring complex real-time processing changes.
Solution Approach 2:
The system applies local quality by implementing context-aware vocabulary selection where different vocabulary lists are applied to different contexts, intents, or slots. The ASR system dynamically selects appropriate custom vocabulary based on the specific interaction context, ensuring high transcription accuracy for domain-specific terms while maintaining system simplicity through targeted rather than universal vocabulary enhancement.
2Measurement precision
If a custom vocabulary is implemented to improve domain-specific word recognition, then transcription accuracy improves, but system complexity and computational requirements increase
Solution Approach 1:
The system applies partial action by selectively applying custom vocabulary only when needed for specific intents, slots, or contexts rather than universally for all speech recognition tasks. The ASR system dynamically determines when to activate custom vocabulary based on runtime hints and context, reducing unnecessary computational overhead while maintaining high accuracy for domain-specific recognition where required.
Solution Approach 2:
The system changes parameters by dynamically adjusting vocabulary selection based on runtime conditions such as intent type, slot requirements, and context information. The ASR system modifies its recognition parameters by switching between standard and custom vocabulary lists, optimizing the balance between transcription accuracy and computational resource consumption according to the specific interaction context.
3Measurement precision
If context-aware dynamic vocabulary selection is implemented, then recognition accuracy for specific intents improves, but system complexity increases
Solution Approach 1:
The system implements feedback mechanisms where the ASR system receives runtime hints from downstream components (NLU, dialog state tracker) about the current intent and context, then uses this feedback to dynamically select the appropriate custom vocabulary list. This closed-loop approach ensures high intent detection accuracy by adapting vocabulary selection to the specific interaction context while managing system complexity through structured feedback integration.
Solution Approach 2:
The system applies segmentation by dividing the custom vocabulary into multiple organized lists based on different domains, intents, or slots. The ASR system selectively applies relevant vocabulary segments based on the current interaction context, improving intent detection accuracy for specific domains while managing system complexity through modular, organized vocabulary structures that can be independently selected and maintained.
Data Source
AI summary
Techniques for at least the generation of a chatbot built from a custom vocabulary and to use runtime hints during inference are described. In some examples, the generation of the chatbot includes receiving a request to build a chatbot using a bot definition and a custom vocabulary, wherein the chatbot is to use runtime hints during usage; building the chatbot from the bot definition and custom vocabulary by at least: generating automatic speech recognition (ASR) artifacts to be used in decoding audio input into the chatbot into text for at least one other component of the chatbot to use in determining a next act to be performed, the ASR artifacts including artifacts that use the custom vocabulary and artifacts that do not use the custom vocabulary, and storing the ASR artifacts.


