IVR Language Model Adaptation via Service Data Weights
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Interactive voice response (IVR) systems face challenges in efficiently recognizing regional and national language variations, dialects, and accents due to inefficient initial analysis, intrusive user interaction, and the requirement for user profiles, which can slow down interactions and be unfeasible in certain settings.
Innovation Solution
The IVR system collects customer-specific service data to generate country or dialect-specific weights, allowing it to set an internal language model before user interaction, enabling accurate speech recognition without user profiling or manual language selection.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If the IVR system analyzes user speech on the fly to determine language model, then it can adapt to different language variations, but speech recognition efficiency deteriorates during initial analysis and accuracy may be compromised
Solution Approach 1:
The system performs preliminary language model selection using customer service data before actual speech recognition occurs. By pre-determining the appropriate language model based on aggregated service data, the system eliminates the need for on-the-fly analysis during speech recognition, thereby maintaining both adaptability and efficiency.
2Measurement precision
If the system asks users to indicate language preferences, then language model selection accuracy improves, but interaction time increases and user experience deteriorates due to additional iterations
Solution Approach 1:
The system automatically determines the appropriate language model by analyzing customer service data without requiring user input or manual language selection. This self-service approach eliminates additional interaction iterations while maintaining accurate language model selection based on aggregated service information.
3Measurement precision
If the system requires users to create user profiles with language preferences, then language model selection becomes more accurate, but device complexity increases and ease of operation deteriorates due to profile creation requirements
Solution Approach 1:
The system leverages existing customer service data to automatically infer and apply appropriate language models without requiring users to create profiles or manually configure language preferences. This approach maintains accuracy while eliminating the operational burden of profile creation.
4Measurement precision
If the IVR system uses customer service data to pre-determine language models, then speech recognition accuracy improves before user interaction, but data processing complexity increases
Solution Approach 1:
The system uses a unified approach of aggregating customer service data across multiple services to determine language models. This multi-functional data aggregation process serves both accuracy improvement and complexity management by leveraging existing service data infrastructure rather than creating separate processing systems.
Data Source
AI summary
Disclosed herein are systems, methods, and non-transitory computer-readable storage media for approximating an accent source. A system practicing the method collects data associated with customer specific services, generates country-specific or dialect-specific weights for each service in the customer specific services list, generates a summary weight based on an aggregation of the country-specific or dialect-specific weights, and sets an interactive voice response system language model based on the summary weight and the country-specific or dialect-specific weights. The interactive voice response system can also change the user interface based on the interactive voice response system language model. The interactive voice response system can tune a voice recognition algorithm based on the summary weight and the country-specific weights. The interactive voice response system can adjust phoneme matching in the language model based on a possibility that the speaker is using other languages.


