Adaptive IVR Voice Menus Using Speech Characteristics
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing computer telephony integration (CTI) systems with static and prerecorded voice menus are often overly generic and unresponsive to user requests, leading to inefficiencies in call handling and increased processing loads.
Innovation Solution
A system and method that dynamically adjusts interactive voice response (IVR) features based on user speech characteristics using a generative AI system with a centralized dialogue manager, which includes a machine-learning model to identify speech and voice characteristics and generate personalized voice interactions.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Ease of operation
If static and prerecorded voice menus are used in CTI systems, then system complexity is reduced and ease of operation is improved, but adaptability to user requests deteriorates and call handling efficiency decreases
Solution Approach 1:
The patent implements dynamic voice menus that adapt in real-time to user speech characteristics. The system analyzes user input patterns, speech rate, and language preferences, then dynamically adjusts menu options and response timing. This transforms the static voice menu into a dynamic system that evolves during the interaction, resolving the contradiction between operational simplicity and adaptability.
Solution Approach 2:
The system changes multiple parameters including speech rate, language, menu depth, and response timing based on user characteristics. By dynamically adjusting these parameters during the call, the system maintains ease of operation while significantly improving adaptability to individual user needs and preferences.
2Device complexity
If static voice menus are used, then device complexity is reduced, but processing workloads increase due to inefficient call handling and rerouting
Solution Approach 1:
The patent implements self-service capabilities where the voice menu system automatically analyzes user speech characteristics and adjusts itself without requiring external intervention. The system performs real-time speech analysis, intent recognition, and dynamic menu reconfiguration autonomously, reducing the need for complex backend processing and call rerouting to human agents.
Solution Approach 2:
The system performs preliminary speech analysis and user characteristic identification early in the interaction, before the main service delivery. By pre-adapting the voice menu based on initial user input patterns, the system reduces subsequent processing workloads and eliminates the need for complex call rerouting operations.
3Adaptability or versatility
If dynamic adjustment of IVR features is implemented, then adaptability to user requests is improved, but device complexity and processing workloads increase
Solution Approach 1:
The patent segments the voice response system into modular components: speech characteristic analysis module, intent recognition module, dynamic menu generation module, and response synthesis module. Each module performs a specific function independently, making the overall complex system manageable and maintainable while achieving high adaptability through coordinated module operation.
4Adaptability or versatility
If personalized voice interactions are generated in real-time, then adaptability and user satisfaction are improved, but execution time and latency increase
Solution Approach 1:
The system performs preliminary speech characteristic analysis and user profiling during initial interactions or idle periods. By pre-processing and storing user preferences, speech patterns, and language characteristics, the system minimizes real-time processing requirements and reduces latency when generating personalized voice interactions during actual calls.
Solution Approach 2:
The patent applies different processing levels to different parts of the interaction. Critical path elements like intent recognition use optimized fast-processing algorithms, while less time-sensitive elements like detailed speech pattern analysis use more computationally intensive methods. This local quality approach balances adaptability with execution time requirements.
Data Source
AI summary
A system includes a memory configured to store user profiles associated with a plurality of users and an interactive voice response (IVR) system configured to service calls. The system includes processors configured to receive a call from a first user, generate a first voice interaction configured to prompt the first user to perform an utterance of a second voice interaction, and detect the utterance of the second voice interaction. The processors are configured to detect the utterance of the second voice interaction, execute a machine-learning model trained to identify speech and voice characteristics of the first user and to generate a third voice interaction based on the identified speech and voice characteristics, dynamically adjust IVR response features associated with the third voice interaction based on the identified speech and voice characteristics, and output the third voice interaction in accordance with the dynamically adjusted one or more IVR response features.


