Adaptive IVR Speech Generation Using User Speech Feedback
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Interactive voice response systems (IVRs) often generate speech that annoys or frustrates users due to inappropriate speech rates, loudness, language, accent, or other characteristics, leading to terminated communication sessions and increased resource consumption as users attempt to avoid interacting with IVRs.
Innovation Solution
A communication platform establishes a communication session with a user device to generate and adapt speech based on user feedback, updating speech generation settings such as rate, language, and accent to match user preferences using natural language processing and machine learning techniques.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Ease of operation
If the IVR uses a fixed speech generation rate, then the system operation is simple, but the user satisfaction deteriorates due to inappropriate speech rates
Solution Approach 1:
The speech generation rate is changed from a fixed static value to a dynamic value that adapts in real-time based on user characteristics. The system continuously monitors user speech and adjusts the speech generation rate dynamically to match user preferences, resolving the contradiction between operational simplicity and adaptability.
Solution Approach 2:
The system implements a feedback loop where user speech is analyzed to determine characteristics such as speech rate, and this feedback is used to adjust the speech generation rate. The feedback mechanism enables the system to learn from user interactions and continuously optimize speech generation parameters without complex manual configuration.
2Productivity
If the IVR generates speech with inappropriate characteristics, then resource consumption increases due to repeated calls, but the speech generation quality deteriorates
Solution Approach 1:
The system uses feedback from user speech to continuously improve speech generation quality. By analyzing user speech characteristics and adjusting generation parameters accordingly, the system reduces the likelihood of user frustration and repeated calls, thereby improving both user satisfaction and resource consumption efficiency.
Solution Approach 2:
The system automatically adjusts speech generation parameters without requiring external intervention or manual configuration. It serves itself by learning from user interactions and autonomously optimizing speech characteristics to prevent user dissatisfaction and reduce resource waste from repeated calls.
3Adaptability or versatility
If the IVR uses multiple speech generation settings, then the speech quality adaptability improves, but the device complexity increases
Solution Approach 1:
The system automatically manages multiple speech generation settings without requiring manual configuration or complex system setup. It self-adjusts parameters such as speech rate, language, and accent based on user speech analysis, thereby maintaining high adaptability while keeping the system configuration simple and intuitive.
Solution Approach 2:
The system changes speech generation parameters dynamically based on user characteristics rather than requiring complex pre-configuration. By adjusting parameters such as speech rate, pitch, and language in real-time based on user feedback, the system achieves high adaptability with minimal system complexity.
Data Source
AI summary
A device causes a communication session to be established between the device and a user device to allow the device and the user device to communicate speech, and receives user speech from the user device. The device processes the user speech using a natural language processing technique to determine a plurality of characteristics of the user speech, and updates a speech generation setting of a plurality of speech generation settings based on the plurality of characteristics of the user speech. The device generates, after updating the speech generation setting, device speech using a text-to-speech technique based on the speech generation setting, and sends the device speech to the user device.


