Adaptive IVR Speech Settings Based on User Voice Characteristics

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Interactive voice response systems (IVRs) often generate speech that annoys or frustrates users due to inappropriate speech rates, loudness, language, accent, or other characteristics, leading to terminated communication sessions and increased resource consumption as users attempt to avoid interacting with IVRs.

Innovation Solution

A communication platform establishes a communication session with a user device, generates speech based on initial settings, receives user speech, determines its characteristics, and updates speech generation settings to match user preferences, continuously adapting to ensure more pleasing speech characteristics.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Ease of operation

If the IVR uses fixed speech generation settings, then the system is simple to operate, but the speech characteristics may annoy or frustrate users leading to terminated sessions

Engineering Contradiction:
Improvesystem operation simplicityVSAvoidspeech characteristic adaptability
Core Design Contradiction:
Ease of operationVSAdaptability or versatility

Solution Approach 1:

The speech generation settings are transformed from static fixed values to dynamic adjustable parameters. The system continuously monitors user speech characteristics (rate, loudness, language, accent) and automatically adjusts speech generation settings in real-time to match user preferences, making the system adaptive rather than fixed

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

The system implements a feedback loop where user speech is analyzed to determine characteristics such as rate, loudness, language, and accent. This feedback is then used to update speech generation settings, creating a closed-loop control system that continuously optimizes speech output based on user response

Inventive Principle:
Principle #23Feedback

2Productivity

If the IVR continuously adapts speech settings based on user speech, then user engagement is maintained, but processing complexity and resource consumption increase

Engineering Contradiction:
Improveuser engagement maintenanceVSAvoidprocessing complexity
Core Design Contradiction:
ProductivityVSDevice complexity

Solution Approach 1:

The system performs preliminary analysis of user speech characteristics early in the interaction to establish baseline preferences. By determining user speech rate, loudness, language, and accent at the beginning and during the session, the system proactively configures appropriate speech generation settings before user frustration occurs, rather than waiting for problems to manifest

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The system modifies specific speech generation parameters (rate, volume, language, accent) based on analyzed user speech characteristics. By changing these individual parameters independently and systematically, the system achieves complex adaptation through manageable parameter adjustments rather than overhauling the entire processing system

Inventive Principle:
Principle #35Parameter changes

3Adaptability or versatility

If the IVR analyzes user speech characteristics in real-time, then speech generation can be optimized, but processing time and computational resources increase

Engineering Contradiction:
Improvespeech optimization capabilityVSAvoidprocessing time
Core Design Contradiction:
Adaptability or versatilityVSLoss of time

Solution Approach 1:

The system analyzes only the most critical speech characteristics (rate, loudness, language, accent) rather than performing comprehensive speech analysis. By focusing on these key parameters that have the greatest impact on user experience, the system achieves effective optimization with minimal processing overhead, avoiding unnecessary analysis of less important features

Inventive Principle:
Principle #16Partial or excessive action

Data Source

PatentUS10468014B1Updating a speech generation setting based on user speech
Publication Date: 2019.11.05 CAPITAL ONE SERVICES LLC
  • US10468014B1 patent drawing
  • US10468014B1 patent drawing
  • US10468014B1 patent drawing

AI summary

A device causes a communication session to be established between the device and a user device to allow the device and the user device to communicate speech, and receives user speech from the user device. The device processes the user speech using a natural language processing technique to determine a plurality of characteristics of the user speech, and updates a speech generation setting of a plurality of speech generation settings based on the plurality of characteristics of the user speech. The device generates, after updating the speech generation setting, device speech using a text-to-speech technique based on the speech generation setting, and sends the device speech to the user device.