Voice Prompt Optimization via Dynamic Speech Parameter Adjustment

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing automated telephony systems lack the ability to optimize speech parameters effectively, resulting in unclear and unnatural voice prompts that fail to engage a wide range of listeners, particularly in terms of indicating questions and sensitive information.

Innovation Solution

The system optimizes speech parameters such as pitch, duration, and loudness to produce voice prompts that are more natural and effective, using tailored pitch accents and other parameters based on context and demographics, allowing for clear differentiation between questions and information, and adjustment for different languages and populations.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Manufacturing precision

If automated telephony systems use standard voice prompts, then the system operation is simple, but the clarity and naturalness of speech is poor

Engineering Contradiction:
Improveclarity of voice promptsVSAvoidsystem complexity
Core Design Contradiction:
Manufacturing precisionVSDevice complexity

Solution Approach 1:

The system dynamically adjusts multiple speech parameters including pitch contours, pause durations, syllable lengths, and loudness levels to optimize the naturalness and clarity of voice prompts. This involves analyzing the semantic content and modifying acoustic parameters accordingly, such as lowering pitch for important information and extending pauses before questions.

Inventive Principle:
Principle #35Parameter changes

Solution Approach 2:

The voice synthesis system transitions from static, pre-recorded prompts to dynamic, adaptive speech generation that responds to contextual factors. The system continuously adjusts speech parameters based on real-time analysis of the prompt content, ensuring optimal delivery for different types of information (questions, statements, warnings).

Inventive Principle:
Principle #15Dynamics

2Loss of information

If speech parameters are optimized for naturalness, then listener comprehension improves, but processing time increases

Engineering Contradiction:
Improveinformation comprehensionVSAvoidprocessing time
Core Design Contradiction:
Loss of informationVSLoss of time

Solution Approach 1:

The system performs preliminary analysis of the voice prompt content to identify key semantic elements, information types (question, statement, warning), and important words before generating the speech. This pre-processing allows the system to pre-determine optimal pitch contours, pause locations, and emphasis patterns, reducing real-time processing delays while maintaining high comprehension levels.

Inventive Principle:
Principle #10Preliminary action

3Adaptability or versatility

If pitch accents and speech parameters are tailored to context and demographics, then adaptability improves, but device complexity increases

Engineering Contradiction:
Improveadaptability to demographicsVSAvoidsystem complexity
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The system applies different speech parameter settings tailored to specific demographic groups and contextual situations. This includes adjusting pitch ranges, speech rates, and pause durations based on the target audience characteristics and the specific type of information being conveyed, creating locally optimized speech delivery for different scenarios.

Inventive Principle:
Principle #3Local quality

Data Source

PatentUS10229668B2Systems and techniques for producing spoken voice prompts
Publication Date: 2019.03.12 ELIZA CORP
  • US10229668B2 patent drawing
  • US10229668B2 patent drawing
  • US10229668B2 patent drawing

AI summary

Methods and systems are described in which spoken voice prompts can be produced in a manner such that they will most likely have the desired effect, for example to indicate empathy, or produce a desired follow-up action from a call recipient. The prompts can be produced with specific optimized speech parameters, including duration, gender of speaker, and pitch, so as to encourage participation and promote comprehension among a wide range of patients or listeners. Upon hearing such voice prompts, patients/listeners can know immediately when they are being asked questions that they are expected to answer, and when they are being given information, as well as the information that considered sensitive.