TTS Prosody Modification and Random Frequency Overlay for IVR Security

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional text-to-speech (TTS) systems generate speech signals that are recognizable by interactive voice response (IVR) systems, leading to potential security breaches and misappropriation of information, as these systems cannot differentiate between various speech outputs.

Innovation Solution

A method and apparatus that modify at least one prosody characteristic of the speech signal using a prosody sample and overlay a random frequency signal, preventing IVR systems from comprehending the speech while maintaining human intelligibility, by incorporating a prosody modifier and frequency overlay subsystem.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If conventional TTS systems generate speech signals with standard prosody characteristics, then the speech is intelligible to human users, but IVR systems can recognize and comprehend the speech signals leading to security breaches

Engineering Contradiction:
Improvespeech securityVSAvoidspeech intelligibility for IVR systems
Core Design Contradiction:
ReliabilityVSLoss of information

Solution Approach 1:

The patent modifies prosody parameters (pitch, duration, intensity) of the synthesized speech signal to create variations that prevent IVR systems from recognizing the speech while maintaining human intelligibility. The system applies random frequency shifts and temporal modifications to speech segments, changing the acoustic characteristics without affecting semantic content.

Inventive Principle:
Principle #35Parameter changes

Solution Approach 2:

The patent introduces a prosody modification layer as an intermediary between the TTS synthesis engine and the audio output. This intermediary component processes the synthesized speech through multiple transformations including pitch shifting, duration modification, and frequency overlay to create a protected speech signal that preserves meaning for humans but obscures recognition for automated systems.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Reliability

If random frequency signals are overlayed on speech signals to prevent IVR recognition, then speech security is improved, but speech quality may be degraded for human understanding

Engineering Contradiction:
Improvespeech securityVSAvoidspeech quality
Core Design Contradiction:
ReliabilityVSManufacturing precision

Solution Approach 1:

The patent applies random frequency signals at controlled amplitude levels that are sufficient to disrupt IVR recognition algorithms but remain below the threshold of human perception. The frequency overlay is applied partially - only to specific frequency ranges and with limited intensity - ensuring security protection without degrading audible speech quality.

Inventive Principle:
Principle #16Partial or excessive action

Solution Approach 2:

The patent applies different modifications to different parts of the speech signal. Random frequency signals are overlayed selectively on certain frequency bands and time segments rather than uniformly across the entire signal. This localized application ensures that critical speech frequencies remain clear for human understanding while less critical frequencies are modified to prevent automated recognition.

Inventive Principle:
Principle #3Local quality

Data Source

PatentUS7558389B2Method and system of generating a speech signal with overlayed random frequency signal
Publication Date: 2009.07.07 NUANCE COMMUNICATIONS INC
  • US7558389B2 patent drawing
  • US7558389B2 patent drawing
  • US7558389B2 patent drawing

AI summary

A method and apparatus utilizing prosody modification of a speech signal output by a text-to-speech (TTS) system to substantially prevent an interactive voice response (IVR) system from understanding the speech signal without significantly degrading the speech signal with respect to human understanding. The present invention involves modifying the prosody of the speech output signal by using the prosody of the user's response to a prompt. In addition, a randomly generated overlay frequency is used to modify the speech signal to further prevent an IVR system from recognizing the TTS output. The randomly generated frequency may be periodically changed using an overlay timer that changes the random frequency signal at a predetermined intervals.