Speech-to-Tone Conversion for IVR Accessibility

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing IVR systems that do not employ speech recognition require users to manually select numbers on a keypad, which can be distracting, cumbersome, and inaccessible for tactile-challenged individuals, and upgrading these systems to include voice recognition can be costly and time-consuming.

Innovation Solution

Implementing a speech-to-tone application on user devices that converts spoken words into DTMF tones, allowing users to interact with DTMF-enabled systems, such as IVR, without the need to press keys, by automatically launching during calls and transmitting relevant tones to the system.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Ease of operation

If users manually select numbers on a keypad to respond to IVR, then the system can process the response, but the user experiences distraction and cumbersome operation

Engineering Contradiction:
Improveease of responding to IVRVSAvoiduser distraction
Core Design Contradiction:
Ease of operationVSObject-affected harmful factors

Solution Approach 1:

The patent replaces the mechanical keypad input system with an acoustic/speech recognition system. The user speaks commands instead of pressing physical keys, substituting mechanical interaction with voice-based interaction. This is implemented through speech-to-text conversion that translates spoken words into the DTMF tone sequences required by the IVR system.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

2Adaptability or versatility

If IVR systems are upgraded to include speech recognition, then voice command recognition is enabled, but the cost and time for upgrading increase

Engineering Contradiction:
Improvevoice command recognition capabilityVSAvoidupgrade cost and time
Core Design Contradiction:
Adaptability or versatilityVSEase of manufacture

Solution Approach 1:

The patent introduces an intermediary speech-to-text conversion layer between the user and the existing IVR system. This intermediary component translates speech into DTMF tones, allowing voice command recognition without modifying the core IVR system. The intermediary handles the complexity of speech recognition while the legacy system remains unchanged, avoiding expensive and time-consuming upgrades.

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The speech-to-text conversion component serves multiple functions: it enables voice command recognition, maintains compatibility with existing DTMF systems, and can be applied to various IVR implementations without system-specific customization. This universal approach allows the same solution to work across different platforms and reduces overall upgrade costs.

Inventive Principle:
Principle #6Universality (Multi-functionality)

3Reliability

If users are required to press keys on a keypad, then the IVR system can receive input, but tactile-challenged individuals face accessibility difficulties

Engineering Contradiction:
Improveinput reception reliabilityVSAvoidaccessibility for tactile-challenged users
Core Design Contradiction:
ReliabilityVSEase of operation

Solution Approach 1:

The patent replaces the mechanical keypad interface with a voice-based interface, eliminating the need for tactile key pressing. Users with tactile challenges can speak commands instead of attempting to locate and press physical keys, making the system accessible to a broader range of users while maintaining reliable input reception through speech-to-text conversion.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Data Source

PatentEP3667661B1Speech to dual-tone multifrequency system and method
Publication Date: 2024.07.03 MITEL CORP
  • EP3667661B1 patent drawingFigure 1
  • EP3667661B1 patent drawingFigure 2
  • EP3667661B1 patent drawingFigure 3

AI summary

A method and system for converting speech to tones and transmitting the tones to another device are disclosed. The method can include determining when a communication is initiated, automatically launching a speech-to-tone application on the communication device, determining when pre-defined words are spoken, performing one or more of converting the pre-defined words to a signal comprising a tone using the communication device and converting a stored key sequence to a signal comprising a tone using the communication device, and transmitting the signal to another device. The system can include one or more devices to perform the method.