Pronunciation Interface for IVR Audio Generation

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Interactive Voice Response (IVR) systems struggle with proper names and places whose pronunciations do not follow predictable rules, leading to faulty audio files that are unrecognizable, making them harder to use and less engaging, especially in international contexts.

Innovation Solution

A pronunciation interface is provided that allows users to edit and set the pronunciation of text strings, breaking them down into parts and offering alternatives, enabling users to select preferred pronunciations and stress patterns, which are then used to generate improved audio files that can be generalized across contexts.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Ease of operation

If fully automated IVR systems are used to generate audio files, then system complexity is reduced and operation is simplified, but pronunciation accuracy deteriorates leading to unrecognizable audio output

Engineering Contradiction:
Improveautomation levelVSAvoidpronunciation accuracy
Core Design Contradiction:
Ease of operationVSManufacturing precision

Solution Approach 1:

The patent segments the word into multiple parts (e.g., syllables or phonemes) and presents alternatives for each part separately. This allows users to precisely control the pronunciation of each segment while maintaining an automated interface. The segmentation enables detailed pronunciation control without requiring full manual audio editing.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent introduces an intermediary interface between the automated system and the user. This interface presents pronunciation alternatives and collects user preferences, acting as a mediator that translates user input into improved audio output. The intermediary maintains automation benefits while incorporating human judgment for accuracy.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Manufacturing precision

If pronunciation alternatives are provided for each part of a word, then pronunciation precision is improved, but interface complexity increases

Engineering Contradiction:
Improvepronunciation precisionVSAvoidinterface complexity
Core Design Contradiction:
Manufacturing precisionVSDevice complexity

Solution Approach 1:

By dividing the word into manageable parts and providing alternatives for each part separately, the interface becomes less overwhelming. Users can focus on selecting pronunciations for individual segments rather than choosing from all possible combinations at once, reducing perceived complexity while maintaining precision.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The system provides more pronunciation alternatives than strictly necessary (excessive action), allowing users to select from comprehensive options. This approach ensures precision by offering all possible correct pronunciations, while the structured presentation prevents interface complexity from becoming unmanageable.

Inventive Principle:
Principle #16Partial or excessive action

3Manufacturing precision

If user input is collected for pronunciation preferences, then linguistic quality is improved, but processing time increases

Engineering Contradiction:
Improvelinguistic qualityVSAvoidprocessing time
Core Design Contradiction:
Manufacturing precisionVSLoss of time

Solution Approach 1:

The system collects pronunciation preferences in advance during interface interaction, storing them for future use. This preliminary action allows the pronunciation data to be captured before actual audio generation, so that subsequent audio files can be generated quickly using the pre-collected preferences without repeated user input.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The system implements feedback loops where user pronunciation selections are captured and used to improve future audio generations. The feedback mechanism learns from user inputs and applies the preferences systematically, reducing processing time for subsequent operations while maintaining high linguistic quality.

Inventive Principle:
Principle #23Feedback

Data Source

PatentUS8160881B2Human-assisted pronunciation generation
Publication Date: 2012.04.17 MICROSOFT TECHNOLOGY LICENSING LLC
  • US8160881B2 patent drawing
  • US8160881B2 patent drawing
  • US8160881B2 patent drawing

AI summary

Pronunciation generation may be provided. First, a pronunciation interface may be provided. The pronunciation interface may be configured to display a word and a plurality of alternatives corresponding to a one of a plurality of parts of the word. The plurality of parts may comprise phonemes or syllables of the word. Next, pronunciation data may be received through the pronunciation interface. The pronunciation data may indicate one of the plurality of alternatives. Then a pronunciation of the word may be generated based upon the received pronunciation data. The pronunciation may correspond to the indicated one of the plurality of alternatives. In addition, the pronunciation data may indicate which one of the plurality of parts of the word is stressed. This stress indication may be received in response to a user sliding a user selectable element to indicate which one of the plurality of parts of the word is stressed.