Speech Recognition System Using Confidence Scoring for IVR Routing

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Traditional speech recognition systems face challenges in accurately processing complex data entries due to limitations in dual-tone multi-frequency (DTMF) interfaces and the intricacies of spoken language, including phonetic variability, environmental noise, and user variability, leading to impracticality in accommodating complex transactions.

Innovation Solution

A communication system incorporating a speech recognition system that utilizes a grammar database and confidence database, integrated with an interactive voice response (IVR) unit, processes voice communications to match user utterances against stored phrases, employing phonetic transcriptions and prosodic units, and includes a verification system for secure transactions, while addressing acoustic and speaker variability through advanced signal processing techniques.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If DTMF interfaces are used for data entry, then system simplicity is maintained, but transaction speed becomes impractically slow for complex data entry

Engineering Contradiction:
Improvetransaction speedVSAvoidinterface complexity
Core Design Contradiction:
ProductivityVSDevice complexity

Solution Approach 1:

The patent replaces the mechanical DTMF interface system with a speech recognition system that processes natural language inputs. This substitution enables complex data entry transactions to be completed through voice commands rather than sequential key presses, dramatically improving transaction speed while managing interface complexity through automated speech processing algorithms

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

2Productivity

If voice based systems are used to replace DTMF input, then transaction speed improves, but speech recognition accuracy deteriorates due to phonetic variability and environmental noise

Engineering Contradiction:
Improvetransaction speedVSAvoidspeech recognition accuracy
Core Design Contradiction:
ProductivityVSMeasurement precision

Solution Approach 1:

The patent implements confidence scoring and verification mechanisms that provide feedback loops for speech recognition. When speech inputs fall below confidence thresholds, the system requests clarification or alternative inputs, thereby maintaining high accuracy despite phonetic variability and environmental noise while preserving the speed benefits of voice-based interaction

Inventive Principle:
Principle #23Feedback

Solution Approach 2:

The patent introduces grammar databases and speech-to-text conversion intermediaries that mediate between raw speech inputs and system processing. These intermediaries normalize varied speech patterns into standardized formats, filtering out phonetic variability and environmental noise to improve recognition accuracy while maintaining transaction speed

Inventive Principle:
Principle #24Intermediary (Mediator)

3Ease of operation

If speech recognition systems process complex transactions, then user interaction efficiency improves, but error rates increase due to the intricacies of spoken language

Engineering Contradiction:
Improveuser interaction efficiencyVSAvoiddata entry accuracy
Core Design Contradiction:
Ease of operationVSReliability

Solution Approach 1:

The patent pre-configures grammar databases with expected speech patterns, command structures, and data formats before user interactions occur. This preliminary preparation enables the system to efficiently interpret complex transactions with high accuracy, reducing errors while maintaining ease of operation by guiding users through anticipated speech patterns

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS8077835B2Method and system of providing interactive speech recognition based on call routing
Publication Date: 2011.12.13 VERIZON PATENT & LICENSING INC
  • US8077835B2 patent drawing
  • US8077835B2 patent drawing
  • US8077835B2 patent drawing

AI summary

A speech recognition process and system are used for interactive telecommunication. A caller is prompted for input. Each of the phrases represents a destination for routing the call. The response utterance is matched by the system to one of the phrases and the call is routed to the corresponding destination. If the call thereafter has been redirected to a destination representing another of the phrases, speech recognition training data are generated for mapping the utterance to the redirected destination.