Speech Recognition System Using Confidence Scoring for IVR Routing
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Traditional speech recognition systems face challenges in accurately processing complex data entries due to limitations in dual-tone multi-frequency (DTMF) interfaces and the intricacies of spoken language, including phonetic variability, environmental noise, and user variability, leading to impracticality in accommodating complex transactions.
Innovation Solution
A communication system incorporating a speech recognition system that utilizes a grammar database and confidence database, integrated with an interactive voice response (IVR) unit, processes voice communications to match user utterances against stored phrases, employing phonetic transcriptions and prosodic units, and includes a verification system for secure transactions, while addressing acoustic and speaker variability through advanced signal processing techniques.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If DTMF interfaces are used for data entry, then system simplicity is maintained, but transaction speed becomes impractically slow for complex data entry
Solution Approach 1:
The patent replaces the mechanical DTMF interface system with a speech recognition system that processes natural language inputs. This substitution enables complex data entry transactions to be completed through voice commands rather than sequential key presses, dramatically improving transaction speed while managing interface complexity through automated speech processing algorithms
2Productivity
If voice based systems are used to replace DTMF input, then transaction speed improves, but speech recognition accuracy deteriorates due to phonetic variability and environmental noise
Solution Approach 1:
The patent implements confidence scoring and verification mechanisms that provide feedback loops for speech recognition. When speech inputs fall below confidence thresholds, the system requests clarification or alternative inputs, thereby maintaining high accuracy despite phonetic variability and environmental noise while preserving the speed benefits of voice-based interaction
Solution Approach 2:
The patent introduces grammar databases and speech-to-text conversion intermediaries that mediate between raw speech inputs and system processing. These intermediaries normalize varied speech patterns into standardized formats, filtering out phonetic variability and environmental noise to improve recognition accuracy while maintaining transaction speed
3Ease of operation
If speech recognition systems process complex transactions, then user interaction efficiency improves, but error rates increase due to the intricacies of spoken language
Solution Approach 1:
The patent pre-configures grammar databases with expected speech patterns, command structures, and data formats before user interactions occur. This preliminary preparation enables the system to efficiently interpret complex transactions with high accuracy, reducing errors while maintaining ease of operation by guiding users through anticipated speech patterns
Data Source
AI summary
A speech recognition process and system are used for interactive telecommunication. A caller is prompted for input. Each of the phrases represents a destination for routing the call. The response utterance is matched by the system to one of the phrases and the call is routed to the corresponding destination. If the call thereafter has been redirected to a destination representing another of the phrases, speech recognition training data are generated for mapping the utterance to the redirected destination.


