Single Utterance Semantic Extraction via Iterative Recognition Constraints
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional voice response systems require multiple interactions with users due to their rigidly structured questioning approach, which is undesirable in scenarios with significant latency, such as communication over cellular data networks.
Innovation Solution
The system extracts semantically distinct items from a single utterance by repeatedly recognizing the same utterance using constraints provided by previously recognized semantic items, allowing for less structured recognition while maintaining accuracy through repeated processing with locale-specific grammars.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If conventional voice response systems use rigidly structured questioning to extract multiple pieces of information, then recognition accuracy is maintained, but the number of interactions with the user increases
Solution Approach 1:
The patent segments a single utterance into multiple semantic items (e.g., business name, address, phone number) and processes them separately through multiple recognition passes. Each pass focuses on extracting specific semantic categories using targeted grammars, allowing the system to maintain high recognition accuracy for each item type while reducing the total number of user interactions from multiple conversational turns to a single utterance processing sequence
Solution Approach 2:
The system performs preliminary semantic classification and constraint generation before the actual recognition of each semantic item. By pre-processing the utterance to identify potential semantic categories and generating appropriate grammatical constraints in advance, the system prepares recognition templates that improve accuracy for subsequent semantic extraction without requiring additional user interactions
2Measurement precision
If multiple interactions are required with the user to extract information, then structured questioning ensures accurate information extraction, but efficiency decreases in latency-prone environments like cellular networks
Solution Approach 1:
The patent merges multiple information extraction tasks into a single utterance processing operation. Instead of requiring separate interactions for different semantic items (business name, address, phone number), the system processes all semantic categories simultaneously from one utterance by performing multiple recognition passes with different grammatical constraints, thereby improving productivity in latency-prone environments while maintaining extraction accuracy
Solution Approach 2:
The system dynamically changes recognition parameters (grammars, constraints, semantic categories) between different recognition passes of the same utterance. By adjusting the grammatical constraints and focus areas for each pass, the system optimizes information extraction for different semantic categories without requiring additional user input, thus improving efficiency while preserving accuracy
3Ease of operation
If the system uses less structured single utterance recognition, then user convenience increases, but recognition accuracy may deteriorate
Solution Approach 1:
The patent implements dynamic grammar selection and constraint application based on the semantic category being extracted. The system adapts the grammatical constraints and recognition parameters for each semantic item (business name, address, phone number) while maintaining a flexible, unstructured interface for the user. This dynamic adjustment of recognition parameters preserves accuracy for each semantic category while allowing the user to speak naturally in a single utterance
Data Source
AI summary
Semantically distinct items are extracted from a single utterance by repeatedly recognizing the same utterance using constraints provided by semantic items already recognized. User feedback for selection or correction of partially recognized utterance may be used in a hierarchical, multi-modal, or single step manner. An accuracy of recognition is preserved while the less structured and more natural single utterance recognition form is allowed to be used.


