Semantic Categorization Using Lexical Chaining Confidence Scores
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current semantic categorization in IVR systems is labor-intensive and restrictive, requiring manual definition of grammars and tags, limiting flexibility and accuracy in handling user utterances outside pre-defined domains.
Innovation Solution
A system and method for automatic semantic categorization using lexical chaining confidence scores, where each word in category descriptions is paired with semantically related words in WordNet, enabling the categorization algorithm to match user utterances with category descriptions based on semantic similarity scores.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If manual grammar definition and semantic tagging are used for semantic categorization, then the system can accurately understand pre-defined utterances, but the process becomes labor-intensive and restrictive to pre-defined domains
Solution Approach 1:
The system automatically generates grammars and semantic tags by analyzing training utterances and extracting patterns, eliminating the need for manual grammar development. The semantic categorizer self-configures by learning from example utterances and automatically creating the necessary linguistic structures for categorization.
Solution Approach 2:
The patent replaces manual mechanical processes of grammar writing and tag assignment with automated computational processes. Statistical language models and machine learning algorithms automatically analyze utterances and generate grammatical structures, substituting human labor with computational automation.
2Reliability
If fixed-grammar directed-dialog systems are used, then semantic categorization can be performed with pre-defined rules, but the system becomes restrictive and cannot handle utterances outside pre-defined domains
Solution Approach 1:
The system transitions from static pre-defined grammars to dynamic adaptive grammars that automatically adjust based on training data. The statistical language models are trained on diverse utterances and adapt to different domains and speaking styles, allowing the system to handle varied user inputs while maintaining consistent categorization through learned patterns.
Solution Approach 2:
The patent creates a universal semantic categorization system that can handle multiple domains and utterance types through a single automated framework. The statistical language models and semantic analysis engine work across different contexts (banking, flight information, etc.) without requiring domain-specific manual configuration, providing both consistency and versatility.
3Measurement precision
If Statistical Language Models are trained on specific domain data, then transcription accuracy is greatly improved for that domain, but utterances outside the domain receive low confidence scores
Solution Approach 1:
The system performs preliminary training of statistical language models on domain-specific data before deployment, ensuring high transcription accuracy for expected utterances. The grammars and semantic structures are pre-configured through automated training on training sets, so the system is ready to accurately process domain-specific inputs while maintaining the ability to adapt to new domains through retraining.
Data Source
AI summary
There is disclosed a system and method for automatically performing semantic categorization. In one embodiment at least one text description pertaining to a category set is accepted along with words that are anticipated to be uttered by a user pertaining to that category set; lexical chaining confidence score is attached to each pair matched between the anticipated words and the accepted text description. These confidence scores are used subsequently by a categorization circuit that accepts a text phrase utterance from an input source along with a category set pertaining to the accepted utterance. The categorization circuit, in one embodiment, creates word pairs matched between the accepted text phrase utterance and the accepted category set. From these word scores, the category pertaining to the utterance is determined based, at least in part, on the assigned lexical chaining confidence scores as previously determined.


