Word Classification Engine Using Correlation Measures
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Manual determination of psycholinguistic properties or classes of words and expressions is costly and results in datasets of limited size, making it inefficient for AI applications and content selection.
Innovation Solution
A method involving training a classifier using positive and negative training data to determine the measure of correlation between a word and a psycholinguistic class, allowing for automatic classification and content selection, utilizing machine learning techniques such as bidirectional recurrent neural networks and nearest neighbor algorithms.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If manual determination of psycholinguistic properties is used, then measurement precision is improved, but productivity deteriorates and loss of time increases
Solution Approach 1:
The patent uses contextual patterns from manually classified sentences as templates to automatically classify new sentences. The system copies the classification logic embedded in human-labeled examples and applies it systematically to large datasets, achieving both accuracy and scalability.
Solution Approach 2:
The patent replaces manual mechanical classification processes with an automated computational system. Instead of human experts manually analyzing each sentence, the system uses algorithmic processing of contextual patterns to determine psycholinguistic properties, dramatically increasing productivity while maintaining precision.
2Measurement precision
If manual determination of psycholinguistic properties is used, then measurement precision is improved, but loss of time increases
Solution Approach 1:
The patent performs preliminary classification by identifying contextual patterns in training sentences that are already manually labeled. The system pre-extracts features and relationships from these examples, building a classification model in advance that can quickly process new sentences without requiring real-time manual analysis.
Solution Approach 2:
The system copies classification patterns from pre-labeled sentences and applies them to new data, eliminating the need for repeated manual analysis and significantly reducing processing time while maintaining accuracy.
3Productivity
If automatic classification methods are used, then productivity is improved, but measurement precision deteriorates
Solution Approach 1:
The patent enables the classification system to self-improve by learning from manually labeled examples. The system automatically extracts contextual patterns and classification rules from training data, then applies these self-derived rules to classify new sentences, achieving both automation and accuracy.
Solution Approach 2:
The patent replaces inaccurate simple automatic methods with a sophisticated pattern-matching system that uses contextual analysis. The system substitutes basic keyword matching with comprehensive contextual pattern recognition, maintaining precision while achieving automated productivity.
Data Source
AI summary
Method and apparatus for training and using a classifier for words. Embodiments include receiving a first plurality of sentences comprising a first word that is associated with a class and a second plurality of sentences comprising a second word that is not associated with the class. Embodiments include training a classifier using positive training data for the class that is based on the first plurality of sentences and negative training data for the class that is based on the second plurality of sentences. Embodiments include determining a measure of correlation between a third word and the class by using a sentence comprising the third word as an input to the classifier. Embodiments include using the measure of correlation to perform an action selected from the following list: selecting content to provide to a user; determining an automatic chat response; or filtering a set of content.


