Unsupervised Text Classification via Lexical Label Functions

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current methods for annotating and categorizing text data are cumbersome, time-consuming, and unsuitable for real-time applications, particularly when dealing with large amounts of unorganized text data from various sources, as they require manual annotation and large, frequently updated rulesets, which hinders efficient data privacy and security processes.

Innovation Solution

The solution involves generating labelling functions based on contextual patterns from a lexical database, applying these to unsupervised machine learning algorithms to determine probabilistic labels, and using transformer-based machine learning algorithms for efficient and scalable text data categorization, reducing the need for manual oversight and improving real-time classification accuracy.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If manual annotation methods are used to categorize text data, then classification accuracy can be maintained, but the process becomes cumbersome, time-consuming, and unsuitable for real-time applications

Engineering Contradiction:
Improveclassification accuracyVSAvoidannotation time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The system enables text data to be automatically categorized through unsupervised machine learning algorithms that self-organize data into clusters without human intervention. The algorithm independently determines categories by analyzing contextual patterns and semantic relationships, eliminating the need for manual annotation while maintaining classification quality.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The patent replaces manual mechanical annotation processes with computational machine learning systems. Transformer-based models and unsupervised clustering algorithms substitute human annotators, enabling automated text categorization that operates at machine speed while capturing complex linguistic patterns that manual methods could identify.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

2Adaptability or versatility

If large rulesets are used to categorize text data, then coverage can be improved, but the system becomes complex and requires frequent updates

Engineering Contradiction:
Improvedata coverageVSAvoidruleset complexity
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The system transitions from static ruleset parameters to dynamic learned parameters through machine learning. The model adapts its internal parameters automatically by learning from data distributions and contextual patterns, providing comprehensive coverage without requiring explicit rule maintenance or frequent updates.

Inventive Principle:
Principle #35Parameter changes

Solution Approach 2:

The unsupervised learning algorithm serves multiple functions simultaneously: it performs clustering, categorization, and pattern recognition across diverse text domains. This universal approach replaces multiple specialized rulesets, achieving broad data coverage through a single adaptive system that handles various text types and categories.

Inventive Principle:
Principle #6Universality (Multi-functionality)

3Productivity

If unsupervised machine learning is used for text categorization, then processing speed and scalability improve, but the need for contextual understanding becomes more challenging

Engineering Contradiction:
Improveprocessing speedVSAvoidcontextual understanding
Core Design Contradiction:
ProductivityVSDifficulty of detecting and measuring

Solution Approach 1:

The patent introduces transformer-based language models as intermediaries that bridge unsupervised learning and contextual understanding. These models pre-learn linguistic patterns and semantic relationships from large corpora, enabling the unsupervised clustering algorithm to operate on enriched representations that capture contextual nuances without requiring manual labeling.

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The system performs preliminary processing through pre-trained transformer models that encode contextual information before the unsupervised clustering occurs. This preliminary action embeds contextual understanding into the data representations, allowing the subsequent unsupervised algorithm to work efficiently with already-contextualized inputs, thus maintaining both speed and comprehension.

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS11914630B2Classifier determination through label function creation and unsupervised learning
Publication Date: 2024.02.27 PAYPAL INC
  • US11914630B2 patent drawing
  • US11914630B2 patent drawing
  • US11914630B2 patent drawing

AI summary

Software architectures relating to machine learning (e.g., relating to classifying sequential text data. Unlabeled sequential text data may be produced by a variety of sources such as text messages, email messages, message chats, social media applications, and web pages. Classifying such data may be difficult due to the freeform and unlabeled nature of text data from these sources. Thus, techniques for training a machine learning algorithm to classify unlabeled text data in freeform format. Training is based on generation of labelling functions from lexical databases, applying the labelling functions to unlabeled text data in an unsupervised manner, and generating trained classifiers that accurately classify the unlabeled text data. The trained classifiers may then be implemented classify text data accessed from the variety of sources. The present techniques provide high-quality and efficient labeling of unlabeled text data in freeform formats.