Transductive Document Classifier Adapting to Drifting Concepts

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current automatic classification systems rely on rule-based or inductive machine learning methods that require significant manual setup and cannot adapt to dynamically changing environments without manual effort, as they typically use small sets of labeled training examples and struggle with generalizing effectively.

Innovation Solution

The implementation of a transductive machine learning system that uses a processor to classify documents by training a classifier through iterative calculations with unlabeled documents and adjusting cost factors based on expected label values, allowing for adaptation to changing environments and improved classification accuracy.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If rule-based or inductive machine learning methods are used for document classification, then the system can be implemented with conventional approaches, but the system cannot adapt to dynamically changing environments without manual effort and requires significant manual setup

Engineering Contradiction:
Improveadaptability to changing environmentsVSAvoidmanual setup effort
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The transductive classifier automatically adapts to changing classification concepts by utilizing unlabeled documents and iteratively adjusting cost factors based on expected label values, eliminating the need for manual reconfiguration when environments change

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The system dynamically adjusts cost factors during iterative calculations based on expected label values, allowing the classification model to adapt to drifting concepts without manual intervention

Inventive Principle:
Principle #15Dynamics

2Ease of manufacture

If small sets of labeled training examples are used, then the manual effort is reduced, but the system struggles with generalizing effectively

Engineering Contradiction:
Improvemanual effort for training dataVSAvoidgeneralization capability
Core Design Contradiction:
Ease of manufactureVSReliability

Solution Approach 1:

Unlabeled documents serve as intermediaries between the small set of labeled training examples and the final classification model, enabling the system to leverage abundant unlabeled data to improve generalization without requiring extensive manual labeling

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The system changes parameters (cost factors) during iterative calculations based on expected label values, allowing the model to adapt and generalize better from limited labeled examples by learning from the structure of unlabeled data

Inventive Principle:
Principle #35Parameter changes

3Measurement precision

If transductive classifier with iterative calculations is used, then the classification accuracy improves and adaptability to changing environments is achieved, but the computational complexity increases

Engineering Contradiction:
Improveclassification accuracyVSAvoidcomputational complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The system performs iterative calculations where cost factors are adjusted based on expected label values, performing partial retraining iterations that balance computational effort with improved classification accuracy and adaptability

Inventive Principle:
Principle #16Partial or excessive action

Data Source

PatentUS8719197B2Data classification using machine learning techniques
Publication Date: 2014.05.06 TUNGSTEN AUTOMATION CORPORATION
  • US8719197B2 patent drawing
  • US8719197B2 patent drawing
  • US8719197B2 patent drawing

AI summary

Systems, methods and computer program products for classifying documents are presented. Systems, methods and computer program products for analyzing documents, e.g., associated with legal discovery are also presented. Systems, methods and computer program products for cleaning up data are also presented. Systems, methods and computer program products for verifying an association of an invoice with an entity are also presented. Systems, methods and computer program products for managing medical records are presented. Systems, methods and computer program products for face recognition are presented.