Assisted Document Classification via Confidence-Guided User Training

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing document classification techniques face challenges with adoption and accuracy due to the need for complete and accurate learning sets, which are often unavailable in real-world scenarios, and users struggle to provide sufficient positive and negative documents for training algorithms.

Innovation Solution

The proposed method involves analyzing a document repository to identify a set of documents corresponding to a sample document, presenting them to the user for manual classification, and calculating a confidence measure to determine the accuracy of the document classification algorithm, allowing for iterative user input until a satisfactory confidence level is reached.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If machine learning algorithms are used for document classification, then classification accuracy can be improved, but the requirement for complete and accurate learning sets makes adoption difficult in real-world scenarios

Engineering Contradiction:
Improveclassification accuracyVSAvoidadoption ease
Core Design Contradiction:
Measurement precisionVSEase of operation

Solution Approach 1:

The system performs preliminary actions by automatically analyzing the document repository and pre-identifying candidate documents before presenting them to the user. This preliminary processing reduces the manual effort required from users by pre-filtering and organizing potential training documents, thereby making adoption easier while maintaining high classification accuracy through comprehensive learning set construction

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The system introduces an intermediary layer between the user and the machine learning algorithm. This intermediary automatically generates, filters, and presents candidate documents to the user for classification, serving as a mediator that handles the complexity of learning set construction. This intermediary approach maintains high accuracy by ensuring comprehensive training data while reducing user burden and improving adoption ease

Inventive Principle:
Principle #24Intermediary (Mediator)

2Measurement precision

If users manually classify documents to create training sets, then the learning set can be more accurate, but users may not have sufficient resources or ability to provide adequate positive and negative documents

Engineering Contradiction:
Improvelearning set accuracyVSAvoidnumber of training documents
Core Design Contradiction:
Measurement precisionVSQuantity of substance

Solution Approach 1:

The system performs preliminary analysis of the entire document repository to pre-identify and filter candidate documents before they are presented to the user. This preliminary action ensures that users receive a curated list of high-quality candidates that are most likely to be useful for training, thereby achieving high learning set accuracy with fewer user-provided documents

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The system applies partial action by presenting only a subset of the most relevant candidate documents to the user for manual classification, rather than requiring users to review all possible documents. This partial approach maintains learning set accuracy by focusing on the most informative documents while significantly reducing the quantity of documents users must process

Inventive Principle:
Principle #16Partial or excessive action

3Measurement precision

If users are presented with many documents for classification, then the training set can be more comprehensive, but the time and effort required for users increases

Engineering Contradiction:
Improveclassification accuracyVSAvoiduser time for classification
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The system performs preliminary filtering and analysis of the document repository to pre-identify the most relevant candidate documents. This preliminary action ensures that users receive a curated, high-quality subset of documents that are most useful for training, thereby achieving comprehensive learning sets with minimal user time investment

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The system applies partial action by presenting only the most relevant subset of documents to users rather than all available documents. This partial approach maintains training comprehensiveness by focusing on high-value documents while significantly reducing the time and effort required from users

Inventive Principle:
Principle #16Partial or excessive action

Data Source

PatentUS9342795B1Assisted learning for document classification
Publication Date: 2016.05.17 EMC IP HLDG CO LLC
  • US9342795B1 patent drawing
  • US9342795B1 patent drawing
  • US9342795B1 patent drawing

AI summary

Methods, apparatus and articles of manufacture for assisted learning for document classification are provided herein. A method includes analyzing a collection of documents within a document repository to identify a set of multiple documents corresponding to a sample document, presenting at least a portion of the set of multiple documents to a user for user classification, and calculating a confidence measure based on the user classification of the at least a portion of the set of multiple documents, wherein said confidence measure corresponds to a level of accuracy by which a document classification algorithm detects one or more documents related to the sample document.