Expert-Guided Text Classifier for Document Relevance Scoring

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current computerized systems for electronic document analysis are highly dependent on human intuition and creativity, leading to uncontrolled variance in errors, risks, and costs, with manual keyword methods being inefficient and lacking graduated relevance scores, resulting in clumsy and rigid culling and review processes.

Innovation Solution

An expert-guided system that reviews sample documents, learns to score relevance, and iteratively improves accuracy, providing graduated relevance scores, automated prioritization, and keyword generation, enabling more efficient culling and review processes through machine learning and text classification.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Ease of manufacture

If manual keyword methods are used for document culling, then the process is simple to implement, but the precision of relevance identification deteriorates

Engineering Contradiction:
Improveease of implementationVSAvoidrelevance identification precision
Core Design Contradiction:
Ease of manufactureVSMeasurement precision

Solution Approach 1:

The patent replaces manual keyword-based mechanical classification with machine learning-based automated text classification. The system uses trained classifiers to automatically analyze document content and assign relevance scores, substituting human expert intuition with algorithmic processing that provides both automation and improved precision through learned patterns from training data.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Solution Approach 2:

The patent transforms the binary relevant/not-relevant classification into a graduated relevance scoring system. By changing the output parameter from discrete categories to continuous scores, the system provides more nuanced relevance identification while maintaining automated processing. This parameter change enables more precise differentiation between documents of varying relevance levels.

Inventive Principle:
Principle #35Parameter changes

2Device complexity

If manual keyword methods are used for document culling, then the process is rigid and binary, but the adaptability to budget changes and case evolution deteriorates

Engineering Contradiction:
Improveprocess flexibilityVSAvoidadaptability to budget and case changes
Core Design Contradiction:
Device complexityVSAdaptability or versatility

Solution Approach 1:

The patent implements dynamic adaptability by allowing the system to adjust culling thresholds and re-run classification based on changing budget constraints or case priorities. The graduated scoring system enables flexible threshold adjustment without requiring process redesign, and the system can be retrained on new training data to adapt to evolving case requirements, providing both structural simplicity and operational flexibility.

Inventive Principle:
Principle #15Dynamics

3Extent of automation

If expert-based manual review is used, then creativity and intuition can be applied, but the variance in errors and costs increases

Engineering Contradiction:
Improveautomation levelVSAvoiderror consistency
Core Design Contradiction:
Extent of automationVSReliability

Solution Approach 1:

The patent implements self-service through automated machine learning classification that performs document review without requiring human expert intervention for each document. The system trains on initial data and then autonomously classifies subsequent documents, eliminating the need for continuous human expert involvement while maintaining consistent application of classification criteria, thereby reducing variance in errors and costs.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The patent incorporates feedback mechanisms where classification results can be reviewed and used to retrain and improve the classifier. This feedback loop allows the system to learn from corrections and improve its reliability over time while maintaining automated operation, reducing error variance through continuous improvement based on performance feedback.

Inventive Principle:
Principle #23Feedback

4Measurement precision

If graduated relevance scores are implemented, then document prioritization precision improves, but the system complexity increases

Engineering Contradiction:
Improverelevance scoring precisionVSAvoidsystem complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent segments the document review process into distinct phases: training phase where classifiers are trained on labeled data, and deployment phase where trained classifiers automatically score documents. This segmentation allows the complex machine learning operations to be confined to the initial training phase, while the operational phase remains relatively simple, requiring only threshold application and score generation without continuous complex processing.

Inventive Principle:
Principle #1Segmentation

Data Source

PatentUS8914376B2System for enhancing expert-based computerized analysis of a set of digital documents and methods useful in conjunction therewith
Publication Date: 2014.12.16 MICROSOFT ISRAEL RES & DEV 2002 LTD
  • US8914376B2 patent drawing
  • US8914376B2 patent drawing
  • US8914376B2 patent drawing

AI summary

An electronic document analysis method receiving N electronic documents pertaining to a case encompassing a set of issues including at least one issue and establishing relevance of at least the N documents to at least one individual issue in the set of issues, the method comprising, for at least one individual issue from among the set of issues, receiving an output of a categorization process applied to each document in training and control subsets of the at least N documents, the output including, for each document in the subsets, one of a relevant-to-the-individual issue indication and a non-relevant-to-the-individual issue indication; building a text classifier simulating the categorization process using the output for all documents in the training subset of documents; and running the text classifier on the at least N documents thereby to obtain a ranking of the extent of relevance of each of the at least N documents to the individual issue. The method may also comprise evaluating the text classifier's quality using the output for all documents in the control subset.