Document Classification via Stack Grouping and Prime Document Review

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

The current human review process for document classification in litigation is inefficient, leading to high costs, inconsistencies, and inaccuracies due to the sheer volume of data and subjective judgments, which can result in significant time consumption and potential legal repercussions.

Innovation Solution

A system that groups electronic documents into stacks based on similarity using algorithms, with a prime document representing each stack, allowing for targeted quality control and human review, thereby reducing the time and effort required for document classification and ensuring consistency and accuracy.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If human reviewers manually review all documents individually, then classification accuracy can be maintained through subjective judgment, but the time and cost required increases exponentially with document volume

Engineering Contradiction:
Improveclassification accuracyVSAvoidreview time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The patent segments the large document set into multiple stacks based on similarity grouping. Documents are clustered into stacks where each stack contains documents with similar characteristics, and a representative document is selected from each stack for human review. This segmentation reduces the number of documents requiring individual human review while maintaining classification accuracy through the representative sampling approach.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent creates representative copies or summaries of document stacks that capture the essential characteristics of multiple documents. Instead of reviewing every original document, reviewers examine these representative copies which are generated through automated analysis, significantly reducing review time while preserving classification accuracy.

Inventive Principle:
Principle #26Copying

2Productivity

If multiple reviewers work in parallel to increase productivity, then document review throughput increases, but consistency and accuracy decrease due to subjective judgments and fatigue

Engineering Contradiction:
Improvereview throughputVSAvoidclassification consistency
Core Design Contradiction:
ProductivityVSReliability

Solution Approach 1:

The patent transforms the classification task from subjective human judgment to objective algorithmic parameter-based grouping. Documents are clustered using quantitative similarity metrics and parameters, eliminating subjective variability. This allows multiple reviewers to work in parallel on different stacks with consistent, reproducible results based on algorithmic parameters rather than subjective interpretation.

Inventive Principle:
Principle #35Parameter changes

Solution Approach 2:

The patent implements a feedback mechanism where reviewer decisions on representative documents are used to refine and validate the automated grouping algorithms. The system learns from reviewer feedback to improve future grouping accuracy, ensuring that as productivity increases with more reviewers, consistency is maintained through continuous algorithmic refinement based on collective reviewer input.

Inventive Principle:
Principle #23Feedback

3Adaptability or versatility

If traditional human review processes are used, then flexibility in handling diverse document types is maintained, but the cost and time required for reviewing terabytes of data becomes prohibitive

Engineering Contradiction:
Improvehandling flexibilityVSAvoidreview cost
Core Design Contradiction:
Adaptability or versatilityVSLoss of energy

Solution Approach 1:

The patent creates a universal automated grouping system that can handle multiple document types and classification categories simultaneously. The same algorithmic framework applies to different document formats, legal issues, and classification schemes, providing flexibility across diverse document sets without requiring separate manual review processes for each type, thereby reducing overall costs while maintaining adaptability.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Data Source

PatentUS9703863B2Document classification and characterization
Publication Date: 2017.07.11 CONSILIO LLC
  • US9703863B2 patent drawing
  • US9703863B2 patent drawing
  • US9703863B2 patent drawing

AI summary

Data is received that characterizes each of a plurality of documents within a document set. Based on this data, the plurality of documents are grouped into a plurality of stacks using one or more grouping algorithms. A prime document is identified for each stack that includes attributes representative of the entire stack. Subsequently, provision of data is provided that characterizes documents for each stack including at least the identified prime document to at least one human reviewer. User-generated input from the human reviewer is later received that categorized each provided document and data characterizing the user-generated input can then be provided. Related apparatus, systems, techniques and articles are also described.