Iterative Statistical Sampling for Electronic Document Review

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Organizations face challenges in efficiently reviewing billions of documents and petabytes of data within a reasonable time and budget to identify sensitive information, especially in compliance with data protection regulations.

Innovation Solution

A computer-implemented method involving iterative statistical test processes, including elusion sampling and random sampling, to separate and analyze subsets of documents, ensuring compliance and risk assessment with predefined criteria, utilizing statistical metrics and intelligent separation processes to identify and isolate sensitive data.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If traditional rule-based detection methods are used to review all documents, then comprehensive sensitivity detection is achieved, but the review time and resource consumption become unacceptably high for billions of documents

Engineering Contradiction:
Improvedetection accuracyVSAvoidreview time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The patent divides the document review process into multiple stages: first applying statistical sampling to identify a subset of potentially non-compliant documents, then applying rule-based detection only to this smaller subset. This segmentation maintains detection accuracy while dramatically reducing the time and resources needed compared to reviewing all documents.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent changes the parameter of detection scope from 100% of documents to a statistically determined subset. By using confidence levels and error rates as parameters, the system identifies a smaller portion of documents that are most likely to contain sensitive information, thereby reducing review time while maintaining acceptable detection accuracy.

Inventive Principle:
Principle #35Parameter changes

2Productivity

If statistical sampling methods are used to review only a subset of documents, then review time and resource consumption are reduced, but detection accuracy and compliance assurance may be compromised

Engineering Contradiction:
Improvereview efficiencyVSAvoiddetection accuracy
Core Design Contradiction:
ProductivityVSMeasurement precision

Solution Approach 1:

The patent incorporates feedback mechanisms where the results from statistical sampling inform subsequent rule-based detection. The statistical analysis provides feedback about which documents are most likely to contain sensitive information, allowing the system to focus rule-based detection resources on high-risk documents and thereby maintain detection accuracy while improving efficiency.

Inventive Principle:
Principle #23Feedback

Solution Approach 2:

The patent applies partial action by using statistical sampling to identify a subset of documents for detailed review. Rather than applying full rule-based detection to all documents, the system applies it only to the subset identified as potentially non-compliant, achieving acceptable detection accuracy with reduced resource consumption.

Inventive Principle:
Principle #16Partial or excessive action

3Reliability

If rule-based detection is applied to all documents, then comprehensive compliance verification is achieved, but computational resources and processing time are excessively consumed

Engineering Contradiction:
Improvecompliance assuranceVSAvoidcomputational resource consumption
Core Design Contradiction:
ReliabilityVSUse of energy by moving object

Solution Approach 1:

The patent segments the document processing into two phases: statistical sampling phase and rule-based detection phase. By dividing the workload this way, the system maintains compliance assurance through rule-based detection while reducing computational resource consumption by limiting rule-based analysis to only the subset of documents identified as potentially non-compliant.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent performs preliminary statistical sampling and analysis before applying rule-based detection. This preliminary action identifies which documents are most likely to contain sensitive information, allowing the system to focus computational resources on high-risk documents and thereby reduce overall resource consumption while maintaining compliance assurance.

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS11783088B2Processing electronic documents
Publication Date: 2023.10.10 INTERNATIONAL BUSINESS MACHINE CORPORATION
  • US11783088B2 patent drawing
  • US11783088B2 patent drawing
  • US11783088B2 patent drawing

AI summary

A method for processing electronic documents comprises an iteration including: (i) applying, by a computer device, a first statistical test process to a first subset of the documents, the first statistical test process estimating whether or not content of the documents of the first subset comply with a predefined criterion; (ii) in response to a result of the first statistical test process, estimating, by the computer device, that the documents of the first subset do not comply with the criterion, selecting, by the computer device, a part of the documents of the first subset, and moving, by the computer device, the part of the documents to a second subset of the documents; and (iii) applying, by the computer device, a second statistical test process to the second subset of the documents, the second statistical test process calculating at least one statistical metric related to the documents of the second subset.