Iterative Statistical Sampling for Electronic Document Review
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Organizations face challenges in efficiently reviewing billions of documents and petabytes of data within a reasonable time and budget to identify sensitive information, especially in compliance with data protection regulations.
Innovation Solution
A computer-implemented method involving iterative statistical test processes, including elusion sampling and random sampling, to separate and analyze subsets of documents, ensuring compliance and risk assessment with predefined criteria, utilizing statistical metrics and intelligent separation processes to identify and isolate sensitive data.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If traditional rule-based detection methods are used to review all documents, then comprehensive sensitivity detection is achieved, but the review time and resource consumption become unacceptably high for billions of documents
Solution Approach 1:
The patent divides the document review process into multiple stages: first applying statistical sampling to identify a subset of potentially non-compliant documents, then applying rule-based detection only to this smaller subset. This segmentation maintains detection accuracy while dramatically reducing the time and resources needed compared to reviewing all documents.
Solution Approach 2:
The patent changes the parameter of detection scope from 100% of documents to a statistically determined subset. By using confidence levels and error rates as parameters, the system identifies a smaller portion of documents that are most likely to contain sensitive information, thereby reducing review time while maintaining acceptable detection accuracy.
2Productivity
If statistical sampling methods are used to review only a subset of documents, then review time and resource consumption are reduced, but detection accuracy and compliance assurance may be compromised
Solution Approach 1:
The patent incorporates feedback mechanisms where the results from statistical sampling inform subsequent rule-based detection. The statistical analysis provides feedback about which documents are most likely to contain sensitive information, allowing the system to focus rule-based detection resources on high-risk documents and thereby maintain detection accuracy while improving efficiency.
Solution Approach 2:
The patent applies partial action by using statistical sampling to identify a subset of documents for detailed review. Rather than applying full rule-based detection to all documents, the system applies it only to the subset identified as potentially non-compliant, achieving acceptable detection accuracy with reduced resource consumption.
3Reliability
If rule-based detection is applied to all documents, then comprehensive compliance verification is achieved, but computational resources and processing time are excessively consumed
Solution Approach 1:
The patent segments the document processing into two phases: statistical sampling phase and rule-based detection phase. By dividing the workload this way, the system maintains compliance assurance through rule-based detection while reducing computational resource consumption by limiting rule-based analysis to only the subset of documents identified as potentially non-compliant.
Solution Approach 2:
The patent performs preliminary statistical sampling and analysis before applying rule-based detection. This preliminary action identifies which documents are most likely to contain sensitive information, allowing the system to focus computational resources on high-risk documents and thereby reduce overall resource consumption while maintaining compliance assurance.
Data Source
AI summary
A method for processing electronic documents comprises an iteration including: (i) applying, by a computer device, a first statistical test process to a first subset of the documents, the first statistical test process estimating whether or not content of the documents of the first subset comply with a predefined criterion; (ii) in response to a result of the first statistical test process, estimating, by the computer device, that the documents of the first subset do not comply with the criterion, selecting, by the computer device, a part of the documents of the first subset, and moving, by the computer device, the part of the documents to a second subset of the documents; and (iii) applying, by the computer device, a second statistical test process to the second subset of the documents, the second statistical test process calculating at least one statistical metric related to the documents of the second subset.


