ML Classifiers for Automated Duplicate Document Detection
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing mechanisms for detecting duplicate documents and document variants require significant human involvement, making them slower and more expensive than desired.
Innovation Solution
The development and utilization of machine-learning generated classifiers that automate the detection of duplicate documents by using human-generated duplicate decision data to train models, allowing for reduced human intervention and faster deployment in production environments.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If machine-learning generated classifiers are used to automate duplicate document detection, then productivity and speed of duplicate detection are improved, but device complexity increases due to the need for training data collection and model development infrastructure
Solution Approach 1:
The system performs preliminary actions by collecting human-generated duplicate decision data in advance and using it to train machine-learning classifiers before actual duplicate detection operations. This pre-training phase enables the system to automatically detect duplicates without requiring human involvement during the actual detection process, thereby improving productivity while managing complexity through upfront preparation.
Solution Approach 2:
The patent introduces machine-learning classifiers as intermediary components between human operators and the duplicate detection process. These classifiers act as mediators that learn from human-generated training data and then automatically perform duplicate detection, reducing the need for direct human involvement in each detection case while maintaining detection quality.
2Measurement precision
If machine-learning classifiers are trained using human-generated duplicate decision data, then measurement precision of duplicate detection is improved, but loss of time increases during the classifier development phase
Solution Approach 1:
The system performs preliminary training actions by collecting human-generated duplicate decision data and training machine-learning classifiers in advance before deployment. This pre-training approach allows the system to achieve high measurement precision in duplicate detection while confining the time investment to the initial classifier development phase, after which the classifiers can automatically detect duplicates without requiring additional human time.
Solution Approach 2:
The patent creates copies of human judgment patterns by training machine-learning classifiers on human-generated duplicate decision data. The classifiers learn to replicate human duplicate detection accuracy, thereby achieving high measurement precision without requiring continuous human involvement in the detection process.
3Ease of operation
If automated machine-learning classifiers are deployed for duplicate detection, then ease of operation is improved by reducing manual effort, but device complexity increases due to automated processing infrastructure requirements
Solution Approach 1:
The system implements self-service by deploying automated machine-learning classifiers that independently perform duplicate detection without requiring manual human operation for each case. The classifiers automatically process documents, apply learned patterns from training data, and identify duplicates autonomously, thereby improving ease of operation while managing complexity through automation.
Solution Approach 2:
The patent replaces the mechanical system of manual human review with an automated machine-learning-based system. The classifiers substitute for human operators by automatically analyzing documents and detecting duplicates based on patterns learned from training data, thereby reducing manual effort while introducing automated processing infrastructure.
Data Source
AI summary
Technologies are disclosed herein for generating and utilizing machine-learning generated classifiers configured to identify document relationships. Manually-generated data is captured that indicates if documents in a document corpus have a relationship with one another, such as duplicates or variations. A determination may then be made as to whether a classifier is to be generated based on the duplicate decision data. If a classifier is to be generated, machine learning may be performed using training documents from the document corpus and the duplicate decision data to generate a classifier. The machine-learning generated classifier may then be utilized in a production environment to determine whether a new document is a duplicate of documents in the document corpus and/or to identify other relationships between documents in the document corpus, such as documents that are similar or are variations of one another.


