ML Classifiers for Automated Duplicate Document Detection

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing mechanisms for detecting duplicate documents and document variants require significant human involvement, making them slower and more expensive than desired.

Innovation Solution

The development and utilization of machine-learning generated classifiers that automate the detection of duplicate documents by using human-generated duplicate decision data to train models, allowing for reduced human intervention and faster deployment in production environments.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If machine-learning generated classifiers are used to automate duplicate document detection, then productivity and speed of duplicate detection are improved, but device complexity increases due to the need for training data collection and model development infrastructure

Engineering Contradiction:
Improveduplicate detection speedVSAvoidsystem complexity
Core Design Contradiction:
ProductivityVSDevice complexity

Solution Approach 1:

The system performs preliminary actions by collecting human-generated duplicate decision data in advance and using it to train machine-learning classifiers before actual duplicate detection operations. This pre-training phase enables the system to automatically detect duplicates without requiring human involvement during the actual detection process, thereby improving productivity while managing complexity through upfront preparation.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent introduces machine-learning classifiers as intermediary components between human operators and the duplicate detection process. These classifiers act as mediators that learn from human-generated training data and then automatically perform duplicate detection, reducing the need for direct human involvement in each detection case while maintaining detection quality.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Measurement precision

If machine-learning classifiers are trained using human-generated duplicate decision data, then measurement precision of duplicate detection is improved, but loss of time increases during the classifier development phase

Engineering Contradiction:
Improveduplicate detection accuracyVSAvoidclassifier development time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The system performs preliminary training actions by collecting human-generated duplicate decision data and training machine-learning classifiers in advance before deployment. This pre-training approach allows the system to achieve high measurement precision in duplicate detection while confining the time investment to the initial classifier development phase, after which the classifiers can automatically detect duplicates without requiring additional human time.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent creates copies of human judgment patterns by training machine-learning classifiers on human-generated duplicate decision data. The classifiers learn to replicate human duplicate detection accuracy, thereby achieving high measurement precision without requiring continuous human involvement in the detection process.

Inventive Principle:
Principle #26Copying

3Ease of operation

If automated machine-learning classifiers are deployed for duplicate detection, then ease of operation is improved by reducing manual effort, but device complexity increases due to automated processing infrastructure requirements

Engineering Contradiction:
Improvemanual effort requirementVSAvoidautomated processing complexity
Core Design Contradiction:
Ease of operationVSDevice complexity

Solution Approach 1:

The system implements self-service by deploying automated machine-learning classifiers that independently perform duplicate detection without requiring manual human operation for each case. The classifiers automatically process documents, apply learned patterns from training data, and identify duplicates autonomously, thereby improving ease of operation while managing complexity through automation.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The patent replaces the mechanical system of manual human review with an automated machine-learning-based system. The classifiers substitute for human operators by automatically analyzing documents and detecting duplicates based on patterns learned from training data, thereby reducing manual effort while introducing automated processing infrastructure.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Data Source

PatentUS9818066B1Automated development and utilization of machine-learning generated classifiers
Publication Date: 2017.11.14 AMAZON TECH INC
  • US9818066B1 patent drawing
  • US9818066B1 patent drawing
  • US9818066B1 patent drawing

AI summary

Technologies are disclosed herein for generating and utilizing machine-learning generated classifiers configured to identify document relationships. Manually-generated data is captured that indicates if documents in a document corpus have a relationship with one another, such as duplicates or variations. A determination may then be made as to whether a classifier is to be generated based on the duplicate decision data. If a classifier is to be generated, machine learning may be performed using training documents from the document corpus and the duplicate decision data to generate a classifier. The machine-learning generated classifier may then be utilized in a production environment to determine whether a new document is a duplicate of documents in the document corpus and/or to identify other relationships between documents in the document corpus, such as documents that are similar or are variations of one another.