2D Visual Fingerprinting for Duplicate Document Detection

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

In litigation discovery, the manual inspection of millions of documents for relevant information is costly, time-consuming, and prone to errors due to the presence of duplicate and near-duplicate documents, especially when they contain handwritten text or annotations that Optical Character Recognition (OCR) is unreliable.

Innovation Solution

A system and method using two-dimensional visual fingerprints to automatically detect and highlight duplicate or different document content, allowing for efficient identification and comparison of documents regardless of format or location, by extracting and comparing visual fingerprints and highlighting differences.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If manual inspection of documents is performed to identify duplicate content, then accuracy in detecting handwritten text and annotations is improved, but time consumption and cost increase significantly

Engineering Contradiction:
Improvedetection accuracyVSAvoidreview time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The patent introduces visual fingerprints as an intermediary representation of document content. Instead of directly comparing entire documents or relying on OCR, the system extracts visual fingerprints (key visual features) from documents and compares these compact representations. This intermediary approach enables accurate detection of duplicate content including handwritten text and annotations while dramatically reducing processing time and computational resources.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Productivity

If OCR is used to process document content, then text extraction speed is improved, but reliability deteriorates due to inability to accurately recognize handwritten text and annotations

Engineering Contradiction:
Improveprocessing speedVSAvoidrecognition accuracy
Core Design Contradiction:
ProductivityVSReliability

Solution Approach 1:

The patent replaces the OCR mechanical recognition system with a visual fingerprinting approach. Instead of attempting to transcribe and recognize text characters through OCR, the system extracts visual features directly from the document image. This substitution maintains high processing speed while achieving reliable detection of both printed and handwritten content, as the visual fingerprinting method operates on image data rather than requiring text transcription.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

3Loss of information

If all documents in a large collection are manually reviewed to find relevant information, then completeness of information retrieval is improved, but cost and time consumption increase exponentially

Engineering Contradiction:
Improveinformation completenessVSAvoidreview time
Core Design Contradiction:
Loss of informationVSLoss of time

Solution Approach 1:

The patent extracts visual fingerprints from documents and stores them in a searchable index. When searching for duplicate or similar documents, the system queries this pre-extracted fingerprint index rather than reviewing entire documents. This extraction approach enables rapid identification of relevant documents from large collections by comparing compact fingerprint representations, significantly reducing review time while maintaining the ability to detect all duplicate content through systematic fingerprint matching.

Inventive Principle:
Principle #2Taking out (Extraction)

Data Source

PatentUS8750624B2Detection of duplicate document content using two-dimensional visual fingerprinting
Publication Date: 2014.06.10 GENESEE VALLEY INNOVATIONS LLC
  • US8750624B2 patent drawing
  • US8750624B2 patent drawing
  • US8750624B2 patent drawing

AI summary

A system and method of detecting duplicate document content in a large document collection and automatically highlighting duplicate or different document content among the detected document content using two-dimensional visual fingerprints.