AI Document Similarity Identification Using Text Hashing

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Document searches in legal discovery processes are time-consuming due to the large number of documents that need to be searched, reviewed, and identified for relevant text, especially in e-discovery applications.

Innovation Solution

A method and system that identifies similar documents by determining a seed document, receiving a search request, and generating a graphical user interface with a similarity panel to provide data on identical and similar text across multiple documents, using pattern recognition and allowing for filtering and redaction of documents based on legal issues.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If traditional document search methods are used to review all documents in legal discovery, then comprehensive document review can be achieved, but the time required for the process increases significantly

Engineering Contradiction:
Improvedocument review completenessVSAvoidsearch time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The patent creates a digital representation (hash) of the selected text portion and uses this copy to search for identical or similar text across all other documents. Instead of manually reviewing each document, the system copies the text signature and efficiently matches it against the entire document collection, dramatically reducing review time while maintaining completeness.

Inventive Principle:
Principle #26Copying

Solution Approach 2:

The patent replaces the mechanical manual review process with an automated computational system that uses hashing algorithms and pattern recognition. The system automatically identifies similar text portions across documents using computational methods, substituting human manual searching with automated text analysis that is both faster and more consistent.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

2Measurement precision

If manual review of each document is performed to identify relevant text, then accurate identification can be achieved, but the complexity of the process increases

Engineering Contradiction:
Improvetext identification accuracyVSAvoidprocess complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent transforms the document identification problem from a complex semantic analysis task into a simpler parameter-matching problem. By converting text into hash values and comparing these parameters, the system maintains high identification accuracy while dramatically simplifying the underlying process complexity. The parameter transformation allows efficient comparison without requiring complex understanding of document content.

Inventive Principle:
Principle #35Parameter changes

3Reliability

If all documents are searched individually to find similar text, then thorough search coverage is achieved, but the productivity of the search process decreases

Engineering Contradiction:
Improvesearch coverageVSAvoiddocument processing speed
Core Design Contradiction:
ReliabilityVSProductivity

Solution Approach 1:

The patent merges the search process by combining multiple document comparisons into a single efficient operation. Instead of searching each document separately, the system extracts text portions from multiple documents, creates hash representations, and performs batch comparisons. This merging of operations maintains thorough search coverage across all documents while significantly improving processing speed through consolidated operations.

Inventive Principle:
Principle #5Merging (Combining)

Data Source

PatentUS10733193B2Similar document identification using artificial intelligence
Publication Date: 2020.08.04 CASEPOINT LLC
  • US10733193B2 patent drawing
  • US10733193B2 patent drawing
  • US10733193B2 patent drawing

AI summary

Implementations generally relate to processing similar documents. In some implementations, a method includes receiving a plurality of documents related to e-discovery. The method further includes determining a seed document from the plurality of documents. The method further includes receiving a search request to search at least one selection of text in the seed document. The method further includes identifying other documents from the plurality of documents based on a similarity between text in the other documents and the at least one selection of text in the seed document. The method further includes generating a graphical user interface that includes a similarity panel that provides similarity data between text in the other documents and the at least one selection of text in the seed document.