Document Data Segment Identification Using OCR and NLP

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current computer vision and process automation technologies are unable to effectively review and extract data from complex, data-rich electronic documents, such as those containing a mixture of text, tables, images, and other content, leading to the need for manual review which is time-consuming and prone to human error.

Innovation Solution

A computer-implemented method using a trained natural language processing model and optical character recognition processor to extract text data, determine candidate entity data, access n-gram words from a knowledge base, and calculate similarity scores to identify and select the most accurate entity data, enabling the review of complex data sets 25 times faster than manual methods.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If existing computer vision and process automation technology is used, then simple electronic documents can be reviewed and extracted, but complex data-rich electronic documents cannot be effectively processed

Engineering Contradiction:
Improvecapability to process complex data-rich documentsVSAvoidautomation of document review
Core Design Contradiction:
Adaptability or versatilityVSExtent of automation

Solution Approach 1:

The patent segments complex data-rich documents into multiple data segments including text segments, table segments, and image segments. Each segment type is processed by specialized extraction modules, enabling the system to handle diverse content types that existing automation technology cannot process effectively.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent creates a universal document processing system that can handle multiple document types and content formats (text, tables, images) through a single integrated platform. The system uses multiple extraction modules that work together to process various data segments, achieving multi-functionality that replaces manual review across different document complexities.

Inventive Principle:
Principle #6Universality (Multi-functionality)

2Measurement precision

If manual review of complex documents is performed, then accurate extraction is achieved, but processing is time-consuming and expensive

Engineering Contradiction:
Improveaccuracy of data extractionVSAvoidprocessing speed
Core Design Contradiction:
Measurement precisionVSProductivity

Solution Approach 1:

The system performs self-service by automatically identifying, segmenting, and extracting data from complex documents without requiring manual intervention. The multiple extraction modules work autonomously to process different data segments, achieving both high accuracy through specialized processing and high productivity through automation.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The patent replaces the mechanical manual review process with an automated computer-based system that uses multiple extraction modules to process document segments. This substitution maintains or improves accuracy while dramatically increasing processing speed and reducing costs associated with manual review.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

3Productivity

If existing automation technology is used, then processing speed is maintained, but accuracy and reliability of extraction deteriorate for complex documents

Engineering Contradiction:
Improveprocessing speedVSAvoidaccuracy of extraction
Core Design Contradiction:
ProductivityVSReliability

Solution Approach 1:

By segmenting complex documents into distinct data segments (text, tables, images) and applying specialized extraction methods to each segment type, the system maintains high accuracy for complex documents while preserving fast automated processing speeds. Each segment is processed by the most appropriate extraction module.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent applies local quality by using different extraction approaches for different data segments based on their specific characteristics. Text segments use natural language processing, table segments use structured data extraction, and image segments use optical character recognition, ensuring each segment is processed with the appropriate method for maximum accuracy.

Inventive Principle:
Principle #3Local quality

Data Source

PatentUS11487798B1Method for identifying a data segment in a data set
Publication Date: 2022.11.01 ALTADA TECH SOLUTIONS LTD
  • US11487798B1 patent drawing
  • US11487798B1 patent drawing
  • US11487798B1 patent drawing

AI summary

A data processing system receives a plurality of electronic documents in image format. For each signature segment, the system determines associated surrounding text using an optical character recognition processor. The system accesses first data stored in a signature knowledge base. The system determines first similarity scores based on the associated surrounding text and the first data using a statistical measure technique. The system selects an optimum signature segment based on a distance metric between each signature segment and each associated surrounding text. For each stamp segment, the system extracts text from the stamp segment using an optical character recognition processor. The system accesses second data stored in a stamp knowledge base. The system determines second similarity scores based on the extracted text and the second data using the statistical measure technique. The system selects an optimum stamp segment based on the second similarity scores.