Unstructured Data Extraction via OCR and Machine Learning

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Manual interpretation and execution of trust and estate documents are error-prone and inefficient, requiring significant manual effort and risking financial and reputational issues due to complex linkages and relationships in unstructured data across numerous documents.

Innovation Solution

A trust and estate smart application module that uses machine learning models and optical character recognition (OCR) to automatically extract and structure data from unstructured documents, linking entities with events in a hierarchical format for automated downstream application execution.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If manual interpretation of trust and estate documents is used, then human judgment and flexibility are applied, but errors occur and efficiency is low

Engineering Contradiction:
Improveprocessing speedVSAvoiderror rate
Core Design Contradiction:
ProductivityVSReliability

Solution Approach 1:

The patent replaces manual mechanical interpretation with an automated system comprising optical character recognition (OCR) for document digitization, machine learning models for data extraction, and natural language processing for entity-event linkage. This substitution eliminates human errors while maintaining accurate interpretation through sophisticated algorithms that process complex legal documents systematically.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Solution Approach 2:

The system performs self-service by automatically extracting, validating, and structuring data from trust and estate documents without requiring manual intervention. The machine learning models independently identify entities, events, and relationships, while the system self-corrects through validation mechanisms, enabling autonomous processing that maintains high reliability.

Inventive Principle:
Principle #25Self-service

2Productivity

If manual extraction of data from unstructured documents is performed, then flexibility in handling complex linkages is maintained, but significant manual effort and time are required

Engineering Contradiction:
Improveprocessing speedVSAvoidtime for manual extraction
Core Design Contradiction:
ProductivityVSLoss of time

Solution Approach 1:

The system performs preliminary actions by pre-processing documents through OCR to convert scanned images to editable text, and by pre-segmenting documents into manageable sections using machine learning models. This preparation enables rapid automated extraction of entities and events without requiring manual review of entire documents, significantly reducing processing time.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent segments the complex task of document processing into distinct manageable stages: OCR processing for digitization, machine learning-based sectioning for identifying relevant portions, pattern matching for extracting entities and events, and natural language processing for linking relationships. This segmentation enables parallel processing and reduces overall time requirements.

Inventive Principle:
Principle #1Segmentation

3Productivity

If automated extraction systems are implemented, then processing speed increases, but system complexity increases

Engineering Contradiction:
Improveprocessing speedVSAvoidsystem complexity
Core Design Contradiction:
ProductivityVSDevice complexity

Solution Approach 1:

The system achieves universality by implementing a multi-functional platform that handles document digitization through OCR, extracts data using machine learning models, links entities and events through natural language processing, and structures output in standardized formats. This integrated approach consolidates multiple functions into a single system, managing complexity through functional integration rather than separate components.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The patent introduces intermediary components including a knowledge-based database for storing patterns and rules, and a processing pipeline that mediates between input documents and output structured data. These intermediaries simplify the overall system architecture by providing standardized interfaces and data structures, making the complex automated extraction process manageable and maintainable.

Inventive Principle:
Principle #24Intermediary (Mediator)

4Loss of information

If unstructured data is stored in full detail, then complete information is preserved, but storage requirements increase

Engineering Contradiction:
Improveinformation completenessVSAvoidstorage requirements
Core Design Contradiction:
Loss of informationVSQuantity of substance

Solution Approach 1:

The system extracts only the essential information from unstructured documents, including entities, events, and their relationships, while discarding redundant contextual details. This extraction process creates a condensed structured representation that preserves critical information needed for trust and estate administration while significantly reducing storage requirements compared to storing complete unstructured documents.

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The patent transforms data from unstructured to structured format by changing the organizational parameters of information storage. Instead of storing complete document text, the system reorganizes data into standardized fields, categories, and relationships, reducing the quantity of stored information while maintaining information completeness for administrative purposes.

Inventive Principle:
Principle #35Parameter changes

Applied Scientific Principles

This section explains which scientific principles are used to turn an abstract innovation direction into a practical engineering solution.

Function Achieved in This Case

Enables efficient, error-free extraction and processing of trust and estate data within seconds, reducing storage requirements and automating administrative tasks, thus minimizing operational risks.

Implementation Method 1

applying optical character recognition (OCR) processing to the sectioned digitized data by utilizing an OCR device

Methodology Applied
Scientific EffectOptical character recognition:

Data Source

PatentUS11710099B2Method and apparatus for automatically extracting information from unstructured data
Publication Date: 2023.07.25 JPMORGAN CHASE BANK NA
  • US11710099B2 patent drawing
  • US11710099B2 patent drawing
  • US11710099B2 patent drawing

AI summary

Various methods, apparatuses/systems, and media for automatically extracting information from unstructured data are provided. A receiver receives digitized data of a document having unstructured data format. A processor applies machine learning models for sectioning the digitized data. An OCR device applies an OCR processing to the sectioned digitized data. The processor matches the sectioned digitized data to patterns and rules; applies classification models to the matched digitized data to identify entities and events from the sectioned digitized data; automatically link each entity with corresponding event in a hierarchical format to generate a document having structured data format; and output the document having the structured data with metadata having the linked entity with corresponding event in the hierarchical format to downstream applications.