Document Data Extraction via Spatial Anchor Blocks

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current technologies face challenges in efficiently extracting and processing large volumes of unstructured and semi-structured data from documents, as they often lack the ability to accurately identify and associate specific data within vast electronic repositories, leading to undetected time-sensitive information.

Innovation Solution

A computer system that uses optical character recognition (OCR) and post-OCR processing to identify and group words and contender values into anchor blocks based on spatial positioning, and then determines confidence levels through multi-dimensional analysis, including spatial, typographical, and keyword associations, to extract positively associated values from documents.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If traditional data processing methods are used on unstructured and semi-structured documents, then the system can maintain simplicity in processing, but the ability to accurately identify and extract specific data from large document volumes deteriorates

Engineering Contradiction:
Improvedata extraction accuracyVSAvoidprocessing system complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent segments the document processing task into multiple distinct modules: OCR processing module, spatial relationship analysis module, anchor block identification module, and data extraction module. Each module handles a specific aspect of the processing pipeline, allowing the system to achieve high extraction accuracy through specialized processing at each stage while managing overall complexity through modular architecture.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent introduces spatial dimension analysis by examining the two-dimensional positioning of text elements, anchor blocks, and data fields within documents. This spatial dimension enables the system to accurately identify relationships between elements and extract data from unstructured layouts, transforming a traditionally one-dimensional text processing problem into a multi-dimensional analysis that significantly improves extraction precision.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

2Productivity

If manual review and specialized expertise are used to extract data from documents, then extraction accuracy is maintained, but processing time and operational costs increase significantly

Engineering Contradiction:
Improvedocument processing speedVSAvoidtime for data extraction
Core Design Contradiction:
ProductivityVSLoss of time

Solution Approach 1:

The patent implements self-service through automated OCR processing, spatial relationship analysis, and pattern recognition algorithms that enable the system to extract data from documents autonomously without requiring manual review or specialized human expertise. The system automatically identifies anchor blocks, determines spatial relationships, and extracts relevant data fields, achieving both high productivity and rapid processing speeds while eliminating time-consuming manual operations.

Inventive Principle:
Principle #25Self-service

3Measurement precision

If comprehensive multi-dimensional analysis is performed on each document, then data extraction accuracy improves, but processing time and computational resources increase

Engineering Contradiction:
Improvevalue association accuracyVSAvoidcomputational resource consumption
Core Design Contradiction:
Measurement precisionVSUse of energy by stationary object

Solution Approach 1:

The patent applies preliminary action by performing OCR processing and initial text recognition before conducting spatial relationship analysis. The system pre-processes documents to identify candidate regions and anchor blocks, then uses these pre-identified elements as input for subsequent spatial and semantic analysis. This staged approach ensures high value association accuracy through comprehensive multi-dimensional analysis while reducing computational resource consumption by avoiding redundant processing of already-identified elements.

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS11860865B2Analytic systems, methods, and computer-readable media for structured, semi-structured, and unstructured documents
Publication Date: 2024.01.02 TUNGSTEN AUTOMATION CORPORATION
  • US11860865B2 patent drawing
  • US11860865B2 patent drawing
  • US11860865B2 patent drawing

AI summary

A computer system extracts contender values as positively associated with a pre-defined value from a compilation of one or more electronically stored semi-structured document(s) and/or one or more electronically stored unstructured document(s). The computer system performs a multi-dimensional analysis to narrow the universe of contender values from all words on a page of the compilation to the contender value(s) with the highest likelihood of being associated with the pre-defined value. The system's platform allows every user of the system to customize the system according to the user's needs. Various aspects can enable users to mine document stores for information that can be charted, graphed, studied, and compared to help make better decisions.