Document Data Extraction Using Spatial Context and Grammar Rules

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current systems fail to accurately extract meaningful information from complex electronic documents like invoices due to their inability to interpret spatial elements such as location, font, and shading, leading to inefficiencies in data extraction and processing.

Innovation Solution

A system comprising a data acquisition engine and a processing engine that uses a backward tracking search algorithm and context-sensitive grammars to extract and interpret data from electronic documents, considering spatial properties like location, font, and shading, and applies rules to determine the meaning and format of extracted data.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Loss of information

If optical character recognition (OCR) is used to recognise text in a PDF file, then text extraction is enabled, but the extracted text lacks context and is not immediately usable for automated processing

Engineering Contradiction:
Improvecontext informationVSAvoidautomated processing usability
Core Design Contradiction:
Loss of informationVSExtent of automation

Solution Approach 1:

The patent transitions from one-dimensional text extraction to multi-dimensional data extraction by incorporating spatial coordinates, font properties, and contextual relationships. Each extracted element includes not just the text content but also its position (x, y coordinates), font characteristics, and hierarchical context within the document structure, enabling automated processing while preserving meaning.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

Solution Approach 2:

The patent introduces an intermediary processing layer between OCR and final data extraction. This layer applies context-sensitive grammar rules and spatial relationship analysis to interpret the extracted text, adding contextual meaning and structured relationships that make the data immediately usable for automated processing while maintaining the original context.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Productivity

If proprietary business to business solutions are used to automate invoice handling, then data extraction is enabled, but high compliance costs and system incompatibility arise

Engineering Contradiction:
Improveinvoice handling automationVSAvoidsystem compatibility
Core Design Contradiction:
ProductivityVSDevice complexity

Solution Approach 1:

The patent creates a universal document processing system that can handle multiple document types (invoices, forms, contracts) using the same context-sensitive grammar engine. The system extracts spatial properties and applies grammatical rules universally across different document formats and layouts, eliminating the need for proprietary solutions for each document type and reducing compliance costs.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The patent changes the fundamental parameters of document processing from format-specific recognition to parameter-based extraction. By focusing on spatial coordinates, font properties, and grammatical relationships rather than document format, the system achieves high productivity across diverse document types without requiring complex proprietary solutions for each format.

Inventive Principle:
Principle #35Parameter changes

3Loss of information

If context-sensitive grammar and geographic/word association are used for data extraction, then information extraction is enabled, but inability to interpret spatial elements like location, font, and shading reduces accuracy

Engineering Contradiction:
Improvemeaningful informationVSAvoiddata extraction accuracy
Core Design Contradiction:
Loss of informationVSMeasurement precision

Solution Approach 1:

The patent segments the document analysis into distinct components: spatial property extraction (coordinates, font, shading), grammatical structure analysis, and meaning interpretation. By separating these functions, the system can accurately capture spatial elements and their relationships, then apply context-sensitive grammar to interpret their meaning, significantly improving extraction accuracy.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent performs preliminary extraction of spatial properties (location, font, shading) before applying grammatical analysis. This preliminary action captures the contextual cues that will be used by the context-sensitive grammar rules to accurately interpret the meaning of extracted elements, improving overall data extraction accuracy.

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS9558295B2System for data extraction and processing
Publication Date: 2017.01.31 ADVANCED BUSINESS SOFTWARE AND SOLUTIONS LTD
  • US9558295B2 patent drawing
  • US9558295B2 patent drawing
  • US9558295B2 patent drawing

AI summary

A system for extracting and interpreting information received in a human-readable format, typically PDF, assigning field tags to the extracted information and transferring the tagged information to a data processing system so that the tagged information can be uploaded to the system automatically. The system provides an incoming document with a time stamp to enable differentiation of the incoming document from other incoming documents, then, the incoming document may be spilt into sections to enable processing of each section individually. Subsequently, context and information are extracted by allowing a processing engine to apply a predetermined set of rules so that the extracted information to be ascribed meaning and assigned a field tag depending on its meaning. The system generates an editable output which is sent to a user.