Unified Data Extraction Platform for Diverse Document Processing

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional data extraction and processing systems are labor-intensive, prone to errors, and inefficient in handling diverse document types, often converting data into proprietary formats that are not easily transferable across different devices and contact management systems, and struggle to recognize physical relationships within documents.

Innovation Solution

A unified data extraction platform utilizing a micro-service-based architecture with an API gateway, ingestion, extraction, classification, validation, and transformation engines, which ingests documents from various sources, applies Optical Content Recognition (OCR), classifies data using machine learning, and converts data into standard formats for seamless transfer across applications.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If conventional scanned document recognition systems are used to extract information from diverse document types, then data extraction can be performed, but the systems are labor-intensive and prone to errors

Engineering Contradiction:
Improvedata extraction efficiencyVSAvoidextraction accuracy
Core Design Contradiction:
ProductivityVSReliability

Solution Approach 1:

The system segments the data extraction process into distinct modular components: an extraction engine that identifies and extracts data, a validation engine that verifies extracted data against predefined rules, and a classification engine that categorizes document types. This segmentation allows each component to specialize in specific tasks, improving both efficiency and accuracy while reducing labor-intensive manual intervention.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The system implements feedback mechanisms where the validation engine continuously verifies extracted data against predefined business rules and constraints. When validation fails, the system provides feedback to the extraction engine to correct errors. This closed-loop feedback system significantly reduces extraction errors and improves reliability without requiring manual review of every document.

Inventive Principle:
Principle #23Feedback

2Ease of manufacture

If conventional systems convert information into proprietary formats, then data can be processed, but the formats are not easily transferable to other contact management systems

Engineering Contradiction:
Improvedata processing capabilityVSAvoidformat compatibility
Core Design Contradiction:
Ease of manufactureVSAdaptability or versatility

Solution Approach 1:

The system implements a universal data output format that is designed to be compatible with multiple contact management systems and applications. The transformation engine converts extracted data into standardized formats (such as JSON, XML, or CSV) that can be easily integrated with various downstream systems, eliminating the need for proprietary formats and improving adaptability across different platforms.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The system introduces a transformation engine as an intermediary component between the extraction engine and downstream applications. This intermediary translates extracted data into universally compatible formats, acting as a mediator that enables seamless data transfer between different systems without requiring system-specific proprietary formats, thereby enhancing versatility and interoperability.

Inventive Principle:
Principle #24Intermediary (Mediator)

3Ease of operation

If conventional systems are tethered to a particular electronic device, then processing can be performed, but the document cannot be readily processed on dependent devices

Engineering Contradiction:
Improvedevice-specific processingVSAvoidcross-device processing capability
Core Design Contradiction:
Ease of operationVSAdaptability or versatility

Solution Approach 1:

The system is designed as a device-agnostic platform that can process documents on any electronic device with appropriate computing resources. The extraction, validation, and classification engines are implemented as software modules that can run on various devices (desktops, laptops, mobile devices, servers), enabling cross-device processing capability while maintaining ease of operation through consistent user interfaces and workflows.

Inventive Principle:
Principle #6Universality (Multi-functionality)

4Device complexity

If conventional systems cannot determine physical relationships among items in documents with varied layouts, then processing is simplified, but layout recognition fails for diverse document types

Engineering Contradiction:
Improveprocessing simplicityVSAvoidlayout recognition capability
Core Design Contradiction:
Device complexityVSAdaptability or versatility

Solution Approach 1:

The classification engine dynamically adapts to different document layouts by learning from training data representing various document types and their physical relationships. The system uses machine learning algorithms that can identify and adapt to different layout patterns (such as invoice layouts, receipt layouts, form layouts) without requiring manual configuration for each layout type, thereby achieving versatile layout recognition while maintaining processing simplicity through automated adaptation.

Inventive Principle:
Principle #15Dynamics

Data Source

PatentUS11934421B2Unified extraction platform for optimized data extraction and processing
Publication Date: 2024.03.19 COGNIZANT TECH SOLUTIONS INDIA PVT LTD
  • US11934421B2 patent drawing
  • US11934421B2 patent drawing
  • US11934421B2 patent drawing

AI summary

The present invention provides for a system and a method for optimized data extraction of different document types. First digitised data is extracted from ingested documents based on extraction rules and is classified into first classified data based on pre-defined rules. Confidence score is assigned to first classified data based on comparison of first classified data with pre-defined data. A second digitised data is extracted from classified document types corresponding to first classified data via a tool selected from multiple integrated tools based on extraction rules. An extraction score is determined for second digitised data. Classified document types are validated based on pre-defined requirements. In the event the pre-determined requirements are met the confidence score and the extraction score are compared with pre-defined parameters. If the result is above a pre-determined threshold the second digitized data is transmitted as executable files to applications for execution.