AI Document Processor for Unstructured Data Extraction

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing systems face difficulties in extracting information from unstructured electronic documents, as they lack well-defined data models, making it challenging to programmatically parse and utilize the information for operations in enterprise systems.

Innovation Solution

A document processing system employing AI and machine learning techniques, including optical character recognition (OCR), natural language processing (NLP), and domain-specific models, to convert and extract information from structured and unstructured documents, enabling automatic execution of processes and discrepancy resolution.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If traditional programming methods are used to extract information from structured documents, then extraction is relatively easy and efficient, but this approach fails when dealing with unstructured documents that lack well-defined data models

Engineering Contradiction:
Improvedocument format adaptabilityVSAvoidextraction system complexity
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The patent introduces an intermediary layer consisting of domain models and natural language processing components that mediate between the extraction system and unstructured documents. This intermediary layer translates unstructured document content into structured data formats that enterprise systems can process, enabling the system to handle both structured and unstructured documents through a unified approach.

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The system dynamically changes its processing parameters based on the document type being analyzed. For structured documents, it uses traditional field-based extraction parameters, while for unstructured documents, it switches to NLP-based parameters including entity recognition, relationship extraction, and contextual analysis, allowing adaptive handling of different document formats.

Inventive Principle:
Principle #35Parameter changes

2Measurement precision

If manual processing methods are used for unstructured documents, then information extraction accuracy can be maintained through human judgment, but processing speed and efficiency are significantly reduced

Engineering Contradiction:
Improveinformation extraction accuracyVSAvoiddocument processing speed
Core Design Contradiction:
Measurement precisionVSProductivity

Solution Approach 1:

The patent segments the document processing task into multiple specialized components: text preprocessing, entity recognition, relationship extraction, and validation modules. Each segment handles a specific aspect of information extraction, allowing the system to process documents rapidly while maintaining accuracy through specialized processing at each stage rather than requiring complete manual review.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The system implements feedback mechanisms where extracted information is validated against domain models and business rules, with incorrect or uncertain extractions being flagged for review. This feedback loop enables the system to automatically correct most errors while maintaining high processing speed, and only requires manual intervention for ambiguous cases.

Inventive Principle:
Principle #23Feedback

3Ease of operation

If enterprise systems directly process unstructured documents without conversion, then system integration is simplified, but the systems cannot reliably parse and extract needed information from documents lacking well-defined data models

Engineering Contradiction:
Improvesystem integration easeVSAvoidinformation extraction reliability
Core Design Contradiction:
Ease of operationVSReliability

Solution Approach 1:

The patent creates structured copies of unstructured document content by generating internal representations that mirror the document's information in a standardized format. This copying process transforms unstructured text into structured data objects that enterprise systems can reliably process, maintaining the original document's information while presenting it in a format suitable for automated processing and integration.

Inventive Principle:
Principle #26Copying

Data Source

PatentUS11003796B2Artificial intelligence based document processor
Publication Date: 2021.05.11 ACCENTURE GLOBAL SOLUTIONS LTD
  • US11003796B2 patent drawing
  • US11003796B2 patent drawing
  • US11003796B2 patent drawing

AI summary

An Artificial Intelligence (AI) based document processing system receives a request including one or more documents related to a process to be automatically executed. The information including the fields and an intent required for the process are extracted from one or more of the request and the documents. The required documents and fields are selected based on the intent and a domain model. The required fields are validated using external knowledge and the discrepancies identified therein are resolved. An internal master document is built based on the required fields. The internal master document is employed for the automatic execution of the process which can include a de-identification process or an appeal process.