Cognitive Data Processing System for Unstructured Information

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current systems face challenges in processing unstructured data from diverse sources, such as call centers, social media, and regulatory databases, due to increased complexity and regulatory hurdles, especially in industries like Banking and Finance, Life science, and Healthcare, where manual processing is impractical and AI-based solutions have probabilistic outcomes.

Innovation Solution

A processor-implemented method and system that extracts metadata from source documents of various forms, processes them using cognitive processing to convert into Enterprise-to Business (E2B) XML, evaluates accuracy, and decides validity using deterministic and probabilistic approaches, mimicking human decision-making processes.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If manual processing is used for unstructured data from diverse sources, then processing accuracy can be maintained through human judgment, but processing productivity decreases due to increased complexity and volume

Engineering Contradiction:
Improveprocessing accuracyVSAvoidprocessing throughput
Core Design Contradiction:
Measurement precisionVSProductivity

Solution Approach 1:

The patent replaces manual mechanical processing with an automated cognitive processing system that uses machine learning models, natural language processing, and entity recognition algorithms to extract and validate data from unstructured documents, thereby maintaining accuracy while dramatically increasing processing throughput

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Solution Approach 2:

The patent introduces an intermediary validation layer that uses confidence scores and multiple processing models to bridge the gap between automated processing speed and human-level accuracy, allowing the system to handle high volumes while maintaining quality through probabilistic validation mechanisms

Inventive Principle:
Principle #24Intermediary (Mediator)

2Productivity

If rule-based content extraction is used for structured documents, then processing speed is maintained, but adaptability to unstructured data and diverse file formats decreases

Engineering Contradiction:
Improveprocessing speedVSAvoidhandling diverse formats
Core Design Contradiction:
ProductivityVSAdaptability or versatility

Solution Approach 1:

The patent implements a dynamic processing architecture that automatically adapts its extraction methodology based on the input document type, format, and structure. The system transitions between rule-based processing for structured data and cognitive processing for unstructured data, maintaining both speed and adaptability

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

The patent creates a universal processing framework that can handle multiple document formats (PDF, Word, Excel, images, emails, social media posts) through a single integrated system that combines format-specific parsers with general-purpose cognitive processing capabilities

Inventive Principle:
Principle #6Universality (Multi-functionality)

3Adaptability or versatility

If cognitive processing is used to extract data from unstructured documents, then adaptability to diverse sources increases, but processing complexity and computational resources increase

Engineering Contradiction:
Improvehandling unstructured dataVSAvoidsystem complexity
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The patent segments the cognitive processing task into distinct modular components including document preprocessing, entity recognition, relationship extraction, validation, and output generation. Each module can be independently optimized and managed, reducing overall system complexity while maintaining high adaptability

Inventive Principle:
Principle #1Segmentation

4Extent of automation

If AI-based probabilistic solutions are used for data extraction, then processing automation increases, but reliability of outcomes decreases due to probabilistic nature

Engineering Contradiction:
Improveautomation levelVSAvoidoutcome certainty
Core Design Contradiction:
Extent of automationVSReliability

Solution Approach 1:

The patent implements feedback mechanisms where confidence scores from probabilistic AI models are continuously evaluated and used to adjust processing strategies. High-confidence extractions are accepted automatically, while low-confidence cases are routed for additional validation or manual review, thereby maintaining high automation levels while improving reliability

Inventive Principle:
Principle #23Feedback

Data Source

PatentEP3462331B1Automated cognitive processing of source agnostic data
Publication Date: 2021.08.04 TATA CONSULTANCY SERVICES LTD
  • EP3462331B1 patent drawingFigure 1A
  • EP3462331B1 patent drawingFigure 1B
  • EP3462331B1 patent drawingFigure 2

AI summary

With the scale of information available today along with the existing diverse channels of communication, manual processing of information is becoming a challenge and companies across industries are under tremendous pressure to lower transactional costs. Artificial Intelligence based automation of business transactions has seen regulatory hurdles due to probabilistic nature of the outcome. The main challenge lies in processing of transactions with unstructured information. Systems and methods of the present disclosure uses deterministic as well as probabilistic approaches to maximize accuracy. The larger use of deterministic approach with configurable components and ontologies helps to improvise accuracy, precision and reduce recall. The probabilistic approach is used when there is absence of quality information or less information for learning. Also, confidence indicators are provided at attribute level of data being processed and at each decision level.