Document Indexing via Statistical Knowledge Graphs

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing document processing systems rely on supervised approaches with uniform text parameters, leading to inaccurate content extraction due to neglect of statistical variations in content, resulting in inefficient handling of large datasets in industries like E-commerce, Education, Pharma, and IT.

Innovation Solution

A processor-implemented method and system for document processing that pre-processes documents to identify unique words and phrases, correlates them with topics based on word patterns, and builds a knowledge graph for accurate indexing and retrieval, enabling the identification of relevant sections and patterns within unstructured documents.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Ease of manufacture

If supervised approach with uniform text parameters is used for document processing, then the system structure is simple and easy to implement, but the content extraction accuracy deteriorates due to neglect of statistical variations

Engineering Contradiction:
Improvesystem implementation simplicityVSAvoidcontent extraction accuracy
Core Design Contradiction:
Ease of manufactureVSMeasurement precision

Solution Approach 1:

The patent changes the processing parameters from uniform fixed parameters to dynamic statistical parameters. It calculates frequency distributions, mean, standard deviation, and other statistical metrics for text parameters like font size, font style, and spacing. This allows the system to adapt to variations in document content while maintaining a relatively simple overall structure.

Inventive Principle:
Principle #35Parameter changes

Solution Approach 2:

The system transitions from static uniform parameters to dynamic statistical parameters that adapt to the specific document being processed. By calculating statistical characteristics (mean, standard deviation, frequency distributions) for each document, the system dynamically adjusts its processing parameters to match the actual content variations, thereby improving extraction accuracy.

Inventive Principle:
Principle #15Dynamics

2Productivity

If section-based data identification is performed without considering statistical variations, then the processing speed is fast, but the content extraction accuracy deteriorates significantly

Engineering Contradiction:
Improvedata processing speedVSAvoidcontent extraction accuracy
Core Design Contradiction:
ProductivityVSMeasurement precision

Solution Approach 1:

The patent applies partial statistical analysis by focusing on key text parameters (font size, font style, spacing, line height) rather than analyzing all possible document features. This selective approach maintains processing speed while capturing the essential statistical variations needed for accurate content extraction.

Inventive Principle:
Principle #16Partial or excessive action

Solution Approach 2:

The system changes from ignoring statistical parameters to using calculated statistical parameters (frequency distributions, mean, standard deviation) for content identification. This allows the system to maintain fast processing by using pre-calculated statistical metrics rather than performing complex real-time analysis.

Inventive Principle:
Principle #35Parameter changes

3Ease of operation

If uniform text parameters are enforced for all documents, then the processing methodology is consistent and simple, but the adaptability to different document types and styles deteriorates

Engineering Contradiction:
Improveprocessing methodology consistencyVSAvoiddocument type adaptability
Core Design Contradiction:
Ease of operationVSAdaptability or versatility

Solution Approach 1:

The system transitions from static uniform parameters to dynamic statistical parameters that automatically adapt to different document types. By calculating statistical characteristics for each document, the system maintains a consistent processing methodology while adapting to variations in document styles, formats, and content structures.

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

The patent creates a universal processing framework that can handle multiple document types and styles. The statistical parameter calculation approach works across different document formats and content types, making the system versatile while maintaining methodological consistency through the unified statistical analysis process.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Data Source

PatentUS11775549B2Method and system for document indexing and retrieval
Publication Date: 2023.10.03 TATA CONSULTANCY SERVICES LTD
  • US11775549B2 patent drawing
  • US11775549B2 patent drawing
  • US11775549B2 patent drawing

AI summary

Existing systems for document processing are either based on a supervised approach using annotated tags, and these systems identify section-based data from the unstructured documents without considering the statistical variations in content, which results in highly inaccurate content extraction. The disclosure herein generally relates to document processing, and, more particularly, to method and system for document indexing and retrieval. The system provides a mechanism to correlate unique words in a document with different topics identified in the document, based on a word pattern identified from the document. The correlations are captured in a knowledge graph, and can be further used in applications such as but not limited to document retrieval.