Unstructured Document Indexing via Text Stream Segmentation

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Traditional document processing and searching technologies are ineffective for non-text based documents like PDFs, which often contain multiple sections and pages, requiring users to manually search within these documents to find specific information.

Innovation Solution

A method for analyzing and indexing unstructured or semistructured documents by converting them into text streams, identifying textual contents, logical sections, and associating context information, and saving the indexing in a data storage device, allowing for efficient retrieval and search functionality.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If traditional indexing and searching techniques are used for non-text based documents, then the documents can be stored and accessed, but the searching effectiveness is poor and users must manually search within documents

Engineering Contradiction:
Improvesearching effectivenessVSAvoidmanual searching time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The patent segments non-text documents into multiple text streams by identifying different content types (main text, captions, tables, figures) and extracting them separately. This segmentation allows each stream to be indexed independently, improving search precision without requiring manual navigation through the entire document.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent introduces an intermediary processing layer that converts non-text document formats into structured text streams with metadata. This intermediary step enables traditional text-based search algorithms to effectively query non-text documents, eliminating the need for manual searching while maintaining search accuracy.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Loss of information

If large documents covering many topics are stored in non-text formats, then comprehensive information is preserved, but users must open documents in specific readers and manually search to find relevant sections

Engineering Contradiction:
Improveinformation completenessVSAvoiddocument access convenience
Core Design Contradiction:
Loss of informationVSEase of operation

Solution Approach 1:

The patent divides large documents into multiple text streams based on content type and logical sections. Each stream can be independently searched and accessed, allowing users to find specific information without opening the entire document in a specialized reader. This maintains information completeness while dramatically improving access convenience.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent adds a new dimension to document access by creating a hierarchical structure with multiple text streams and sections. Users can navigate through sections and streams rather than linearly through the entire document, enabling direct access to relevant information while preserving all original content.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

3Shape

If non-text based documents are stored without conversion to text streams, then original formatting is preserved, but indexing and searching become ineffective

Engineering Contradiction:
Improvedocument formattingVSAvoidindexing effectiveness
Core Design Contradiction:
ShapeVSMeasurement precision

Solution Approach 1:

The patent segments documents into multiple text streams while preserving the association between streams and original formatting elements. Each stream is indexed separately with metadata indicating its source and formatting context, enabling effective searching without losing formatting information.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent implements a nested structure where text streams are nested within sections, which are nested within the overall document structure. This nested organization allows indexing at multiple levels while maintaining the relationship between extracted text and original formatting, improving indexing effectiveness without sacrificing formatting preservation.

Inventive Principle:
Principle #7Nested doll (Nesting)

Data Source

PatentUS8504553B2Unstructured and semistructured document processing and searching
Publication Date: 2013.08.06 WELLS FARGO BANK NAT ASSOC
  • US8504553B2 patent drawing
  • US8504553B2 patent drawing
  • US8504553B2 patent drawing

AI summary

A method for analyzing and indexing an unstructured or semistructured document according to one embodiment includes receiving an unstructured or semistructured document; converting the document to one or more text streams; analyzing the one or more text streams for identifying textual contents of the document; analyzing the one or more text streams for identifying logical sections of the document; associating the textual contents with the logical sections; indexing the textual contents and their association with the logical sections; and saving a result of the indexing in a data storage device.