Document Ingestion Layout Graphs for Heterogeneous Document Extraction

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Organizations face challenges in identifying suitable employees and job seekers due to the lack of standardized formats for resumes, transcripts, and job descriptions, leading to reliance on manual intervention and limited criteria, which hinders efficient data extraction from heterogeneous documents.

Innovation Solution

A system that generates a layout graph from document terms, partitions it into content shapes based on relationship strength, classifies these shapes, and employs extraction models to extract relevant information, using generative AI for classification and pattern recognition, independent of document format.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If manual intervention is used to extract information from heterogeneous documents, then extraction accuracy can be maintained, but productivity is significantly reduced

Engineering Contradiction:
Improveextraction accuracyVSAvoiddata extraction efficiency
Core Design Contradiction:
Measurement precisionVSProductivity

Solution Approach 1:

The patent introduces an intermediary representation layer (structured data format with standardized fields) between the heterogeneous document inputs and the downstream processing systems. This intermediary structure enables automated processing while maintaining data quality, as the standardized format serves as a bridge that translates diverse document types into a common, machine-processable form without requiring manual intervention for each document.

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The system implements a universal extraction framework that handles multiple document types (resumes, transcripts, job descriptions, course descriptions) through a single standardized processing pipeline. The unified data structure and common extraction logic enable the system to process heterogeneous documents automatically, achieving both high productivity and maintained accuracy across diverse input formats.

Inventive Principle:
Principle #6Universality (Multi-functionality)

2Productivity

If standardized formats are imposed on diverse documents, then processing efficiency is improved, but adaptability to different document types is reduced

Engineering Contradiction:
Improveprocessing efficiencyVSAvoiddocument format flexibility
Core Design Contradiction:
ProductivityVSAdaptability or versatility

Solution Approach 1:

The patent segments the document processing into distinct, modular components: document ingestion, content extraction, structured representation, and downstream processing. Each component handles specific aspects of the workflow independently, allowing the system to maintain standardized processing efficiency while adapting to different document types through the flexible segmentation of processing stages.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The system employs parameter-based configuration to adapt the standardized processing pipeline to different document types. By adjusting extraction parameters, field mappings, and processing rules based on document type metadata, the system maintains efficient standardized processing while being versatile enough to handle diverse formats without sacrificing adaptability.

Inventive Principle:
Principle #35Parameter changes

Data Source

PatentUS12361741B1Document ingestion pipeline
Publication Date: 2025.07.15 ASTRUMU INC
  • US12361741B1 patent drawing
  • US12361741B1 patent drawing
  • US12361741B1 patent drawing

AI summary

Embodiments manage data for document ingestion pipelines. Terms in a document may be determined based on a term location within the document. A layout graph that includes term nodes may be generated based on the terms such that each of the terms corresponds to a different term node. Relationships between the term nodes may be determined based on a traversal of the layout graph such that the layout graph may be partitioned into content shapes based on a strength of the relationships. The content shapes may be classified to reduce computational resources for extracting information from the document based on types associated with each content shape such that the types may be associated with extraction models that each support a plurality of different document formats. Extraction models may be employed to extract information from the classified content shapes.