Hierarchical Document Structuring for Accurate AI Knowledge Bases

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing AI computer models struggle with processing unstructured electronic documents, such as PDFs, due to the lack of clear structural demarcations, leading to reduced accuracy in knowledge base construction and inefficient use of resource-limited models like LLMs.

Innovation Solution

A method to extract a hierarchical structure from unstructured documents by converting them to plain text, identifying candidate structural elements, applying semantic consistency rules, and merging them to generate a structured document suitable for AI operations.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If unstructured electronic documents are used directly for AI model training and operation, then data processing can proceed without additional conversion steps, but the accuracy of knowledge base construction deteriorates due to lack of clear structural demarcations

Engineering Contradiction:
Improveaccuracy of knowledge base constructionVSAvoidcomplexity of document processing
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent applies preliminary action by performing data preparation operations on unstructured electronic documents before they are used for AI model training and operation. This includes converting unstructured documents into structured formats with clear hierarchical organization, extracting meaningful structural elements, and organizing content into chapters and sections. By preparing the data in advance with proper structure, the accuracy of knowledge base construction is improved without adding complexity during the actual AI model operation phase.

Inventive Principle:
Principle #10Preliminary action

2Measurement precision

If unstructured documents are processed without structural organization, then processing speed is maintained, but the precision of AI operations such as classification and prediction deteriorates

Engineering Contradiction:
Improveprecision of AI operationsVSAvoidtime for document processing
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The patent implements preliminary action by organizing unstructured documents into structured formats with clear hierarchical relationships before AI processing. This pre-organization into chapters, sections, and structured elements enables more precise classification and prediction operations by AI models, as the semantic relationships and context are preserved. The time investment for structuring is made once during data preparation, rather than being repeated during each AI operation.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent introduces an intermediary structured document format that acts as a mediator between unstructured electronic documents and AI model processing. This intermediate structured format preserves the semantic meaning and hierarchical relationships of the original document while presenting the information in a way that is optimized for AI operations. The structured format includes organized chapters, sections, and semantic elements that serve as an intermediary representation, enabling more precise AI operations without requiring direct processing of raw unstructured data.

Inventive Principle:
Principle #24Intermediary (Mediator)

3Productivity

If resource-limited AI models like LLMs are used with unstructured documents, then model complexity is reduced, but operational efficiency deteriorates due to inability to effectively process unstructured data

Engineering Contradiction:
Improveoperational efficiency of AI modelsVSAvoidcomplexity of data preparation
Core Design Contradiction:
ProductivityVSDevice complexity

Solution Approach 1:

The patent applies preliminary action by performing comprehensive data preparation and structuring operations on unstructured electronic documents before they are fed to resource-limited AI models. This includes converting unstructured content into organized hierarchical structures with clear semantic relationships, extracting key structural elements, and preparing the data in an optimized format. By doing this preparation work in advance, the operational efficiency of resource-limited models is significantly improved, as they can process well-structured data more effectively without needing to perform complex unstructured data analysis themselves.

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS20250252322A1Constructing Structured Knowledge Bases Using Unstructured Documents
Publication Date: 2025.08.07 INTERNATIONAL BUSINESS MACHINE CORPORATION
  • US20250252322A1 patent drawing
  • US20250252322A1 patent drawing
  • US20250252322A1 patent drawing

AI summary

Mechanisms are provided for extracting a structure from an unstructured electronic document. An unstructured electronic document is converted into a plain text document, and candidate structure identification rules are executed on the plain text document to identify candidate structural elements. Candidate structural element semantic consistency rules are executed on the candidate structural elements to remove candidate structural elements that do not match expected semantic ordering of structural elements and generate a first group of structural elements. A plurality of second groupings of candidate structural elements are generated based on the first grouping, where each second grouping comprises a different combination of structural elements than other second groupings. Candidate structural elements of the second groupings are merged to generate a hierarchical structure for the unstructured electronic document. A structured electronic document is generated, corresponding to the unstructured electronic document, based on the hierarchical structure.