Multi-Type Document Processing With Layout-Aware Thought Trees

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing solutions for processing structured format documents and tabular data in large-scale machine learning models suffer from high manual workload, poor transferability, and subpar performance due to loss of original formatting and positional information.

Innovation Solution

A computer-implemented method that organizes structured data into text inputs using tree of thoughts and layout analysis, enabling efficient automation workflows by combining these techniques with large language models.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If manual processing methods are used for structured format documents, then processing accuracy can be maintained, but manual workload increases significantly

Engineering Contradiction:
Improveprocessing accuracyVSAvoidmanual workload
Core Design Contradiction:
Measurement precisionVSEase of operation

Solution Approach 1:

The system enables automated self-processing of structured documents through the tree of thoughts framework, where the language model autonomously analyzes document layouts, extracts tabular data, and performs processing tasks without requiring manual intervention, thereby reducing manual workload while maintaining accuracy through automated validation mechanisms

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The patent introduces a tree of thoughts intermediary structure that mediates between the raw document input and the final processing output. This intermediary organizes the complex processing workflow into structured thought nodes, enabling automated accurate processing while reducing the need for manual intervention in document handling

Inventive Principle:
Principle #24Intermediary (Mediator)

2Productivity

If existing processing methods are applied to structured documents, then some processing capability is achieved, but transferability to different application scenarios is poor

Engineering Contradiction:
Improveprocessing capabilityVSAvoidtransferability
Core Design Contradiction:
ProductivityVSAdaptability or versatility

Solution Approach 1:

The tree of thoughts framework provides a universal processing architecture that can handle multiple document types and application scenarios through the same core mechanisms. The system extracts and processes tabular data from various structured formats using consistent methods, enabling high transferability across different domains while maintaining strong processing capability

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The system employs dynamic adaptation within the tree of thoughts structure, where processing pathways can be flexibly adjusted based on the specific document type and application scenario. This dynamic nature allows the same core system to effectively transfer across different scenarios while maintaining high processing productivity

Inventive Principle:
Principle #15Dynamics

3Speed

If original formatting and positional information are discarded during processing, then processing speed increases, but performance deteriorates due to loss of critical information

Engineering Contradiction:
Improveprocessing speedVSAvoidperformance
Core Design Contradiction:
SpeedVSReliability

Solution Approach 1:

The system performs preliminary extraction and preservation of formatting and positional information during the document analysis phase, storing this metadata alongside the extracted content. This preliminary action ensures that critical structural information is retained before processing begins, enabling both high-speed processing and maintained performance

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent implements a nested structure where formatting and positional information is embedded within the tree of thoughts nodes alongside the extracted content. This nesting allows the system to process content at high speed while simultaneously maintaining access to formatting and positional data when needed, resolving the contradiction between speed and performance

Inventive Principle:
Principle #7Nested doll (Nesting)

4Measurement precision

If complex processing workflows are implemented for structured documents, then processing accuracy improves, but device complexity increases

Engineering Contradiction:
Improveprocessing accuracyVSAvoidsystem complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent segments the complex processing workflow into discrete tree of thoughts nodes, each handling specific processing tasks. This segmentation breaks down complex accuracy-critical operations into manageable, modular units that can be independently processed and validated, improving accuracy while keeping system complexity manageable through structured organization

Inventive Principle:
Principle #1Segmentation

Data Source

PatentUS20250322684A1Processing multi-type document for machine learning comprehension
Publication Date: 2025.10.16 INTERNATIONAL BUSINESS MACHINE CORPORATION
  • US20250322684A1 patent drawing
  • US20250322684A1 patent drawing
  • US20250322684A1 patent drawing

AI summary

A computer-implemented method includes receiving a digital image of a document and a workflow describing an automation task. The method also include converting the digital image of the document into rich text that includes layout information in the document. The method further includes creating, based on the rich text and the workflow, a tree of thoughts that includes nodes and edges connecting the nodes and that binds at least some nodes representing the rich text with a task node representing the automation task. The method also includes converting the nodes and edges of the tree of thoughts into a natural language text. The method further includes inputting the natural language text into a language machine learning model with attention given to a token in the natural language text representing the task node. The language machine learning model, in response, outputs a result of completing the automation task.