Unstructured Document Analysis Using Font and Structure Data

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing technologies face challenges in efficiently analyzing unstructured documents due to their diverse formats and structures, requiring significant time and cost to process them into a usable form.

Innovation Solution

An unstructured document analysis method, device, and system that acquire unstructured document data including font characteristic data and document structure data, extract text based on this data, classify it using a trained neural network model, and generate answers to content queries related to the extracted text.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If traditional methods are used to process unstructured documents, then analysis can be performed, but processing time and cost are considerable

Engineering Contradiction:
Improveprocessing speedVSAvoidprocessing time
Core Design Contradiction:
ProductivityVSLoss of time

Solution Approach 1:

The patent replaces traditional mechanical/manual document processing methods with an AI-based system that uses neural networks and natural language processing to automatically extract, classify, and analyze text from unstructured documents, significantly reducing processing time and human resource requirements

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Solution Approach 2:

The system enables self-service document analysis by automatically processing unstructured documents without requiring manual intervention. The AI model independently performs text extraction, classification, and query answering, allowing the system to serve itself rather than requiring human operators to manually analyze each document

Inventive Principle:
Principle #25Self-service

2Ease of operation

If unstructured documents are processed into operable form, then analysis becomes possible, but considerable time and cost are required

Engineering Contradiction:
ImproveusabilityVSAvoidprocessing time
Core Design Contradiction:
Ease of operationVSLoss of time

Solution Approach 1:

The patent performs preliminary processing of unstructured documents by automatically extracting text and classifying it into predefined categories before the actual analysis. This preliminary action prepares the documents in advance, making them ready for quick query processing without requiring time-consuming manual preparation

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The system replaces manual document preparation and structuring with automated AI-based text extraction and classification mechanisms, transforming unstructured documents into operable formats instantly rather than requiring considerable human time and effort

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

3Measurement precision

If text extraction is performed on unstructured documents, then analysis can be conducted, but accuracy is limited due to diverse formats

Engineering Contradiction:
Improveextraction accuracyVSAvoidformat diversity
Core Design Contradiction:
Measurement precisionVSAdaptability or versatility

Solution Approach 1:

The patent employs a universal AI-based text extraction system that can handle multiple document formats and structures through a single integrated model. The neural network is trained to recognize and extract text from various formats (PDF, Word, scanned documents, etc.), providing consistent accuracy across different document types without requiring format-specific processing methods

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The system adapts to different document formats by dynamically adjusting extraction parameters and processing techniques. The AI model modifies its behavior based on the detected document characteristics, ensuring high extraction accuracy regardless of whether the document uses standard formatting, scanned images, or complex layouts

Inventive Principle:
Principle #35Parameter changes

Data Source

PatentUS20250190685A1Method, device, and system for analyzing unstructured document
Publication Date: 2025.06.12 42 MARU INC
  • US20250190685A1 patent drawing
  • US20250190685A1 patent drawing
  • US20250190685A1 patent drawing

AI summary

An unstructured document analysis method according to an embodiment includes: operations of acquiring unstructured document data including font characteristic data and document structure data, extracting text included in the unstructured document data on the basis of the font characteristic data or the document structure data, classifying the extracted text into a pre-classified item using a trained neural network model, acquiring a content query related to the content included in the unstructured document data and associated with the pre-classified item, and generating an answer to the content query on the basis of the extracted text classified into the item.