Line-Based Text Structuring for OCR Document Field Mapping

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

The inefficiency in managing unstructured documents due to varying document formats and terms across organizations leads to high operational costs and repetitive input processes.

Innovation Solution

A text data structuring method and apparatus that includes a data extraction unit, data processing unit, form classification unit, labeling unit, and relationship identification unit to classify, label, and map text within documents, utilizing OCR, natural language processing, and similarity scoring to improve document management efficiency.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If manual information input is used to process unstructured documents, then data can be entered into systems, but operational costs increase and repetitive input processes occur

Engineering Contradiction:
Improvedocument processing efficiencyVSAvoidtime for repetitive input
Core Design Contradiction:
ProductivityVSLoss of time

Solution Approach 1:

The system performs automatic text extraction, classification, and structuring without human intervention. OCR technology extracts text from images, the classification unit automatically categorizes documents, and the structuring unit organizes data into standardized formats, enabling the system to serve itself and eliminate repetitive manual input

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

Manual mechanical input processes are replaced with automated optical and computational systems. OCR replaces manual text entry, image processing algorithms replace visual inspection, and automated classification replaces manual categorization, substituting mechanical human operations with computational mechanisms

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

2Adaptability or versatility

If different document formats and diverse terms are accommodated, then various document types can be processed, but management complexity increases

Engineering Contradiction:
Improvedocument format compatibilityVSAvoidsystem complexity
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The system is designed to handle multiple document formats and types through a unified processing framework. The classification unit identifies different document types (invoices, contracts, forms), and the structuring unit applies appropriate templates for each type, enabling one system to perform multiple functions without requiring separate processing chains for each format

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The system adapts to different document formats by changing processing parameters dynamically. Based on document classification results, the system adjusts extraction rules, structuring templates, and mapping strategies to match the specific format being processed, allowing flexible adaptation without increasing fundamental system complexity

Inventive Principle:
Principle #35Parameter changes

3Measurement precision

If text extraction and mapping are performed accurately, then data quality improves, but processing time increases

Engineering Contradiction:
Improvetext extraction accuracyVSAvoidprocessing time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The system performs preliminary classification and formatting of extracted text before final mapping. The classification unit pre-organizes extracted content into structured categories, and the structuring unit pre-processes data into standard formats before the mapping unit performs final association, reducing the complexity and time of the final mapping operation

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The processing pipeline is divided into distinct sequential stages: OCR text extraction, image classification, text structuring, and mapping. Each stage handles a specific aspect of processing independently, allowing optimization at each step and preventing the need to reprocess entire documents if errors occur in later stages

Inventive Principle:
Principle #1Segmentation

Data Source

PatentUS12530917B2Text data structuring method and apparatus using line information
Publication Date: 2026.01.20 42 MARU INC
  • US12530917B2 patent drawing
  • US12530917B2 patent drawing
  • US12530917B2 patent drawing

AI summary

A text data structuring apparatus according to the present invention includes: a data extraction unit which extracts text included in an image and position information of the text on the basis of OCR; a data processing unit which extracts line information included in the image by using the text, the position information, and the image; a labeling unit which labels the text as keys or values; and a relationship identification unit which acquires a mapping candidate group including first text, second text, and third text labeled on the basis of the line information, calculates a first similarity score representing meaning similarity between the first text and the third text and a second similarity score representing meaning similarity between the second text and the third text, and decides text to be mapped with the third text among of the first text and the second text.