Line-Based Text Structuring for OCR Key-Value Mapping

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

The processing of unstructured documents between individuals and organizations is inefficient due to the high cost of repetitive input processes caused by different document formats and diverse terms, necessitating improved methods for text data structuring.

Innovation Solution

A text data structuring apparatus and method that includes data extraction, form classification, text labeling, relationship identification, and misrecognition correction, utilizing OCR, natural language processing, and deep learning models to efficiently manage and operate image documents.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If manual text input is performed for unstructured documents, then data accuracy can be maintained, but processing time and operational costs increase significantly

Engineering Contradiction:
Improvedata accuracyVSAvoidprocessing time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The patent replaces manual mechanical text input with an automated OCR-based text extraction system. The OCR unit extracts text from image documents automatically, eliminating the need for manual typing while maintaining data accuracy through subsequent processing units that validate and structure the extracted text.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Solution Approach 2:

The patent introduces intermediate processing units (text extraction, validation, and structuring units) that act as mediators between the raw OCR output and the final structured data. These intermediaries ensure data accuracy by validating extracted text against predefined formats and rules, while significantly reducing processing time compared to manual input.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Reliability

If different document formats are processed manually, then each format can be handled accurately, but the complexity of the processing system increases

Engineering Contradiction:
Improveformat handling accuracyVSAvoidprocessing system complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The patent implements a universal text extraction and structuring system that can handle multiple document formats through a single integrated apparatus. The OCR unit and subsequent processing units are designed to work with various document types (invoices, receipts, forms) without requiring separate manual processing procedures for each format, thereby reducing system complexity while maintaining format-specific accuracy.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The patent uses parameter-based processing where the system adjusts extraction and validation parameters dynamically based on the detected document format. The text structuring unit modifies processing parameters according to the document type, enabling accurate handling of different formats through a unified system rather than requiring separate complex processing chains for each format.

Inventive Principle:
Principle #35Parameter changes

3Productivity

If OCR is used for text extraction from image documents, then processing speed increases, but text recognition accuracy decreases due to misrecognition

Engineering Contradiction:
Improveprocessing speedVSAvoidtext recognition accuracy
Core Design Contradiction:
ProductivityVSMeasurement precision

Solution Approach 1:

The patent implements a feedback mechanism where the text validation unit checks OCR-extracted text against predefined formats, patterns, and business rules. When misrecognition is detected (such as invalid date formats, incorrect numerical patterns, or format violations), the system flags these errors for correction, providing feedback that improves overall text recognition accuracy while maintaining high processing speed through automated validation.

Inventive Principle:
Principle #23Feedback

Solution Approach 2:

The patent performs preliminary validation and format checking on OCR-extracted text before final processing. The text validation unit预先 checks extracted text against expected formats and patterns, identifying potential misrecognition errors early in the process. This preliminary action allows for quick correction of common OCR errors while maintaining the high processing speed advantage of automated extraction.

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentEP4325382B1Text data structuring method and apparatus using line information
Publication Date: 2025.11.12 42 MARU INC
  • EP4325382B1 patent drawingFigure 1
  • EP4325382B1 patent drawingFigure 2(a)
  • EP4325382B1 patent drawingFigure 2(b)

AI summary

A text data structuring apparatus according to the present invention includes: a data extraction unit which extracts text included in an image and position information of the text on the basis of OCR; a data processing unit which extracts line information included in the image by using the text, the position information, and the image; a labeling unit which labels the text as keys or values; and a relationship identification unit which acquires a mapping candidate group including first text, second text, and third text labeled on the basis of the line information, calculates a first similarity score representing meaning similarity between the first text and the third text and a second similarity score representing meaning similarity between the second text and the third text, and decides text to be mapped with the third text among of the first text and the second text.