AI OCR Data Archiving for Text and Table Extraction

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional OCR-based unstructured data conversion devices fail to efficiently extract, identify, and utilize texts and tables in image and PDF files, lack data management and linkage systems, and require manual data collection and analysis, leading to increased worker hours and difficulty in managing big data systematically.

Innovation Solution

An integrated management device utilizing AI OCR, a processor with modules for extraction, preprocessing, search, and display, and a big data platform to automate data collection, analysis, and management, enabling efficient extraction, identification, and utilization of texts and tables, and systematic data management.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If conventional OCR-based unstructured data conversion devices are used to store extracted texts in a relational database, then text extraction is performed, but the function of extracting, identifying, and utilizing texts and tables in image files and PDF files is not achieved

Engineering Contradiction:
Improvedata extraction accuracyVSAvoiddata processing capability
Core Design Contradiction:
Measurement precisionVSAdaptability or versatility

Solution Approach 1:

The patent applies multi-functionality by integrating multiple data processing capabilities into a single system. The device not only performs text extraction using AI OCR but also simultaneously extracts tables from images and PDF files, identifies data types, and converts various formats (images, PDFs, HTML, Excel) into structured relational database formats. This universal approach resolves the contradiction by enabling the system to handle diverse data types while maintaining high extraction accuracy.

Inventive Principle:
Principle #6Universality (Multi-functionality)

2Measurement precision

If conventional OCR-based devices are used, then text extraction is performed, but data management and linkage systems are lacking

Engineering Contradiction:
Improvetext extractionVSAvoiddata management system
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent merges previously separate functions into a unified system. It combines text extraction, table extraction, data type identification, data management, and data linkage capabilities into a single integrated device. The system consolidates data from multiple sources (images, PDFs, HTML, Excel) and stores them in a relational database with proper linkage mechanisms, thereby resolving the contradiction by achieving comprehensive data management without proportionally increasing complexity.

Inventive Principle:
Principle #5Merging (Combining)

3Productivity

If manual data collection and analysis is performed, then data processing is done, but work hours of workers increase

Engineering Contradiction:
Improvedata processing outputVSAvoidwork time
Core Design Contradiction:
ProductivityVSLoss of time

Solution Approach 1:

The patent implements self-service automation where the system performs data collection, extraction, identification, and conversion autonomously without human intervention. The AI OCR engine automatically extracts text and tables from various formats, the data type identification module automatically categorizes data, and the conversion module automatically transforms data into relational database formats. This automation resolves the contradiction by maintaining high productivity while eliminating the time loss associated with manual processing.

Inventive Principle:
Principle #25Self-service

4Measurement precision

If experts in each field are hired to process data, then data processing quality is maintained, but it is difficult to hire experts and systematically manage big data

Engineering Contradiction:
Improvedata processing qualityVSAvoiddata management system
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent replaces the mechanical system of hiring and managing human experts with an automated AI-based system. The AI OCR engine, data type identification module, and conversion module collectively perform functions that would traditionally require multiple specialized experts. The system automatically handles diverse data formats, maintains processing quality through intelligent algorithms, and eliminates the complexity of expert recruitment and management while systematically managing big data.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Data Source

PatentUS20260064950A1Apparatus, system and method for managing archiving for automated big data collection and public data linkage based on artificial intelligence
Publication Date: 2026.03.05 ARCHIVSOFT CO LTD
  • US20260064950A1 patent drawing
  • US20260064950A1 patent drawing
  • US20260064950A1 patent drawing

AI summary

Disclosed is efficiently perform a function of accurately extracting, identifying, and utilizing texts and tables in image files and PDF files.