Data Collection Device Using OCR for Learning Data Extraction

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Collecting learning data for machine learning models, particularly for natural language processing tasks, is challenging due to the need for dedicated loggers, high installation and creation costs, and potential deviations between manually created pseudo data and actual data.

Innovation Solution

A data collection device that acquires data from shared storage areas, determines the file format, and extracts text using suitable libraries or OCR methods, storing it as learning data for machine learning models.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Ease of manufacture

If learning data is manually created, then data creation flexibility is improved, but data creation cost increases enormously

Engineering Contradiction:
Improvedata creation flexibilityVSAvoiddata creation cost
Core Design Contradiction:
Ease of manufactureVSQuantity of substance

Solution Approach 1:

The patent uses OCR technology to copy text from images and documents, creating learning data without manual typing. The system captures screenshots or images containing text, applies OCR to extract the text, and stores it as learning data, thereby eliminating the need for manual data creation while maintaining data quality and flexibility.

Inventive Principle:
Principle #26Copying

2Ease of operation

If a dedicated logger is installed to collect learning data, then data collection capability is improved, but installation cost and security complexity increase

Engineering Contradiction:
Improvedata collection capabilityVSAvoidinstallation cost and security complexity
Core Design Contradiction:
Ease of operationVSDevice complexity

Solution Approach 1:

The patent implements a self-service data collection mechanism where the system automatically captures screenshots, extracts text using OCR, and stores learning data without requiring dedicated logger software or complex installation. The data collection process is integrated into the existing system operations, eliminating the need for separate data collection infrastructure.

Inventive Principle:
Principle #25Self-service

3Loss of time

If manually created pseudo data is used, then data collection effort is reduced, but deviation from actual data increases

Engineering Contradiction:
Improvedata collection effortVSAvoiddata accuracy
Core Design Contradiction:
Loss of timeVSMeasurement precision

Solution Approach 1:

The patent copies actual data from screenshots and documents captured during normal system operation, rather than creating pseudo data manually. This ensures the learning data accurately reflects real-world data patterns and distributions, maintaining high fidelity to actual data while minimizing manual effort through automated OCR extraction.

Inventive Principle:
Principle #26Copying

Data Source

PatentUS20250094453A1Data collection apparatus, data collection method, and program
Publication Date: 2025.03.20 NT T INC
  • US20250094453A1 patent drawing
  • US20250094453A1 patent drawing
  • US20250094453A1 patent drawing

AI summary

A data collection device according to an embodiment includes: an acquisition unit that acquires data when the data is stored in a shared storage area available to one or more users; a determination unit that determines whether or not a format of the data acquired by the acquisition unit is a format in which text included in the data can be extracted by a predetermined library; an extraction unit that extracts the text included in the data by a text extraction method according to a determination result determined by the determination unit; and a storage unit that stores the text extracted by the extraction unit in a database as learning data for a machine learning model that implements a natural language processing task.