Data Collection Device Using OCR for Learning Data Extraction
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Collecting learning data for machine learning models, particularly for natural language processing tasks, is challenging due to the need for dedicated loggers, high installation and creation costs, and potential deviations between manually created pseudo data and actual data.
Innovation Solution
A data collection device that acquires data from shared storage areas, determines the file format, and extracts text using suitable libraries or OCR methods, storing it as learning data for machine learning models.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Ease of manufacture
If learning data is manually created, then data creation flexibility is improved, but data creation cost increases enormously
Solution Approach 1:
The patent uses OCR technology to copy text from images and documents, creating learning data without manual typing. The system captures screenshots or images containing text, applies OCR to extract the text, and stores it as learning data, thereby eliminating the need for manual data creation while maintaining data quality and flexibility.
2Ease of operation
If a dedicated logger is installed to collect learning data, then data collection capability is improved, but installation cost and security complexity increase
Solution Approach 1:
The patent implements a self-service data collection mechanism where the system automatically captures screenshots, extracts text using OCR, and stores learning data without requiring dedicated logger software or complex installation. The data collection process is integrated into the existing system operations, eliminating the need for separate data collection infrastructure.
3Loss of time
If manually created pseudo data is used, then data collection effort is reduced, but deviation from actual data increases
Solution Approach 1:
The patent copies actual data from screenshots and documents captured during normal system operation, rather than creating pseudo data manually. This ensures the learning data accurately reflects real-world data patterns and distributions, maintaining high fidelity to actual data while minimizing manual effort through automated OCR extraction.
Data Source
AI summary
A data collection device according to an embodiment includes: an acquisition unit that acquires data when the data is stored in a shared storage area available to one or more users; a determination unit that determines whether or not a format of the data acquired by the acquisition unit is a format in which text included in the data can be extracted by a predetermined library; an extraction unit that extracts the text included in the data by a text extraction method according to a determination result determined by the determination unit; and a storage unit that stores the text extracted by the extraction unit in a database as learning data for a machine learning model that implements a natural language processing task.


