Template-Based Learning Data Generation for Document Layout Adaptation
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing techniques for extracting character strings from document images, particularly those in quasi-standard forms with varying layouts, face challenges in generating diverse learning data required for effective named entity recognition, as they often rely on pre-defined layouts and require extensive labeled training data.
Innovation Solution
An information processing apparatus and system that generates layout data based on template data to create diverse learning data, allowing for the generation of a learned model capable of extracting named entities from document images with different layouts, using a processor and memory to execute instructions for generating layout data and learning data, which includes an image processing apparatus, a learning apparatus, and an information processing server to handle document image tokenization and token string generation.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Ease of manufacture
If character string data is generated by replacing words in pre-defined character strings, then character string data can be generated, but the data does not correspond to quasi-standard forms in various layouts
Solution Approach 1:
The patent segments the document layout into multiple independent regions (header region, body region, footer region, etc.) that can be independently configured. Each region can contain different types of elements (text, tables, images) with customizable layouts. This segmentation allows the learning data generation system to create diverse layout variations by independently arranging elements in different regions, thereby achieving both ease of data generation and adaptability to various layouts.
Solution Approach 2:
The patent creates a universal template-based learning data generation system that can handle multiple quasi-standard form types (invoices, purchase orders, quotes, etc.) through a single platform. The system uses configurable templates that can be adapted to different document types and layouts, allowing one system to serve multiple functions and generate learning data for various document formats without requiring separate generation processes for each layout type.
2Measurement precision
If extensive labeled training data is used for named entity recognition, then recognition accuracy improves, but data preparation time and cost increase
Solution Approach 1:
The patent applies preliminary action by automatically generating labeled training data through template-based learning data generation before the actual named entity recognition task. The system pre-configures templates with ground truth labels for various document layouts, so when real documents need processing, the model is already trained on diverse, pre-labeled data. This eliminates the need for time-consuming manual labeling of training data while ensuring high recognition accuracy through comprehensive pre-training.
Solution Approach 2:
The system implements self-service by enabling automatic generation of learning data with ground truth labels without requiring manual annotation. The template-based approach automatically creates labeled training examples by defining the expected structure and content of different document regions, allowing the system to generate its own training data autonomously. This self-service capability dramatically reduces the time and resources needed for data preparation while maintaining high labeling accuracy.
Data Source
AI summary
Learning data is generated so as to correspond to documents in various layouts. An information processing apparatus generates layout data indicating a layout of a character string based on template data to define a layout of a document, and generates learning data based on the generated layout data, wherein the generated learning data are used for generating a learned model that extracts a named entity from a document image.


