Template-Based Learning Data Generation for Document Layout Adaptation

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing techniques for extracting character strings from document images, particularly those in quasi-standard forms with varying layouts, face challenges in generating diverse learning data required for effective named entity recognition, as they often rely on pre-defined layouts and require extensive labeled training data.

Innovation Solution

An information processing apparatus and system that generates layout data based on template data to create diverse learning data, allowing for the generation of a learned model capable of extracting named entities from document images with different layouts, using a processor and memory to execute instructions for generating layout data and learning data, which includes an image processing apparatus, a learning apparatus, and an information processing server to handle document image tokenization and token string generation.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Ease of manufacture

If character string data is generated by replacing words in pre-defined character strings, then character string data can be generated, but the data does not correspond to quasi-standard forms in various layouts

Engineering Contradiction:
Improveease of generating learning dataVSAvoidadaptability to various layouts
Core Design Contradiction:
Ease of manufactureVSAdaptability or versatility

Solution Approach 1:

The patent segments the document layout into multiple independent regions (header region, body region, footer region, etc.) that can be independently configured. Each region can contain different types of elements (text, tables, images) with customizable layouts. This segmentation allows the learning data generation system to create diverse layout variations by independently arranging elements in different regions, thereby achieving both ease of data generation and adaptability to various layouts.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent creates a universal template-based learning data generation system that can handle multiple quasi-standard form types (invoices, purchase orders, quotes, etc.) through a single platform. The system uses configurable templates that can be adapted to different document types and layouts, allowing one system to serve multiple functions and generate learning data for various document formats without requiring separate generation processes for each layout type.

Inventive Principle:
Principle #6Universality (Multi-functionality)

2Measurement precision

If extensive labeled training data is used for named entity recognition, then recognition accuracy improves, but data preparation time and cost increase

Engineering Contradiction:
Improvenamed entity recognition accuracyVSAvoiddata preparation time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The patent applies preliminary action by automatically generating labeled training data through template-based learning data generation before the actual named entity recognition task. The system pre-configures templates with ground truth labels for various document layouts, so when real documents need processing, the model is already trained on diverse, pre-labeled data. This eliminates the need for time-consuming manual labeling of training data while ensuring high recognition accuracy through comprehensive pre-training.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The system implements self-service by enabling automatic generation of learning data with ground truth labels without requiring manual annotation. The template-based approach automatically creates labeled training examples by defining the expected structure and content of different document regions, allowing the system to generate its own training data autonomously. This self-service capability dramatically reduces the time and resources needed for data preparation while maintaining high labeling accuracy.

Inventive Principle:
Principle #25Self-service

Data Source

PatentUS20240249546A1Information processing apparatus, information processing system, and storage medium
Publication Date: 2024.07.25 CANON KK
  • US20240249546A1 patent drawing
  • US20240249546A1 patent drawing
  • US20240249546A1 patent drawing

AI summary

Learning data is generated so as to correspond to documents in various layouts. An information processing apparatus generates layout data indicating a layout of a character string based on template data to define a layout of a document, and generates learning data based on the generated layout data, wherein the generated learning data are used for generating a learned model that extracts a named entity from a document image.