Document Type Identification for Field Value Extraction
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing technologies require definition information for each document category to automatically extract field values from documents, which can be cumbersome and inefficient.
Innovation Solution
An information processing apparatus that determines a document type based on its title and uses pre-defined definition information for that type to extract field values, eliminating the need for category-specific rules.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If definition information is prepared for each document category, then field value extraction accuracy is improved, but system complexity and preparation workload increase
Solution Approach 1:
The patent creates a universal definition information structure that can handle multiple document categories through a standardized schema. Instead of creating separate extraction rules for each document type, a single unified definition framework is established that adapts to different categories (invoices, receipts, contracts, etc.) through consistent field identification patterns, thereby reducing system complexity while maintaining extraction accuracy across diverse document types
Solution Approach 2:
The patent segments the document processing system into distinct modular components: document type identification module, definition information storage module, and field value extraction module. This segmentation allows each component to operate independently with well-defined interfaces, reducing overall system complexity while enabling accurate extraction by matching document types with appropriate definition templates
2Measurement precision
If definition information is prepared for each document category, then field value extraction accuracy is improved, but preparation time and resources increase
Solution Approach 1:
The patent performs preliminary document type identification and classification before the actual field value extraction process. By pre-categorizing documents and loading only the relevant definition information for the identified document type, the system avoids the time-consuming process of preparing and processing definition information for all possible document categories, thereby reducing preparation time while maintaining extraction accuracy
Solution Approach 2:
The patent dynamically changes the parameter set (definition information) based on the detected document type. Instead of using a fixed, comprehensive set of definition information for all categories, the system adjusts and loads only the specific parameters and fields relevant to the current document type, reducing preparation time and resource requirements while maintaining high extraction accuracy for the specific category
3Measurement precision
If category-specific rules are used for each document type, then extraction precision is improved, but flexibility and ease of operation decrease
Solution Approach 1:
The patent implements a universal definition information framework that maintains high extraction precision across different document categories without requiring users to manage separate category-specific rules. The system automatically selects and applies the appropriate definition template based on document type identification, providing a unified, easy-to-operate interface while preserving category-specific extraction accuracy through standardized field mapping
Data Source
AI summary
An information processing apparatus includes a processor. The processor is configured to: determine a document type of a document by using a title of the document, the document being classified as the determined document type, the title representing a category of the document and being extracted from a read image of the document; and extract a field value from the document by using an item of definition information prepared in accordance with the determined document type from among items of definition information. The definition information is prepared for each document type and defines a rule for extracting a field value from a document.


