Information Representation Structure Analysis for Atypical Document Extraction

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing document analysis systems fail to efficiently extract target information from atypical documents, as they do not utilize non-textual information such as control characters, spaces, tabs, and external metadata, limiting their ability to interpret document structure accurately.

Innovation Solution

An information representation structure analysis device that identifies information representation grammar and support information types, using templates to generate patterns for extracting target information from atypical documents, incorporating textual and structural data like tables, spaces, and external metadata.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If only text and two-dimensional arrangement information are used for document layout recognition, then the system is simple to implement, but it cannot efficiently extract target information from atypical documents

Engineering Contradiction:
Improveability to extract information from atypical documentsVSAvoidsystem complexity
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The patent segments the document information into multiple types: text information, two-dimensional arrangement information, control character information (tables, spaces, tabs, HTML tags), and external information (headers, footers, invisible document information). This segmentation allows the system to process and utilize diverse information types independently, improving adaptability to atypical documents while managing complexity through modular processing of each information segment.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent creates a universal document analysis system that can handle multiple types of information (text, layout, control characters, external metadata) through a unified approach. The system uses information representation grammar that can accommodate various document formats and structures, making it versatile for different document types while maintaining a consistent processing framework.

Inventive Principle:
Principle #6Universality (Multi-functionality)

2Measurement precision

If multiple types of information (control characters, external metadata, invisible information) are utilized, then extraction accuracy from atypical documents improves, but information processing complexity increases

Engineering Contradiction:
Improveinformation extraction accuracyVSAvoidinformation processing complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent introduces information representation grammar as an intermediary layer between the raw document information and the extraction process. This grammar serves as a mediator that standardizes and structures diverse information types (text, control characters, external metadata), making them easier to process while improving extraction accuracy. The grammar acts as a bridge that manages complexity by providing a unified representation framework.

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The patent changes the parameters of information representation by introducing structured grammar rules that define how different information types relate to each other. By transforming raw, unstructured document information into grammar-based representations with defined relationships and hierarchies, the system improves extraction accuracy while managing complexity through parameterized information structures.

Inventive Principle:
Principle #35Parameter changes

Data Source

PatentUS20240184985A1Information representation structure analysis device, and information representation structure analysis method
Publication Date: 2024.06.06 HITACHI LTD
  • US20240184985A1 patent drawing
  • US20240184985A1 patent drawing
  • US20240184985A1 patent drawing

AI summary

An information representation structure analysis device: stores an information representation template for each combination of an information representation grammar and a support information type, the information representation template being a template used to generate an information representation pattern being a program code for implementing a function of extracting an extraction target, the support information type being a category of support information being information used in extraction of the extraction target; identifies the information representation template to be used to generate the information representation pattern for extraction of the extraction target from an atypycal document, based on the information representation grammar and the support information type identified for the information representation; and generates the information representation pattern by applying the extraction target and basis information to the identified information representation template, the basis information being a basis in extraction of the extraction target from the information representation.