Headingless Table Sequence Reconstruction for Accurate Attribute Extraction
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
When a single table is divided across multiple pages and imaged out of order, the relationship between the divided tables is unclear, and attribute values from a missing heading row cannot be accurately obtained.
Innovation Solution
An information processing apparatus that groups tables with similar structures and arranges them in consecutive order using attribute values from a heading row, even if the tables lack a heading row, to obtain attribute values from the grouped tables.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If divided tables are grouped by structure alone, then tables with same structure can be identified, but the original consecutive order of tables is lost
Solution Approach 1:
The patent applies preliminary action by extracting and storing order attributes from heading rows before the table grouping process. This allows the system to preserve the original consecutive order information (such as page numbers, table sequences, or positional indicators) and use it later to correctly arrange the grouped tables, thus preventing loss of arrangement order while achieving accurate structure-based grouping
Solution Approach 2:
The patent uses order attributes extracted from heading rows as an intermediary element that bridges the gap between structure-based grouping and original table arrangement. This intermediary information (order attributes) enables the system to first group tables by structure and then restore their original consecutive order, effectively mediating between the two requirements
2Device complexity
If tables are arranged without using heading row attributes, then processing can be simplified, but accurate retrieval of attribute values from tables without heading rows becomes impossible
Solution Approach 1:
The patent applies the extraction principle by specifically extracting order attributes from heading rows of tables that do have heading rows. This extracted information is then used to guide the arrangement and attribute value retrieval process for tables without heading rows, allowing the system to maintain reliability without significantly increasing overall processing complexity
3Productivity
If multiple divided tables are processed independently, then processing speed is improved, but the relationship between divided tables cannot be established
Solution Approach 1:
The patent applies segmentation by dividing the table processing into two independent but coordinated stages: first extracting order attributes from tables with heading rows, and second using these extracted attributes to arrange and retrieve data from tables without heading rows. This segmentation allows parallel processing while preserving table relationships through the extracted order attribute information
Data Source
AI summary
An information processing apparatus includes a processor configured to: extract a table group including a first table and at least one second table from plural tables formed as images, the first table and the at least one second table having a common characteristic in terms of a table structure, the first table including a heading row representing attributes of attribute values included in the first table and the at least one second table, the at least one second table not including the heading row; relate the first table and the at least one second table included in the table group to each other so that the first table and the at least one second table are arranged in consecutive order, by using attribute values of an attribute which is included in the heading row and which represents consecutiveness of the first table and the at least one second table; and obtain attribute values corresponding to each of the attributes included in the heading row from a table sequence, the table sequence including the first table and the at least one second table arranged in consecutive order.


