Mixed Data Table Area Extraction via Constant Value Summation
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional techniques fail to properly extract numerical value management areas and header areas from numerical tables when cells contain both numerical values and character strings, leading to inaccurate demarcation.
Innovation Solution
A method that replaces numerical values and character strings with constant values of opposite signs, generates rectangular areas, and calculates sums to identify the numerical value management, row header, and column header areas based on score comparisons, ensuring accurate extraction even when tables contain mixed data types.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If conventional techniques are used to extract areas from numerical tables, then the extraction process is simple, but the accuracy deteriorates when cells contain mixed data types (numerical values and character strings)
Solution Approach 1:
The patent transforms the data type recognition problem into a numerical optimization problem by replacing numerical values with a first constant value and character strings with a second constant value of opposite sign. This parameter transformation enables the use of sum maximization algorithms to automatically identify area boundaries, achieving high accuracy in mixed data type tables without requiring complex rule-based processing.
Solution Approach 2:
The patent replaces manual or rule-based area demarcation methods with an automated algorithmic approach. By substituting the mechanical process of manual inspection with a computational sum-maximization algorithm, the system achieves both high accuracy and efficiency in extracting numerical value management areas and header areas from tables with mixed data types.
2Productivity
If manual demarcation of header areas and numerical value management areas is performed, then accuracy is maintained, but productivity deteriorates when there are a large number of numerical tables
Solution Approach 1:
The patent enables the system to automatically identify and demarcate areas without human intervention. By using the sum maximization algorithm that automatically distinguishes between numerical value management areas and header areas based on the constant value assignments, the system achieves both high productivity for processing large numbers of tables and high accuracy comparable to manual demarcation.
3Adaptability or versatility
If conventional area extraction methods are used, then computational complexity is low, but the ability to handle mixed data types deteriorates
Solution Approach 1:
The patent achieves versatility in handling mixed data types by transforming the heterogeneous data into a homogeneous numerical representation through constant value assignment. This parameter change allows the use of efficient sum maximization algorithms while maintaining the ability to correctly identify areas in tables with mixed numerical and textual data, without requiring complex data type detection logic.
Data Source
AI summary
A processor obtains a table that contains numerical values or character strings in its cells. The processor then replaces each numerical value with a first constant value, and each character string with a second constant value. The two constant values have opposite signs. The processor generates area datasets each including first to third rectangular areas. The right side of the second rectangular area coincides with the left side of the first rectangular area. The bottom side of the third rectangular area coincides with the top side of the first rectangular area. With respect to each generated area dataset, the processor compares a sum of first and second constant values in the first rectangular area with a sum of first and second constant values in the second and third rectangular areas. The processor outputs at least one of the area datasets according to the comparison result.


