Attribute Labeling for Numerical Tabular Data Using Sum and Product Relations
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing methods for attribute labeling in numerical tabular data face challenges in accurately recognizing hierarchical structures, leading to incorrect labeling and high costs in creating exhaustive patterns, especially when dealing with diverse and complex data formats like Excel or CSV, which are not easily machine-processable.
Innovation Solution
A data processing method that specifies regions in a data table based on numerical and character string relationships, using a sum or product input-output relation to associate character strings with numerical values, and applies an attribute labeling pattern that considers nesting and error thresholds to correctly label attributes, reducing incorrect labeling and pattern creation costs.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If existing attribute labeling methods are used, then automatic labeling can be performed, but incorrect labeling occurs and pattern creation costs increase
Solution Approach 1:
The patent segments the attribute labeling process into distinct phases: identifying candidate attributes, determining hierarchical relationships through sum/product relations, and applying labels systematically. This segmentation reduces errors by handling complex labeling in manageable steps rather than attempting exhaustive pattern matching.
Solution Approach 2:
The patent performs preliminary identification of sum and product relations among numerical values before applying attribute labels. By pre-establishing the hierarchical structure through mathematical relations, the system prepares the data framework in advance, reducing the need for complex exhaustive patterns and improving labeling accuracy.
2Adaptability or versatility
If exhaustive attribute labeling patterns are created to handle diverse data formats, then labeling coverage improves, but pattern creation costs and complexity increase significantly
Solution Approach 1:
The patent employs universal sum and product relation detection that works across diverse data formats (Excel, CSV, and other tabular formats) without requiring format-specific patterns. The mathematical relation-based approach is universally applicable to any numerical tabular data, providing adaptability without increasing pattern creation complexity.
Solution Approach 2:
The patent changes the approach from creating exhaustive attribute patterns to detecting mathematical parameters (sum and product relations) among numerical values. This parameter-based method adapts to diverse data formats by focusing on the inherent mathematical relationships rather than format-specific characteristics, reducing pattern creation requirements.
3Extent of automation
If mechanical processing is applied to numerical tabular data with unexplicit hierarchical structures, then automation is achieved, but labeling accuracy decreases
Solution Approach 1:
The patent incorporates feedback mechanisms where the detected sum and product relations inform the hierarchical structure determination. The system uses the mathematical relationships as feedback to automatically infer the correct hierarchical arrangement of attributes, maintaining both automation and accuracy by letting the data structure guide the labeling process.
Solution Approach 2:
The patent enables the numerical tabular data to self-reveal its hierarchical structure through intrinsic sum and product relations among numerical values. Rather than requiring external pattern matching, the data itself provides the structural information through mathematical relationships, achieving automation without sacrificing accuracy.
Data Source
AI summary
A data processing method executed by a computer, the data processing method including specifying a first region range among from a data table, a first region range including a plurality of numerical value regions which are continuously disposed in a first direction, a plurality of numerical values in the plurality of numerical value regions having a relationship with a specified numerical value in an adjacent region, specifying a second region range, the second region range being specified by shifting the first region range in a second direction, the second region range including at least one character string region and at least one blank region, associating a character string in the at least one character string region and the plurality of numerical values, and outputting data that indicates an association between the character string in the at least one character string region and the plurality of numerical values.


