Spreadsheet Data Cluster Identification and Table Transformation
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Determining data clusters in large spreadsheets is tedious and prone to human error, making it difficult for end-users to visually group data for analysis and reporting in business assessment and reporting software applications.
Innovation Solution
The solution involves detecting data clusters in spreadsheet files using computational geometry rules to create a materialized table view, which can be processed by business intelligence applications, automatically reducing the burden on users by identifying cluster types, associating titles and section headers, and providing metadata for further analysis.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If users manually determine data clusters in spreadsheets, then they can visually group data for analysis, but the process becomes tedious and prone to human error
Solution Approach 1:
The system performs automatic data clustering without requiring manual user intervention. The software independently identifies data boundaries, groups related cells, and creates structured table views, allowing the system to serve itself rather than relying on manual user analysis
Solution Approach 2:
The patent replaces the manual mechanical process of visual data grouping with an automated computational system. The software uses algorithms to detect data clusters, identify boundaries, and structure information, substituting human cognitive effort with automated processing
2Adaptability or versatility
If users manually group data cells in large spreadsheets, then they can create data clusters for reporting, but the complexity of determining boundaries increases with spreadsheet size
Solution Approach 1:
The patent segments the spreadsheet into distinct data clusters by automatically detecting boundaries between different data regions. The system divides the large spreadsheet into manageable, logically-grouped sections that can be independently processed and analyzed
Solution Approach 2:
The system changes the parameter of data organization from manual cell-by-cell grouping to automated cluster-based structuring. By transforming the data representation into clustered groups with defined boundaries, the system simplifies the complexity of handling various data types
3Productivity
If automatic data clustering is implemented, then user burden is reduced and human error is minimized, but the system must accurately identify cluster types and associate metadata
Solution Approach 1:
The system uses feedback mechanisms to refine cluster identification. By analyzing the detected data patterns and validating cluster characteristics against expected structures, the system adjusts its detection algorithms to improve accuracy in identifying cluster types and associated metadata
Data Source
AI summary
Various embodiments may operate to access individual lines of information included in a file stored in an electronic storage medium, to detect the existence of data clusters in the file based on neighboring cell content in a horizontal direction (corresponding to the individual lines), and in a vertical direction (orthogonal to the horizontal direction), to identify at least some of the data clusters as being associated with predefined table types (comprising vertical tables, horizontal tables, or cross tables), to merge some of the data clusters into section tables having common properties, and to transform the tables resulting from the merging activity, as well as remaining un-merged data clusters, into a single flat table. The stored file may comprise a spreadsheet file.


