Spreadsheet Data Cluster Identification and Table Transformation

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Determining data clusters in large spreadsheets is tedious and prone to human error, making it difficult for end-users to visually group data for analysis and reporting in business assessment and reporting software applications.

Innovation Solution

The solution involves detecting data clusters in spreadsheet files using computational geometry rules to create a materialized table view, which can be processed by business intelligence applications, automatically reducing the burden on users by identifying cluster types, associating titles and section headers, and providing metadata for further analysis.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If users manually determine data clusters in spreadsheets, then they can visually group data for analysis, but the process becomes tedious and prone to human error

Engineering Contradiction:
Improveaccuracy of data clusteringVSAvoidtime required for data clustering
Core Design Contradiction:
ReliabilityVSLoss of time

Solution Approach 1:

The system performs automatic data clustering without requiring manual user intervention. The software independently identifies data boundaries, groups related cells, and creates structured table views, allowing the system to serve itself rather than relying on manual user analysis

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The patent replaces the manual mechanical process of visual data grouping with an automated computational system. The software uses algorithms to detect data clusters, identify boundaries, and structure information, substituting human cognitive effort with automated processing

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

2Adaptability or versatility

If users manually group data cells in large spreadsheets, then they can create data clusters for reporting, but the complexity of determining boundaries increases with spreadsheet size

Engineering Contradiction:
Improveability to handle various data cluster typesVSAvoidcomplexity of cluster boundary determination
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The patent segments the spreadsheet into distinct data clusters by automatically detecting boundaries between different data regions. The system divides the large spreadsheet into manageable, logically-grouped sections that can be independently processed and analyzed

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The system changes the parameter of data organization from manual cell-by-cell grouping to automated cluster-based structuring. By transforming the data representation into clustered groups with defined boundaries, the system simplifies the complexity of handling various data types

Inventive Principle:
Principle #35Parameter changes

3Productivity

If automatic data clustering is implemented, then user burden is reduced and human error is minimized, but the system must accurately identify cluster types and associate metadata

Engineering Contradiction:
Improveefficiency of data processingVSAvoiddifficulty of identifying cluster characteristics
Core Design Contradiction:
ProductivityVSDifficulty of detecting and measuring

Solution Approach 1:

The system uses feedback mechanisms to refine cluster identification. By analyzing the detected data patterns and validating cluster characteristics against expected structures, the system adjusts its detection algorithms to improve accuracy in identifying cluster types and associated metadata

Inventive Principle:
Principle #23Feedback

Data Source

PatentUS9311371B2Data cell cluster identification and table transformation
Publication Date: 2016.04.12 SAP IRELAND LTD
  • US9311371B2 patent drawing
  • US9311371B2 patent drawing
  • US9311371B2 patent drawing

AI summary

Various embodiments may operate to access individual lines of information included in a file stored in an electronic storage medium, to detect the existence of data clusters in the file based on neighboring cell content in a horizontal direction (corresponding to the individual lines), and in a vertical direction (orthogonal to the horizontal direction), to identify at least some of the data clusters as being associated with predefined table types (comprising vertical tables, horizontal tables, or cross tables), to merge some of the data clusters into section tables having common properties, and to transform the tables resulting from the merging activity, as well as remaining un-merged data clusters, into a single flat table. The stored file may comprise a spreadsheet file.