Interactive Clustering Application for Visual Data Exploration
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current data analytics systems struggle to effectively allow users to visually explore and interactively cluster data sets, particularly in identifying data items that are dissimilar to specific reference clusters, as they often rely on two-dimensional visualizations that may not capture the full complexity and relevance of data attributes.
Innovation Solution
A visual analytics system with a clustering application hosted on a client computer, allowing users to interactively define clusters based on specified degrees of similarity and dissimilarity, using a user interface with sliders and attribute selectors to visually display discriminative clusters, enabling users to query and retrieve data items that meet specific criteria.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Ease of operation
If two-dimensional visualizations are used to display data clusters, then the visualization is simple and easy to understand, but the full complexity and relevance of data attributes cannot be captured
Solution Approach 1:
The patent transitions from two-dimensional visualizations to a hyperdimensional space that can represent multiple data attributes simultaneously. The system projects high-dimensional data clusters into visual space while preserving the relationships and complexities of the original multi-attribute data, allowing users to explore data in dimensions beyond what traditional 2D charts can display.
2Measurement precision
If traditional clustering methods are used, then data items with similar characteristics are grouped together, but users cannot effectively identify data items that are dissimilar to specific reference clusters
Solution Approach 1:
The patent inverts the traditional clustering approach by enabling users to define clusters based on dissimilarity to reference data items. Instead of only grouping similar items together, the system allows users to specify reference clusters and identify items that are deliberately dissimilar to those references, providing a complementary perspective for data exploration.
Solution Approach 2:
The system provides dynamic cluster definition capabilities where users can interactively adjust similarity thresholds, swap reference data items, and reconfigure clusters in real-time. This dynamic approach allows flexible exploration of both similar and dissimilar data relationships without being constrained by fixed clustering parameters.
3Productivity
If interactive visual interfaces are provided for data exploration, then user engagement and analysis capability are improved, but the system complexity and computational requirements increase
Solution Approach 1:
The patent implements self-service capabilities where the system automatically performs computationally intensive tasks such as calculating similarity metrics, projecting high-dimensional data into visual space, and updating cluster configurations. This automation reduces the manual effort required for complex data exploration while maintaining interactive functionality for users.
Solution Approach 2:
The system performs preliminary computations and data processing in advance, such as pre-calculating similarity matrices and preparing projected visualizations, so that interactive operations during user exploration can be executed quickly. This preliminary action reduces real-time computational requirements while maintaining high interactivity.
Data Source
AI summary
A visual analytics system includes a memory and a processor. The processor executes a clustering application having an interactive user-interface rendered on a client computer. The clustering application determines a first cluster of data items of a data set, the data items in the first cluster having first attribute values that are similar to each other within a first degree of similarity and determines a second cluster of data items of the data set, the data items in the second cluster having second attribute values that are similar to each other within a second degree of similarity. For visual analytics, the user interface receives a user selection of a third degree of similarity. In response to which, the clustering application determines a third cluster of data items of the data set, the data items in the third cluster being dissimilar to either the first attribute value of the first reference data item or the second attribute value of the second reference data item by at least the third degree of similarity, and visually displays the third cluster of data items on the user interface.


