Clustering Heterogeneous Data Using Input Output Space Segmentation
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current data clustering methods are inadequate for handling heterogeneous data types, often resulting in clusters with heterogeneous variables, which complicates data analysis and dimensionality reduction in applications like medical monitoring systems.
Innovation Solution
A system that clusters documents by considering both input and output space data, using similarity measures to aggregate and subdivide clusters, ensuring homogeneity within clusters through a combination of input and output space processors and storage devices.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If traditional clustering methods are used on heterogeneous data, then clustering can be performed, but the resulting clusters contain heterogeneous variables which reduces manufacturing precision
Solution Approach 1:
The patent segments the clustering process into two distinct phases: input space clustering that groups documents by textual similarity, and output space refinement that subdivides clusters based on response data homogeneity. This segmentation allows the system to first handle heterogeneous data types broadly, then progressively refine clusters to achieve variable homogeneity within each cluster.
Solution Approach 2:
The patent introduces a second dimension (output space) to the traditional single-dimension clustering approach. By computing similarity measures in both input space (textual similarity) and output space (response data similarity), and combining these dimensions, the system achieves homogeneous clusters from heterogeneous data types that single-dimension methods cannot produce.
2Productivity
If data is clustered without considering output space properties, then processing speed is improved, but data analysis reliability deteriorates due to heterogeneous variables in clusters
Solution Approach 1:
The patent performs preliminary clustering in the input space using textual similarity measures before refining clusters based on output space properties. This preliminary action creates initial groupings that capture broad document similarities quickly, then allows subsequent refinement to ensure reliability without requiring complete analysis from the start.
Solution Approach 2:
The clustering process is made dynamic through iterative refinement. Clusters are initially formed based on input space similarity, then dynamically adjusted by subdividing clusters that contain heterogeneous output space variables. This dynamic approach balances processing efficiency with analysis reliability by refining only when necessary.
3Device complexity
If clusters are formed without subdividing based on output space similarity, then device complexity is reduced, but measurement precision deteriorates
Solution Approach 1:
The patent applies partial refinement by subdividing clusters only when output space similarity measures indicate heterogeneity. Rather than uniformly processing all clusters through both input and output space analysis, the system selectively refines clusters that require it, balancing complexity reduction with measurement precision maintenance.
Data Source
AI summary
A system for clustering a plurality of documents having input and output space data is disclosed that uses both input and output space criteria. The system can aggregate documents into clusters based on input and/or output space similarity measures, and then refine the clusters based on further input and/or output space similarity measures. Aggregation of documents into clusters can include forming a hierarchical tree based on the input and/or output space similarity measures where the hierarchical tree has a root node, branching into intermediate nodes, and branching into leaf nodes covering individual documents, where the hierarchical tree includes a leaf node for each document of the plurality of documents. The system can include forming a forest of sub-trees of the hierarchical tree based on cluster criteria. Textual and numeric similarity measures can be used depending on the type and distribution of data in the input and output spaces.


