Persistence Measures for High-Dimensional Cluster Interpretation

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional techniques for identifying distinguishing features in high-dimensional and semantically connected data, such as medical data, are inadequate, particularly when features are repetitive across clusters.

Innovation Solution

A software and/or hardware facility uses a persistence measure to quantify feature uniqueness across clusters, employing techniques like PageRank and dimensionality reduction to identify features that best distinguish clusters, and predicts cluster assignments for new data items.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If conventional techniques are used to identify distinguishing features in high-dimensional data, then the analysis can be performed with simple methods, but the techniques are inadequate when features are repetitive across clusters and fail to effectively distinguish clusters

Engineering Contradiction:
Improvefeature distinction accuracyVSAvoidanalysis method complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent segments the feature analysis process into multiple components: computing a persistence measure for each feature-cluster pair, ranking features within clusters based on persistence measures, and selectively analyzing features based on their ranks. This segmentation allows the system to handle high-dimensional data with repetitive features by breaking down the complex analysis into manageable steps that progressively identify distinguishing features.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent introduces a persistence measure parameter that quantifies how consistently a feature appears within a specific cluster compared to other clusters. By changing the analysis parameter from simple feature presence to persistence measurement, the system can effectively distinguish clusters even when features are repetitive, thereby improving measurement precision without requiring overly complex methods.

Inventive Principle:
Principle #35Parameter changes

2Loss of information

If all features are analyzed in detail to identify distinguishing features, then comprehensive cluster understanding is achieved, but computational resources and time are excessively consumed

Engineering Contradiction:
Improvefeature information completenessVSAvoidanalysis speed
Core Design Contradiction:
Loss of informationVSProductivity

Solution Approach 1:

The patent applies partial action by computing persistence measures for all features but then ranking and selectively focusing on only the top-ranked features within each cluster. Instead of analyzing all features in equal detail, the system performs comprehensive measurement but selective in-depth analysis, maintaining information completeness while improving analysis speed by concentrating resources on the most distinguishing features.

Inventive Principle:
Principle #16Partial or excessive action

Solution Approach 2:

The patent extracts and isolates the most important distinguishing features by ranking them based on persistence measures. By taking out only the top-ranked features for detailed analysis and presentation, the system reduces computational overhead and increases productivity while still capturing the essential information needed to understand cluster distinctions.

Inventive Principle:
Principle #2Taking out (Extraction)

3Measurement precision

If persistence measures are computed for all feature-cluster pairs, then accurate cluster distinction is achieved, but computational complexity and resource consumption increase

Engineering Contradiction:
Improvecluster distinction accuracyVSAvoidcomputational resources
Core Design Contradiction:
Measurement precisionVSQuantity of substance

Solution Approach 1:

The patent segments the computational process into phases: first computing persistence measures for all feature-cluster pairs, then ranking features within each cluster, and finally selecting only the top-ranked features for detailed analysis. This segmentation allows accurate cluster distinction through comprehensive persistence measurement while reducing overall resource consumption by not processing all features to the same level of detail.

Inventive Principle:
Principle #1Segmentation

Data Source

PatentUS20250307268A1Cluster interpretation using a persistence measure
Publication Date: 2025.10.02 PROVIDENCE ST JOSEPH HEALTH
  • US20250307268A1 patent drawing
  • US20250307268A1 patent drawing
  • US20250307268A1 patent drawing

AI summary

A facility for analyzing the features of data items organized into clusters is described. The facility analyzes the data items of the clusters when the features of the data items are high dimensional and categorical with overlapping values across the clusters. The facility identifies the most distinguishable features that uniquely differentiate the clusters given the above nature of the feature space.