Model-Based Cohort Clustering Visualization for Diverse Data

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing clustering techniques fail to effectively manage large volumes of diverse data with long-tailed distributions and categorical data, leading to overfitting and inadequate cohort identification in IT environments, particularly in systems like CRM and machine-generated data.

Innovation Solution

Implementing a method that includes selecting appropriate input dimensions using a logarithm kernel function for normalization, treating categorical data with mutually exclusive and non-mutually exclusive values, and visualizing cohort structures to determine the optimal number of clusters, while using K-means clustering with enhancements.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If traditional clustering techniques are used on large volumes of diverse data with long-tailed distributions, then the clustering process can be performed, but the results suffer from overfitting and inadequate cohort identification

Engineering Contradiction:
Improvecohort identification accuracyVSAvoiddata diversity complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The patent applies logarithmic transformation to normalize features with long-tailed distributions, converting skewed parameter ranges into comparable scales. This parameter transformation enables traditional clustering algorithms to effectively handle diverse data types without overfitting, directly resolving the contradiction between maintaining reliability and managing data complexity

Inventive Principle:
Principle #35Parameter changes

Solution Approach 2:

The patent segments categorical data into mutually exclusive value groups and assigns weighted representations to each category. This segmentation approach transforms complex categorical variables into structured numerical representations that clustering algorithms can process effectively, improving cohort identification while managing data diversity complexity

Inventive Principle:
Principle #1Segmentation

2Adaptability or versatility

If all generated data is stored for later analysis, then greater flexibility and comprehensive analysis are enabled, but storage costs and data management complexity increase

Engineering Contradiction:
Improveanalysis flexibilityVSAvoiddata storage volume
Core Design Contradiction:
Adaptability or versatilityVSQuantity of substance

Solution Approach 1:

The patent extracts and stores only the cluster assignment labels and essential cohort characteristics from the full raw data after clustering is performed. This extraction approach enables flexible cohort-based analysis without requiring storage of the complete high-volume raw dataset, balancing analysis versatility with manageable storage requirements

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The patent creates simplified copy representations of the data in cluster space, where each data point is represented by its cluster assignment and aggregated cohort features rather than the original high-dimensional raw data. This copying strategy preserves analytical flexibility while dramatically reducing storage volume requirements

Inventive Principle:
Principle #26Copying

3Productivity

If pre-processing extracts specified data items for efficient retrieval, then analysis efficiency improves, but the remainder of the data is discarded losing potential insights

Engineering Contradiction:
Improvedata retrieval efficiencyVSAvoiddiscarded data insights
Core Design Contradiction:
ProductivityVSLoss of information

Solution Approach 1:

The patent transforms the data representation by projecting high-dimensional raw data into cluster space dimensions, where each point is represented by its cluster assignment and aggregated features. This dimensionality transformation enables efficient retrieval and analysis of cohort-based patterns while preserving information from the entire original dataset, avoiding the information loss inherent in traditional pre-processing extraction

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

Data Source

PatentUS12412103B1Monitoring and visualization of model-based clustering definition performance
Publication Date: 2025.09.09 CISCO TECHNOLOGY INC
  • US12412103B1 patent drawing
  • US12412103B1 patent drawing
  • US12412103B1 patent drawing

AI summary

This document discloses methods and systems for cohort identification. The methods and systems include improved calculations to perform cohort identification and practical applications of the improved calculations. Specifically, the systems and methods described herein may utilize key components that include enhancements of existing cohort clustering techniques with regard to selecting a number of cohort input dimensions, normalizing input data using a logarithm kernel-function, treatment of categorical data with mutually exclusive and not-mutually exclusive values, methods and visualization tool to determine appropriate number of cohorts, methods and visualization tool to compare cohorts extracted from different input dimensions, and methods to quantify the difference in cohorts. Beyond improvements to the cohort clustering techniques, also disclosed are ancillary tools to prepare input data by joining CRM and product usage data and facilitate subsequent automated action via an API to retrieve cohort results.