Cohort Clustering Visualization for Input Dimension Selection
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing clustering techniques struggle with managing large volumes of diverse data, particularly in IT environments, as they fail to systematically determine the appropriate number of input dimensions, handle long-tailed distributions, and inadequately treat categorical data, leading to overfitting and ineffective cohort identification.
Innovation Solution
The solution involves selecting optimal input dimensions using a logarithm kernel function for normalization, applying specific weightings to categorical data, and employing visualization tools to determine the appropriate number of cohorts, enabling efficient cohort identification and analysis.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If clustering techniques are applied to large volumes of diverse data, then cohort identification capability is improved, but system complexity and computational burden increase significantly
Solution Approach 1:
The patent segments the clustering process into distinct phases: data preprocessing with logarithm normalization, dimensionality reduction to select optimal input dimensions, and iterative cohort generation. This segmentation allows each phase to be optimized independently, managing system complexity while maintaining cohort identification capability.
Solution Approach 2:
The patent transforms the data by selecting a limited number of optimal input dimensions from the full feature space. This dimensionality reduction creates a simplified representation that maintains cohort identification accuracy while reducing computational complexity and system burden.
2Quantity of substance
If all available dimensions are used for clustering, then clustering completeness is improved, but overfitting occurs and model generalization deteriorates
Solution Approach 1:
The patent extracts only the most relevant input dimensions from the complete set of available features. By selecting a subset of optimal dimensions rather than using all available data, the system achieves better generalization and avoids overfitting while maintaining essential clustering information.
Solution Approach 2:
The patent changes the parameter of input dimension quantity from using all available dimensions to using a selected subset. This parameter change is guided by evaluating clustering performance metrics to identify the optimal number of dimensions that balances completeness with generalization capability.
3Adaptability or versatility
If minimal data processing is performed, then data flexibility and analysis flexibility are improved, but data volume and processing time increase
Solution Approach 1:
The patent performs preliminary data processing including logarithm normalization and optimal dimension selection before the main clustering operation. This preliminary action prepares the data in advance, enabling flexible analysis while improving processing efficiency during the actual clustering execution.
Solution Approach 2:
The patent applies parameter changes through logarithm normalization to transform the data distribution. This preprocessing step handles long-tailed distributions and prepares data for efficient clustering, balancing flexibility with processing productivity.
Data Source
AI summary
This document discloses methods and systems for cohort identification. The methods and systems include improved calculations to perform cohort identification and practical applications of the improved calculations. Specifically, the systems and methods described herein may utilize key components that include enhancements of existing cohort clustering techniques with regard to selecting a number of cohort input dimensions, normalizing input data using a logarithm kernel-function, treatment of categorical data with mutually exclusive and not-mutually exclusive values, methods and visualization tool to determine appropriate number of cohorts, methods and visualization tool to compare cohorts extracted from different input dimensions, and methods to quantify the difference in cohorts. Beyond improvements to the cohort clustering techniques, also disclosed are ancillary tools to prepare input data by joining CRM and product usage data and facilitate subsequent automated action via an API to retrieve cohort results.


