Financial User Cohort Clustering for Data Analysis Efficiency

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Analyzing vast amounts of user data stored in databases is time-consuming and resource-intensive, as comparing and processing data from multiple sources with different user profiles consumes significant computing resources and storage capacity.

Innovation Solution

Forming clusters of user data based on common characteristics or attributes, excluding outliers, and pruning unneeded clusters to reduce the data set analyzed, thereby improving data analysis efficiency and reducing resource usage.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Quantity of substance

If vast amounts of user data are stored in databases, then data completeness and user profile detail are improved, but data analysis time and computing resource consumption increase

Engineering Contradiction:
Improvedata volumeVSAvoiddata analysis efficiency
Core Design Contradiction:
Quantity of substanceVSProductivity

Solution Approach 1:

The patent segments the large database into multiple clusters based on user characteristics, behaviors, and attributes. Each cluster represents a subset of users with similar profiles, allowing the system to analyze only relevant clusters for specific queries rather than processing the entire database, thus improving analysis efficiency while maintaining data completeness

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent extracts and stores cluster definitions and metadata separately from the main user data. By pre-computing and storing cluster information (such as cluster IDs, characteristic summaries, and member counts), the system can quickly identify and retrieve relevant user groups without scanning the entire database, reducing analysis time and resource consumption

Inventive Principle:
Principle #2Taking out (Extraction)

2Adaptability or versatility

If data from multiple sources with different user profiles is processed, then data comprehensiveness is improved, but computing resource consumption and storage capacity requirements increase

Engineering Contradiction:
Improvedata source compatibilityVSAvoidcomputing resource consumption
Core Design Contradiction:
Adaptability or versatilityVSUse of energy by stationary object

Solution Approach 1:

The patent implements a universal cluster framework that can handle multiple data sources and user profiles through a common clustering mechanism. The system uses standardized user attribute schemas and cluster definition formats that work across different data sources, allowing the same clustering infrastructure to process diverse user data without requiring separate processing pipelines for each source

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The patent applies different clustering strategies and algorithms tailored to specific data source characteristics and user profile types. Rather than using a single uniform approach, the system can select appropriate clustering methods for different data sources (e.g., demographic-based clustering for one source, behavior-based clustering for another), optimizing resource usage for each data type while maintaining overall system versatility

Inventive Principle:
Principle #3Local quality

Data Source

PatentEP3701480B1Systems and methods for intelligently grouping financial product users into cohesive cohorts
Publication Date: 2024.05.22 INTUIT INC
  • EP3701480B1 patent drawingFigure 1
  • EP3701480B1 patent drawingFigure 2
  • EP3701480B1 patent drawingFigure 3

AI summary

Systems and methods are provided that, in some embodiments that extract user data from at least one data warehouse. The user data is sorted within each dimension, and partitions each dimension into bins. Clusters are defined as each bin that includes user data for a number of users that exceeds a threshold. Clusters are determined for every combination of dimensions. Each combination of clusters that exceed the threshold is defined as clusters that are formed from multiple dimensions. All clusters and other clusters are stored into a cluster definition table. The clusters are used to analyze the profile of specific users.