Cluster-ensemble model for reducing computational resource utilization in data segmentation

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing data segmentation methods are inefficient and resource-intensive, particularly when dealing with large datasets, due to limitations in processing power and memory, leading to inaccurate and meaningless data segments.

Innovation Solution

The use of a cluster-ensemble model that embeds raw datasets into lower-dimensional embedded datasets, allowing for reduced computational resources while generating accurate and meaningful data segments through a "data first" approach.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Manufacturing precision

If clustering algorithms are applied to raw big data datasets, then data segmentation is performed, but computational resources (processing power and memory) are excessively consumed

Engineering Contradiction:
Improvedata segmentation accuracyVSAvoidcomputational resource utilization
Core Design Contradiction:
Manufacturing precisionVSUse of energy by moving object

Solution Approach 1:

The patent applies dimensionality reduction techniques (e.g., PCA, t-SNE, UMAP) to transform the raw high-dimensional dataset into a lower-dimensional embedded dataset before applying clustering algorithms. This preliminary transformation reduces the computational complexity and memory requirements of subsequent clustering operations while preserving the essential data structure and relationships needed for accurate segmentation.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent changes the dimensional parameter of the dataset by transforming data from high-dimensional raw space to lower-dimensional embedded space. This parameter transformation maintains the essential variance and relationships in the data while significantly reducing the computational burden on processing power and memory resources.

Inventive Principle:
Principle #35Parameter changes

2Quantity of substance

If clustering algorithms process large datasets, then more data points are analyzed, but the accuracy of data segments decreases due to the sheer scale of potential similarities

Engineering Contradiction:
Improvedata volume processedVSAvoidsegmentation accuracy
Core Design Contradiction:
Quantity of substanceVSMeasurement precision

Solution Approach 1:

The patent performs dimensionality reduction before clustering to transform the data landscape, making it easier for clustering algorithms to identify meaningful patterns. By preprocessing the data to reduce dimensions while preserving structure, the system can effectively process large volumes of data without losing segmentation accuracy.

Inventive Principle:
Principle #10Preliminary action

3Adaptability or versatility

If clustering models are specifically designed for given questions, then targeted segmentation is achieved, but computational resources are wasted creating specialized algorithms for each domain

Engineering Contradiction:
Improvedomain-specific customizationVSAvoidalgorithm development efficiency
Core Design Contradiction:
Adaptability or versatilityVSProductivity

Solution Approach 1:

The patent employs universal dimensionality reduction techniques and general-purpose clustering algorithms that can be applied across different domains and data types without requiring domain-specific customization. This universal approach eliminates the need to create specialized algorithms for each question or domain, significantly improving development efficiency while maintaining versatility through the flexibility of the embedding process.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Data Source

PatentUS20250131063A1Reducing utilization of computational resources associated with segmenting datasets via a cluster- ensemble model systems and methods
Publication Date: 2025.04.24 CAPITAL ONE SERVICES LLC
  • US20250131063A1 patent drawing
  • US20250131063A1 patent drawing
  • US20250131063A1 patent drawing

AI summary

In some embodiments, reducing utilization of computational resources associated with segmenting datasets via a cluster-ensemble model may be facilitated. In some embodiments, the system may receive a raw dataset having a first dimension. The system may then embed the raw dataset into an embedded dataset having a second dimension, where the embedded dataset comprises a vector embedding. The system may then provide the embedded dataset to a set of clustering models to generate a set of clusters, where each cluster of the set of clusters corresponds to a respective clustering model of the set of clustering models. The system may provide the set of clusters to a cluster-ensemble model to generate a set of ensemble-clusters. Based on the set of ensemble-clusters, the system may generate a set of data segments corresponding to the set of ensemble-clusters indicating at least one characteristic of a respective ensemble-cluster of the set of ensemble-clusters.