ALBG Data Clustering with Pareto Validation
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional clustering techniques fail to determine the correctness of the clustering process, including the number of clusters, the set of users in clusters, and cluster composition, due to susceptibility to bias from initial sampling and business rules, and lack qualitative assurance.
Innovation Solution
The proposed method and system implement Assurance-enabled Linde Buzo Gray (ALBG) data clustering, which uses a processor to segment user data into clusters based on features and feature values, with a Pareto validation module to ensure validity and a qualitative assurance module to ensure correct assignment of data elements, using a CoP value for iterative refinement and storage in a database.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Ease of manufacture
If conventional sampling-based clustering approaches are used, then the clustering process is simple to implement, but the results lack replicability and are highly susceptible to bias
Solution Approach 1:
The patent implements a feedback mechanism where clustering results are validated against Pareto efficiency criteria. The system iteratively adjusts clustering parameters and validates results to ensure replicability, creating a closed-loop process that continuously improves reliability while maintaining implementation feasibility.
Solution Approach 2:
The patent transforms the clustering approach by changing key parameters from fixed sampling sizes to dynamic parameters optimized through Pareto analysis. This allows the system to adapt clustering parameters based on data characteristics, improving replicability without significantly complicating implementation.
2Ease of operation
If rule-based segmentation approaches are used, then the clustering process is interpretable and controllable, but the results are biased by the selection of business rules
Solution Approach 1:
The patent introduces Pareto efficiency analysis as an intermediary between subjective business rules and clustering results. This intermediary objective criterion validates whether clustering outcomes are truly optimal, reducing bias from rule selection while preserving interpretability through the clear Pareto framework.
Solution Approach 2:
The patent performs preliminary Pareto validation on potential clustering solutions before finalizing them. This preliminary action filters out biased or suboptimal clusterings early in the process, ensuring that only objectively valid results proceed to interpretation and application stages.
3Measurement precision
If iterative segmentation with accuracy criterion is performed, then the clustering accuracy is improved, but the computational complexity and time increase
Solution Approach 1:
The patent applies partial iteration by performing segmentation only to the extent needed to achieve Pareto optimality. Rather than exhaustively iterating through all possible segmentations, the system stops when Pareto criteria are satisfied, achieving sufficient accuracy without excessive computational time.
Solution Approach 2:
The patent performs preliminary assessments of data characteristics to determine appropriate iteration limits and accuracy thresholds before full segmentation. This preliminary action allows the system to achieve good accuracy by focusing computational effort where most beneficial, reducing overall processing time.
Data Source
AI summary
Methods and systems for Assurance-enabled Linde Buzo Gray (ALBG) data clustering is described herein. In an implementation, a user model data from a database available to the processor is obtained. The user model data comprises data elements or users, each of which corresponds to features and feature values associated with the users. These data elements of the user model data are segmented into clusters using our segmentation approach with an initial accuracy criterion parametric value and the output is captured as segment data. The segment data output is checked for initial pareto validity. If successful, iterative segmentation run with incremental accuracy criterion using parameterized value is performed till the segmented clusters are determined valid against pareto validity check. The last successful pareto valid segmented cluster data is considered as the finalized segment output data. For an invalid initial pareto validity check, a segmentation run with a pre-determined accuracy criterion value is done to arrive at the finalized segment output data.


