Hierarchical Clustering with Conflict Resolution
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Merging two database tables to cluster common records is time-consuming and costly, especially when dealing with redundant records that have overlapping but not identical information.
Innovation Solution
The use of hierarchical clustering with conflict resolution, employing an ordinal classifier to evaluate the degree of match between records and assigning hierarchical cluster IDs to each record, allowing for different confidence levels in clustering based on user preferences.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If traditional merging methods are used to cluster database records, then records can be grouped together, but the process becomes time-consuming and costly
Solution Approach 1:
The patent segments the record clustering process into multiple hierarchical levels (tiers). Instead of performing a single comprehensive merge, records are clustered in stages: first by common identifiers, then by additional matching criteria at subsequent tiers. This segmentation reduces computational complexity at each stage while maintaining overall clustering accuracy, directly addressing the time-cost tradeoff.
Solution Approach 2:
The patent performs preliminary actions by pre-processing records to identify and group them by common identifiers before applying more complex clustering algorithms. Records are pre-filtered and organized into candidate groups based on basic matching criteria, which reduces the search space for subsequent clustering operations and significantly decreases processing time.
2Reliability
If records are merged with high confidence only, then accuracy improves, but many valid records may be missed
Solution Approach 1:
The patent implements dynamic confidence thresholds across different hierarchical tiers. Lower tiers use more stringent matching criteria for high-confidence matches, while upper tiers progressively relax criteria to capture additional valid matches. This dynamic adjustment allows the system to balance accuracy and productivity by adapting the confidence requirement based on the clustering stage and record characteristics.
Solution Approach 2:
The patent changes matching parameters (confidence thresholds, similarity criteria) at different hierarchical levels. Each tier uses optimized parameters suited to its specific purpose: early tiers use strict parameter settings for high-precision matching, while later tiers use more permissive settings to increase recall. This parameter variation enables the system to achieve both high confidence and high productivity.
3Reliability
If multiple database tables are merged to capture all customer records, then completeness improves, but the complexity of managing redundant records increases
Solution Approach 1:
The patent segments the management of redundant records by organizing them into hierarchical clusters with clear structural relationships. Each record belongs to specific cluster groups at different tiers, creating a structured framework that simplifies tracking and management. This segmentation transforms the chaotic problem of redundant record management into an organized hierarchical system.
Solution Approach 2:
The patent adds a hierarchical dimension to record management by organizing records across multiple tiers rather than treating them as a flat set. This dimensional transformation allows the system to manage redundancy by assigning records to different cluster levels based on their matching confidence and relationship strength, making complex record relationships more manageable through hierarchical abstraction.
Data Source
AI summary
The present disclosure relates clustering similar data records together in a hierarchical clustering scheme. Each tier in a cluster corresponds to a minimal match score, which reflects a degree of confidence. In this respect, a higher confidence may lead to smaller sized clusters while a lower confidence may lead to larger sized clusters. Ordinal classification may be used to generate hierarchical clusters. In some embodiments, hierarchical clustering with conflict resolution is used to resolve user-defined hard conflicts in each tier of the clustering results.


