Metric-Based Identity Resolution for Digital Records
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
The increasing complexity and noise in digital records from diverse data sources make identity resolution challenging, as existing methods lack a systematic approach to efficiently connect and unify customer identities across fragmented data sources.
Innovation Solution
A system that models metric distances between digital records using common properties and applies threshold criteria to correlate and cluster records, utilizing string metric functions like Jaro-Winkler and Levenshtein distances to associate records with individual identities.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If traditional identity resolution methods are used to connect digital records, then the process can handle basic record linking, but the system cannot efficiently manage increasing data complexity and noise from multiple diverse data sources
Solution Approach 1:
The patent transforms the identity resolution problem by changing the parameter space from direct record matching to metric distance calculation in a multi-dimensional property space. Each digital record is represented as a point with coordinates corresponding to common properties (name, address, phone, etc.), and similarity is measured using standardized metric distances (Euclidean, Manhattan, Cosine). This parameter transformation enables systematic handling of heterogeneous data by projecting diverse attributes into a unified metric space where distance thresholds can reliably identify matching records despite data noise and variation across sources
2Loss of information
If more data sources are integrated to improve customer profile completeness, then the holistic picture of customers improves, but the fragmentation and noise in records increase making resolution more challenging
Solution Approach 1:
The patent introduces an intermediary metric space that mediates between diverse data sources and the final identity resolution. Instead of directly comparing records from different sources, the system first projects all records into a standardized metric space defined by common properties. This intermediary representation acts as a buffer that harmonizes heterogeneous data formats and quality levels, allowing records from multiple sources to be compared using consistent distance measures. The metric space intermediary transforms the difficult direct matching problem into a simpler distance-based clustering problem that can systematically handle information from fragmented sources
3Productivity
If manual or simple automated methods are used for record correlation, then the system can process records, but the efficiency and reliability of identity resolution deteriorate with increasing data volume
Solution Approach 1:
The patent replaces manual or simple automated record correlation methods with a systematic metric-based computational approach. Instead of relying on rule-based matching or human judgment, the system uses mathematical distance metrics (Euclidean, Manhattan, Cosine) to automatically measure similarity between records. This mechanical substitution of computational metrics for manual processes enables the system to efficiently process large volumes of records while maintaining high reliability through mathematically rigorous distance calculations and threshold-based decision rules that consistently identify matching records across diverse data sources
Data Source
AI summary
Digital records are clustered and assigned unique individual identity across a plurality of data sources. Each record contains one or more tagged common properties, which belongs to an underlying identifiable entity. A special parameterized metric distance is calculated for all pairs of records. With adjustment of the metric parameters, the records are clustered based on a predetermined distance threshold. Based on the clustered records, the underlying identity associated with the records can be resolved.


