Identity Resolution Through Slotting and Record Merging
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing identity verification and matching processes lack robustness and accuracy, particularly in handling big, noisy, and unstructured data environments.
Innovation Solution
A method and system for managing entity data through data mining, slotting, and record merging, utilizing a data mining engine to extract and verify characteristics, and a slotting component to create and merge records based on entity data, enhancing accuracy and robustness.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If traditional identity verification processes are used, then the process is simple to implement, but the accuracy and robustness are insufficient in handling big, noisy, and unstructured data
Solution Approach 1:
The system segments the identity verification process into distinct functional modules: a data mining engine for extracting characteristics from unstructured data, a slotting component for organizing extracted data into structured records, and a record merging component for consolidating duplicate entity records. This segmentation allows each module to specialize in specific tasks, improving overall accuracy while managing complexity through modular design.
Solution Approach 2:
The patent introduces intermediary components between raw data and final verification results. The slotting component acts as an intermediary that transforms unstructured extracted characteristics into structured entity records with standardized fields. The record merging component serves as another intermediary that compares and consolidates multiple records before final verification, thereby improving accuracy through intermediate processing steps.
2Loss of information
If multiple data sources are processed to improve entity identification, then the completeness of entity information increases, but the data noise and processing complexity increase
Solution Approach 1:
The data mining engine selectively extracts only relevant characteristics from multiple data sources using predefined schemas and patterns. This extraction process filters out irrelevant information and noise while capturing essential entity attributes. The slotting component further refines this by placing extracted characteristics into appropriate structured fields, ensuring only meaningful data is retained for verification.
Solution Approach 2:
The record merging component implements feedback mechanisms by comparing extracted characteristics against existing entity records and updating or correcting information based on consistency checks. This feedback loop allows the system to identify and eliminate noisy or contradictory data while preserving complete and accurate entity information across multiple data sources.
3Measurement precision
If manual record verification is performed, then the accuracy of entity matching is high, but the processing time and resource consumption increase significantly
Solution Approach 1:
The record merging component enables self-service automated verification by implementing comparison algorithms that automatically match new entity records against existing records in the database. The system autonomously determines whether records represent the same entity based on characteristic comparison, eliminating the need for manual verification while maintaining high accuracy through systematic comparison rules and confidence scoring.
Data Source
AI summary
In an environment containing big data, noisy data, and/or unstructured data, it is desirable to identify an entity referenced by input data. The entity can be identified by generating records corresponding to characteristics of the entity based on the input data. These records can be merged when it is determined that more than one record corresponds to the same entity. By doing so it is possible to more easily identify and classify information related to an entity, though such information may have been obtained in a manner that might otherwise be deemed unstructured or noisy. The method can be applied across large sets of data (“big data”) to obtain meaning from data that may otherwise be unclassifiable to a human observer.


