Entity Resolution via Anonymization for Heterogeneous Data Security

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional entity resolution systems face challenges in handling heterogeneous data sources, maintaining data security, and accurately associating entity references across datasets, especially in scenarios like public sector and federal datasets, where data volume and velocity are high, and semantic relationships are complex.

Innovation Solution

A system comprising an entity reference parsing subsystem, property value standardization and anonymization subsystems, property strength quantification, local and global entity resolution subsystems, and an entity creation and updating subsystem, which parses and standardizes entity references, anonymizes property values, identifies additional properties, and performs entity resolution processes to securely associate and update entity references.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Ease of manufacture

If conventional textual analysis approaches are used for entity resolution, then entity resolution can be accomplished using simple methods, but the system cannot handle heterogeneous data sources and is prone to data breaches

Engineering Contradiction:
Improvesimplicity of entity resolution methodVSAvoidability to handle heterogeneous data sources
Core Design Contradiction:
Ease of manufactureVSAdaptability or versatility

Solution Approach 1:

The patent introduces an intermediary anonymization layer between the raw entity data and the resolution process. Entity references are anonymized before being used in textual analysis, allowing the system to handle heterogeneous data sources without exposing sensitive information. This intermediary step enables versatility while maintaining security.

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The system transforms entity references by changing their parameter representation through anonymization techniques. Instead of using original sensitive data, the system uses anonymized versions that preserve resolution capability while eliminating security risks. This parameter transformation enables handling of diverse data sources with uniform security treatment.

Inventive Principle:
Principle #35Parameter changes

2Reliability

If conventional entity resolution systems process all entity references, then comprehensive entity association can be achieved, but sensitive information leakage and data breaches occur

Engineering Contradiction:
Improvecompleteness of entity associationVSAvoidsensitive information leakage
Core Design Contradiction:
ReliabilityVSObject-affected harmful factors

Solution Approach 1:

The patent converts the potentially harmful exposure of sensitive information into a benefit by using anonymization. The anonymization process removes harmful sensitive details while preserving the essential characteristics needed for entity resolution. This transforms what would be a security risk into a secure resolution mechanism that maintains reliability.

Inventive Principle:
Principle #22Blessing in disguise (Convert harm into benefit)

Solution Approach 2:

The system extracts and removes sensitive information from entity references before processing. By taking out the harmful sensitive components and retaining only the anonymized essential features, the system achieves comprehensive entity association without exposing sensitive data, thus preventing information leakage while maintaining completeness.

Inventive Principle:
Principle #2Taking out (Extraction)

3Device complexity

If the system uses homogeneous entity references with same attribute types, then entity resolution is simpler, but the system cannot adapt to new attributes or inhomogeneous data sources

Engineering Contradiction:
Improvecomplexity of entity resolution systemVSAvoidability to handle new attributes and inhomogeneous data
Core Design Contradiction:
Device complexityVSAdaptability or versatility

Solution Approach 1:

The patent implements a universal anonymization framework that can handle multiple attribute types and heterogeneous data sources through the same process. The anonymization subsystem is designed to work with any entity reference format, making the system multi-functional and adaptable to new attributes without increasing complexity. This universal approach enables the system to process inhomogeneous data while maintaining simple resolution logic.

Inventive Principle:
Principle #6Universality (Multi-functionality)

4Measurement precision

If entity references are assumed to be clean with majority attributes pointing to particular entity, then resolution accuracy is high, but the system fails in noisy or contaminated data environments

Engineering Contradiction:
Improveentity resolution accuracyVSAvoidrobustness to noisy and contaminated data
Core Design Contradiction:
Measurement precisionVSAdaptability or versatility

Solution Approach 1:

The system performs preliminary anonymization processing on entity references before the resolution accuracy assessment. This preliminary action of anonymization cleanses the data by removing noise and contaminated information while preserving essential identifying features. By preparing the data in advance through anonymization, the system maintains high resolution accuracy even when processing noisy or contaminated data sources.

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS12105845B2System and method for entity resolution of a data element
Publication Date: 2024.10.01 SECURITI LLC
  • US12105845B2 patent drawing
  • US12105845B2 patent drawing
  • US12105845B2 patent drawing

AI summary

A system and a method for entity resolution is disclosed. An entity reference parsing subsystem to parse one or more entity references of a corresponding seed set of entity into corresponding one or more personal data properties and property values. A property value standardization subsystem performs one or more standardization operations for standardization of the corresponding one or more property values. A property value anonymization subsystem secures the one or more property values by performing one or more anonymization procedures. A property strength quantification subsystem identifies at least one additional property suspected to belong to the seed set of the entity, assigns a property strength score to the at least one additional property, adds the at least one additional property to the corresponding seed set of entity. A local entity resolution and a global entity resolution subsystem performs a first and a second entity resolution process respectively.