Hashed User Data Distillation for Privacy-Preserving Deduplication

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

The sharing of user data by service providers raises privacy concerns and results in the generation and storage of redundant data, wasting computational resources.

Innovation Solution

A system and method for data distillation that uses a distributed database to merge, combine, and prune user data, maintaining privacy by using anonymized hashes and generating new records, while sharing only public data among entities.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Loss of information

If service providers share user data to maximize information discernment, then information completeness is improved, but user privacy is compromised

Engineering Contradiction:
Improveinformation completenessVSAvoiduser privacy
Core Design Contradiction:
Loss of informationVSObject-affected harmful factors

Solution Approach 1:

The patent extracts identifying information from user data before sharing. Service providers remove personally identifiable information (PII) such as names, addresses, and phone numbers from datasets before exchanging them with other providers. This extraction process retains behavioral and preference data that is valuable for marketing while eliminating privacy risks associated with direct identification.

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The patent introduces hashed identifiers as an intermediary mechanism for data matching. Instead of sharing raw identifying information, providers use hashed versions of user identifiers that cannot be reversed to obtain original data. These hashed identifiers serve as mediators that enable data correlation across different providers without exposing actual user identities, thus resolving the privacy-information completeness contradiction.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Object-affected harmful factors

If service providers remove identifying information to maintain user privacy, then user privacy is protected, but redundant data is generated and stored

Engineering Contradiction:
Improveuser privacyVSAvoidcomputational resources
Core Design Contradiction:
Object-affected harmful factorsVSLoss of energy

Solution Approach 1:

The patent merges multiple anonymized datasets by matching records through hashed identifiers. When the same user interacts with different service providers, their data is correlated using the hashed identifier as a key, consolidating information into unified user profiles. This merging eliminates redundant storage of separate anonymized records for the same individual while preserving privacy through the use of hashed identifiers.

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The patent uses hashed copies of identifiers instead of original identifying information. Rather than storing and processing multiple versions of actual user identities, the system creates and exchanges hashed copies that serve the same functional purpose for data matching and correlation. These cryptographic copies consume minimal computational resources while enabling efficient deduplication across distributed datasets.

Inventive Principle:
Principle #26Copying

3Object-affected harmful factors

If service providers store multiple sets of repetitive anonymized data, then user privacy is maintained, but storage space and processing power are wasted

Engineering Contradiction:
Improveuser privacyVSAvoiddata storage volume
Core Design Contradiction:
Object-affected harmful factorsVSQuantity of substance

Solution Approach 1:

The patent merges multiple anonymized datasets by matching records through hashed identifiers. When the same user interacts with different service providers, their data is correlated using the hashed identifier as a key, consolidating information into unified user profiles. This merging eliminates redundant storage of separate anonymized records for the same individual while preserving privacy through the use of hashed identifiers.

Inventive Principle:
Principle #5Merging (Combining)

Data Source

PatentUS20260086983A1Systems and methods for data distillation
Publication Date: 2026.03.26 COMCAST CABLE COMM LLC
  • US20260086983A1 patent drawing
  • US20260086983A1 patent drawing
  • US20260086983A1 patent drawing

AI summary

Systems and methods are described for distilling data. First data associated with a user may be received. The first data associated with the user may comprise an anonymized hash of an identifier associated with the user. A database may be determined to comprise a first record indicating the anonymized hash. The first record may comprise second data associated with the user. Based on the determining that the database comprises the first record, a second record may be generated. The second record may comprise the first data associated with the user, the second data associated with the user, and the anonymized hash. Based on the determining that the database comprises the first record, the example method may be stored to the database. These and other user and/or data distillation methods and systems are described herein.