Hashed User Data Distillation for Privacy-Preserving Deduplication
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
The sharing of user data by service providers raises privacy concerns and results in the generation and storage of redundant data, wasting computational resources.
Innovation Solution
A system and method for data distillation that uses a distributed database to merge, combine, and prune user data, maintaining privacy by using anonymized hashes and generating new records, while sharing only public data among entities.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Loss of information
If service providers share user data to maximize information discernment, then information completeness is improved, but user privacy is compromised
Solution Approach 1:
The patent extracts identifying information from user data before sharing. Service providers remove personally identifiable information (PII) such as names, addresses, and phone numbers from datasets before exchanging them with other providers. This extraction process retains behavioral and preference data that is valuable for marketing while eliminating privacy risks associated with direct identification.
Solution Approach 2:
The patent introduces hashed identifiers as an intermediary mechanism for data matching. Instead of sharing raw identifying information, providers use hashed versions of user identifiers that cannot be reversed to obtain original data. These hashed identifiers serve as mediators that enable data correlation across different providers without exposing actual user identities, thus resolving the privacy-information completeness contradiction.
2Object-affected harmful factors
If service providers remove identifying information to maintain user privacy, then user privacy is protected, but redundant data is generated and stored
Solution Approach 1:
The patent merges multiple anonymized datasets by matching records through hashed identifiers. When the same user interacts with different service providers, their data is correlated using the hashed identifier as a key, consolidating information into unified user profiles. This merging eliminates redundant storage of separate anonymized records for the same individual while preserving privacy through the use of hashed identifiers.
Solution Approach 2:
The patent uses hashed copies of identifiers instead of original identifying information. Rather than storing and processing multiple versions of actual user identities, the system creates and exchanges hashed copies that serve the same functional purpose for data matching and correlation. These cryptographic copies consume minimal computational resources while enabling efficient deduplication across distributed datasets.
3Object-affected harmful factors
If service providers store multiple sets of repetitive anonymized data, then user privacy is maintained, but storage space and processing power are wasted
Solution Approach 1:
The patent merges multiple anonymized datasets by matching records through hashed identifiers. When the same user interacts with different service providers, their data is correlated using the hashed identifier as a key, consolidating information into unified user profiles. This merging eliminates redundant storage of separate anonymized records for the same individual while preserving privacy through the use of hashed identifiers.
Data Source
AI summary
Systems and methods are described for distilling data. First data associated with a user may be received. The first data associated with the user may comprise an anonymized hash of an identifier associated with the user. A database may be determined to comprise a first record indicating the anonymized hash. The first record may comprise second data associated with the user. Based on the determining that the database comprises the first record, a second record may be generated. The second record may comprise the first data associated with the user, the second data associated with the user, and the anonymized hash. Based on the determining that the database comprises the first record, the example method may be stored to the database. These and other user and/or data distillation methods and systems are described herein.


