Differential Privacy Email Tokenization
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing technologies compromise user privacy when collecting data for learning and spam/malware detection, as users lack control over data collection and analysis, leading to a trade-off between privacy and service usage.
Innovation Solution
A system that allows users to opt-in for data analysis, processes email clear text into tokens with tags, enriches them using user-specific data, and generates feature vectors while ensuring privacy through encryption and differential privacy algorithms, enabling predictive features on client devices without leaking personal data.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If a server collects clear text of user emails for analysis to improve predictive features, then the quality of predictive functionality is improved, but user privacy is compromised
Solution Approach 1:
The patent extracts personally identifiable information (PII) from email clear text through tokenization, separating it from the contextual information needed for predictive features. This allows the server to analyze email patterns and language while removing specific user identifiers, thus improving predictive functionality without compromising user privacy
Solution Approach 2:
The patent introduces an intermediary processing layer that transforms clear text into a privatized format using tokenization and differential privacy algorithms. This intermediary process acts as a mediator between the need for data analysis and privacy protection, enabling feature extraction while maintaining user anonymity
2Productivity
If a server collects user data without user consent for service improvement, then service functionality is enhanced, but user control over personal data is lost
Solution Approach 1:
The patent implements preliminary user consent mechanisms where users are given the option to opt-in or opt-out of data collection before their emails are analyzed. This preliminary action ensures user control is maintained while still allowing service enhancement for those who consent, resolving the contradiction between productivity and user control
3Measurement precision
If a server analyzes clear text of emails for predictive features, then the accuracy of text prediction and natural language processing is improved, but user data security is compromised
Solution Approach 1:
The patent changes the state of user data from clear text to privatized tokens through cryptographic transformations and differential privacy algorithms. This parameter change maintains the statistical properties needed for accurate text prediction while fundamentally altering the data to prevent security compromises and user identification
Data Source
AI summary
Embodiments described herein enable data associated with a large plurality of users to be analyzed without compromising the privacy of the user data. In one embodiment, a user can opt-in to allow analysis of clear text of the user's emails. An analysis process can then be performed in which an analysis service receives clear text of an email of a client device; processes the clear text of the email into one or more tokens having one or more tags; enriches one or more tokens in the processed email using data associated with a user of the client device and the one or more tags; and processes the clear text and one or more enriched tokens to generate a data set of one or more feature vectors.


