Privacy Protected Database Querying via Thresholding and Bucketizing
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing data analytics systems for music streaming services face challenges in providing detailed user insights while protecting listener privacy, as they often share sensitive information that could identify individual users.
Innovation Solution
Implementing a system that uses thresholding, bucketizing, and rounding techniques to generate anonymized and aggregated data counts, allowing for flexible and specific queries without compromising user privacy, by comparing query results to thresholds, assigning them to numerical buckets, and rounding values to protect listener identities.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Loss of information
If detailed user data is shared to provide meaningful insights, then data utility is improved, but user privacy is compromised
Solution Approach 1:
The patent segments user data into aggregated groups rather than sharing individual records. By dividing the data universe into cohorts based on shared characteristics and analyzing segments collectively, the system maintains data utility for insights while preventing identification of individual users within each segment.
Solution Approach 2:
The patent introduces an intermediary processing layer between raw user data and shared insights. This intermediary applies thresholding, bucketizing, and rounding transformations that mediate between detailed data and privacy protection, allowing useful information to pass through while filtering out identifying characteristics.
2Object-affected harmful factors
If data is aggregated and anonymized to protect privacy, then user privacy is improved, but data precision deteriorates
Solution Approach 1:
The patent changes the parameters of data representation through thresholding (setting minimum counts), bucketizing (grouping into ranges), and rounding (approximating values). These parameter transformations protect privacy by removing precise identifying information while retaining sufficient precision for meaningful analytics through controlled approximation.
Solution Approach 2:
The patent applies partial anonymization rather than complete aggregation. By selectively applying privacy techniques only where necessary (through configurable thresholds and bucket sizes), the system maintains high precision for most data while applying privacy protection only where identification risk exists, avoiding excessive loss of information.
3Object-affected harmful factors
If thresholding and bucketizing are applied to protect privacy, then privacy protection is improved, but data specificity deteriorates
Solution Approach 1:
The patent makes privacy protection dynamic through configurable thresholds and bucket sizes rather than applying fixed anonymization. The system can adjust the level of aggregation based on query context, data sensitivity, and required specificity, allowing optimal balance between privacy protection and data usefulness for different scenarios.
Solution Approach 2:
The patent applies different levels of aggregation to different data segments based on local requirements. Rather than uniformly anonymizing all data, the system applies thresholding and bucketizing selectively to specific cohorts or data types where privacy risk exists, maintaining high specificity in areas where privacy is not a concern.
Data Source
AI summary
Systems, methods, and devices are described herein to query data in an entity data database. A query for information about a set of entities is received from a requesting system. The query is in a predefined format and includes search conditions. A querying strategy is determined based on the received query. The entity data database is queried by identifying a set of user records that fulfill a first search condition. A numerical value of the set of user records is next compared to a threshold. Depending on the numerical value, the set of user records is assigned to a first numerical bucket. Depending on the bucket, the numerical value is changed to a second numerical value, which is used to generate an aggregated count value. The aggregated count value is shared with the requesting system.


