Confidential Data Backend for Smoothed Posterior Distribution
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
There is a challenge in collecting and maintaining confidential data in computer systems while ensuring user privacy, as users are hesitant to share sensitive information due to concerns about data security and misuse, particularly in balancing the quality of statistical insights with data coverage.
Innovation Solution
A system is developed that securely collects and processes confidential data by using a confidential data frontend and backend architecture, employing encryption techniques to separate user identification from confidential data, anonymizing data points, and employing log-linear models to estimate posterior distributions for cohorts with small sample sizes, thereby enhancing data privacy and quality.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Loss of information
If confidential data is collected and stored in a computer system, then statistical insights and data coverage are improved, but user privacy and data security are compromised
Solution Approach 1:
The patent introduces a confidential data backend as an intermediary component that acts as a trusted mediator between users and the data collection system. This backend securely stores confidential data, performs statistical computations, and returns only aggregated insights without exposing raw data, thus resolving the contradiction between gaining statistical value and protecting user privacy
Solution Approach 2:
The patent extracts and separates user identification information from confidential data through anonymization techniques. By removing or masking identifying elements, the system maintains statistical utility while eliminating privacy risks, allowing data to be used for insights without compromising user identity
2Measurement precision
If groupings are selected from a larger number of samples, then the quality of statistical insights is improved, but the coverage is reduced
Solution Approach 1:
The patent implements dynamic grouping strategies that can adjust the granularity and composition of data groupings based on available sample sizes and statistical requirements. This allows the system to optimize between insight quality and coverage by adapting groupings to current data conditions rather than using fixed groupings
Solution Approach 2:
The patent merges multiple data groupings and samples through the confidential data backend, combining information from various cohorts and segments. This merging approach enables the system to maintain comprehensive coverage while still producing high-quality statistical insights by aggregating data across multiple groups
Data Source
AI summary
In an example, a set of cohort types and an anonymized set of confidential data data values for a plurality of cohorts having cohort types in the set of cohort types are obtained. Then it is determined, from a set of candidate data transformations, a best fitting data transformation for the anonmyized set of confidential data data values. The anonymized set of confidential data data values is transformed using the best fitting data transformation. Optimal smoothing parameters are computed for each cohort type. Then, for each cohort in the set of cohort types having a small sample size, a best parent for the cohort is determined and a posterior distribution for the cohort is determined based on the best parent for the cohort and the optimal smoothing parameters for a cohort type for the cohort.


