Confidential Data Insights with Confidence Values
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Users are reluctant to share confidential data, such as salary information, due to privacy concerns about data security and usage within computer systems, making it challenging to collect and maintain accurate and reliable confidential data for statistical analysis.
Innovation Solution
A system that securely collects and tracks confidential data by using a confidential data frontend to gather information from users, encrypting it separately from identifying information, and anonymizing it to ensure security and accuracy, with mechanisms to incentivize users by providing insights based on their contributions.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If confidential data is collected and stored for statistical analysis, then the accuracy and reliability of insights are improved, but user privacy concerns and security risks increase
Solution Approach 1:
The patent segments confidential data into multiple independent tables (submission table, cohort table, insights table) with different access permissions. Each table stores specific aspects of the data (raw submissions, processed cohorts, final insights) and can be accessed by different user roles (admin, user, anonymous) according to their needs, thereby maintaining data accuracy while protecting user privacy through structural separation.
2Reliability
If confidential data is encrypted and anonymized, then user security is improved, but the complexity of data management increases
Solution Approach 1:
The patent introduces an intermediary processing layer that automatically transforms raw confidential data into anonymized cohort data through automated generalization rules. This intermediary system handles the complexity of encryption and anonymization internally, presenting a simplified interface to users while maintaining strong security protections. The system mediates between raw data collection and final insight generation, managing security complexity automatically.
3Quantity of substance
If users are incentivized to share confidential data, then the quantity of data is improved, but the risk of data misuse increases
Solution Approach 1:
The patent implements a feedback mechanism where users receive personalized insights generated from aggregated confidential data in return for their submissions. This feedback loop incentivizes continued data sharing while the system maintains security through automated anonymization and access controls. The feedback principle creates a mutually beneficial relationship: users get value from the service while the system accumulates more data for improved insights.
4Reliability
If data is stored in separate encrypted tables, then data security is improved, but the difficulty of detecting and measuring data relationships increases
Solution Approach 1:
The patent performs preliminary processing of confidential data by automatically generalizing specific user attributes into broader cohort categories before storage. This preliminary action creates pre-defined relationships between data points that can be efficiently detected later without requiring complex real-time analysis of encrypted raw data. The system prepares the data structure in advance to facilitate relationship detection while maintaining security.
Data Source
AI summary
In an example, a plurality of previously submitted confidential data values of a first confidential data type retrieved for a slice having one or more attributes. For a confidential data type, one or more submitted confidential data values of the confidential data type from the slice that are considered outliers based on an external data set or internal data set. A confidence score is calculated by multiplying a support score for the confidential data type in the slice by a non-outlier score for the confidential data type in the slice, the support score being equal to n′/(n′+c), where c is a smoothing constant and n′ is the number of non-excluded submitted confidential data values of the confidential data type in the slice and the non-outlier score being equal to n′/n, where n is the total number of non-null submitted confidential data value of the confidential data type in the slice.


