Confidential Data Anomaly Detection via Matrix Factorization
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
There is a challenge in collecting and maintaining confidential data in computer systems, particularly in ensuring the security and anonymity of sensitive information like salary compensation data, where users are hesitant to share due to privacy concerns, and sparse data combinations hinder meaningful statistical insights.
Innovation Solution
A system utilizing computerized matrix factorization and completion techniques, along with secure data encryption and anonymization methods, to collect, track, and utilize confidential data while ensuring security and anonymity, by using a confidential data frontend and backend architecture that separates and encrypts user identification and data, and employs matrix factorization to infer median/mean values for sparse data combinations.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Loss of information
If confidential data is collected and stored in a computer system, then statistical analysis capabilities are improved, but user privacy security deteriorates due to concerns over data access and utilization
Solution Approach 1:
The patent segments confidential data into two separate encrypted components: encrypted cohort identifiers and encrypted confidential values. This segmentation prevents reconstruction of individual user identities while preserving statistical analysis capabilities through secure query processing that operates on segmented data structures.
Solution Approach 2:
The patent introduces encrypted cohort identifiers as an intermediary layer between user identities and confidential data values. This intermediary enables statistical queries to be performed on aggregated data without exposing individual user identities or raw confidential values, thus maintaining privacy security while enabling analysis.
2Reliability
If data encryption and anonymization methods are implemented, then user privacy security is improved, but data accuracy deteriorates due to potential loss of meaningful insights
Solution Approach 1:
The patent changes the parameter representation by using encrypted cohort identifiers instead of raw user identities, and encrypted confidential values instead of plain data. This parameter transformation maintains data accuracy for statistical analysis while ensuring privacy security through encryption that preserves the mathematical relationships needed for meaningful insights.
3Reliability
If separate encryption of user identification and data is implemented, then data security is improved, but system complexity deteriorates due to additional encryption and decryption operations
Solution Approach 1:
The patent performs preliminary encryption of both cohort identifiers and confidential values during data ingestion, and pre-establishes the encrypted data structure with associated metadata. This preliminary action eliminates the need for complex real-time encryption/decryption operations during query processing, reducing system complexity while maintaining security.
4Loss of information
If matrix factorization techniques are used to infer values for sparse cohorts, then statistical analysis capability is improved, but data reliability deteriorates due to potential introduction of erroneous inferences
Solution Approach 1:
The patent applies matrix factorization techniques selectively only to sparse cohorts where data is insufficient, rather than to all data. This partial action allows the system to infer values for specific missing data points while leaving dense cohorts unchanged, thus improving statistical analysis capability without compromising the reliability of well-established data.
Data Source
AI summary
In an example, for each value of a plurality of values of a first attribute of members of a social networking service who have submitted confidential data, an allowed range for normalized confidential data values submitted by members having the value for the first attribute, across all values of a second attribute, is calculated, and then shifted based on an inferred median confidential data value relative to a median of confidential data values. Then, anomalous confidential data values can be detected using this information.


