Fuzzy Graph Compression for Privacy-Preserving Machine Learning
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current machine learning systems face challenges in protecting user data privacy, especially with the growing concern of data privacy across various nations, as existing methods do not adequately ensure that data used in these systems remains anonymous while maintaining predictive power.
Innovation Solution
A method involving a server computer that receives network data, generates graphs based on transaction data, determines fuzzy values for data elements, and creates models using these fuzzy values to anonymize user data, making it difficult for malicious parties to identify specific individuals while retaining predictive capabilities.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If data is anonymized to protect user privacy, then data privacy is improved, but predictive power of machine learning models deteriorates
Solution Approach 1:
The patent transforms precise data values into fuzzy values by changing the parameter representation from exact numbers to linguistic variables with membership functions. This allows data to retain statistical properties and predictive power while becoming inherently less identifiable, thus resolving the contradiction between privacy protection and predictive capability
Solution Approach 2:
The patent introduces fuzzy logic as an intermediary layer between raw data and machine learning models. This intermediary transformation preserves essential patterns and relationships needed for prediction while removing direct identifiability, enabling both privacy protection and model effectiveness
2Loss of information
If precise data values are used to maintain model accuracy, then predictive power is improved, but user identification becomes easier
Solution Approach 1:
By changing data parameters from precise numerical values to fuzzy linguistic variables with ranges and membership degrees, the system maintains statistical integrity for modeling while eliminating exact identifiers that could lead to user re-identification
3Productivity
If data is compressed to reduce storage and processing requirements, then productivity is improved, but data quality and precision deteriorate
Solution Approach 1:
The compression transforms detailed numerical data into fuzzy linguistic variables, significantly reducing data size and processing requirements while preserving the essential statistical properties and patterns needed for machine learning, thus achieving both efficiency and adequate precision
Data Source
AI summary
A disclosed method includes a) receiving by a server computer network data comprising a plurality of transaction data for a plurality of transactions. Each transaction data comprises a plurality of data elements with data values. At least one of the plurality of data elements comprises a user identifier for a user. The server computer can then b) generate one or more graphs comprising a plurality of communities based on the network data. The server computer can c) determine fuzzy values for at least some of the data values for each transaction of the plurality of transactions. For each user. the server computer can d) determine fuzzy values for communities within the plurality of communities. The server computer can then e) generate a model using the fuzzy values obtained in steps c) and d), and at least some of the data values.


