Fuzzy Graph Compression for Privacy-Preserving Machine Learning

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current machine learning systems face challenges in protecting user data privacy, especially with the growing concern of data privacy across various nations, as existing methods do not adequately ensure that data used in these systems remains anonymous while maintaining predictive power.

Innovation Solution

A method involving a server computer that receives network data, generates graphs based on transaction data, determines fuzzy values for data elements, and creates models using these fuzzy values to anonymize user data, making it difficult for malicious parties to identify specific individuals while retaining predictive capabilities.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If data is anonymized to protect user privacy, then data privacy is improved, but predictive power of machine learning models deteriorates

Engineering Contradiction:
Improvedata privacyVSAvoidpredictive power
Core Design Contradiction:
ReliabilityVSLoss of information

Solution Approach 1:

The patent transforms precise data values into fuzzy values by changing the parameter representation from exact numbers to linguistic variables with membership functions. This allows data to retain statistical properties and predictive power while becoming inherently less identifiable, thus resolving the contradiction between privacy protection and predictive capability

Inventive Principle:
Principle #35Parameter changes

Solution Approach 2:

The patent introduces fuzzy logic as an intermediary layer between raw data and machine learning models. This intermediary transformation preserves essential patterns and relationships needed for prediction while removing direct identifiability, enabling both privacy protection and model effectiveness

Inventive Principle:
Principle #24Intermediary (Mediator)

2Loss of information

If precise data values are used to maintain model accuracy, then predictive power is improved, but user identification becomes easier

Engineering Contradiction:
Improvepredictive powerVSAvoiduser identification risk
Core Design Contradiction:
Loss of informationVSObject-affected harmful factors

Solution Approach 1:

By changing data parameters from precise numerical values to fuzzy linguistic variables with ranges and membership degrees, the system maintains statistical integrity for modeling while eliminating exact identifiers that could lead to user re-identification

Inventive Principle:
Principle #35Parameter changes

3Productivity

If data is compressed to reduce storage and processing requirements, then productivity is improved, but data quality and precision deteriorate

Engineering Contradiction:
Improvedata processing efficiencyVSAvoiddata precision
Core Design Contradiction:
ProductivityVSManufacturing precision

Solution Approach 1:

The compression transforms detailed numerical data into fuzzy linguistic variables, significantly reducing data size and processing requirements while preserving the essential statistical properties and patterns needed for machine learning, thus achieving both efficiency and adequate precision

Inventive Principle:
Principle #35Parameter changes

Data Source

PatentUS12079814B2Privacy-preserving graph compression with automated fuzzy variable detection
Publication Date: 2024.09.03 VISA INTERNATIONAL SERVICE ASSOCIATION
  • US12079814B2 patent drawing
  • US12079814B2 patent drawing
  • US12079814B2 patent drawing

AI summary

A disclosed method includes a) receiving by a server computer network data comprising a plurality of transaction data for a plurality of transactions. Each transaction data comprises a plurality of data elements with data values. At least one of the plurality of data elements comprises a user identifier for a user. The server computer can then b) generate one or more graphs comprising a plurality of communities based on the network data. The server computer can c) determine fuzzy values for at least some of the data values for each transaction of the plurality of transactions. For each user. the server computer can d) determine fuzzy values for communities within the plurality of communities. The server computer can then e) generate a model using the fuzzy values obtained in steps c) and d), and at least some of the data values.