Data Leakage Prevention via Blockchain Validation for ML Engines
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing electronic devices face challenges in preventing data leakage to machine learning (ML) engines, which can compromise user privacy by using personal data without user knowledge or consent.
Innovation Solution
A method is introduced that involves detecting data requests from ML engines, creating data blocks based on the requested data and the ML engine's category and tag, and using an activity blockchain to determine and enforce data sharing validity, ensuring that only valid data blocks are shared with the ML engines.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If ML engines access user data directly for training purposes, then the ML engine can learn from diverse data sources, but user privacy and data security are compromised
Solution Approach 1:
The patent introduces an intermediary data validation mechanism that sits between the ML engine and user data. The system validates data requests, creates data blocks with metadata, and controls data flow to ML engines without direct access. This intermediary layer prevents data leakage while still allowing trained ML engines to access necessary data through controlled interfaces.
Solution Approach 2:
The patent segments data into discrete data blocks with associated metadata and validation tags. Each data block is independently validated and tracked, allowing the system to control access to specific data portions while preventing unauthorized access to the entire dataset. This segmentation enables selective data sharing while maintaining security.
2Adaptability or versatility
If multiple ML engines are trained separately on similar data, then each engine can be optimized for specific tasks, but resource usage increases due to redundant data extraction
Solution Approach 1:
The patent creates a universal data validation and management system that serves multiple ML engines simultaneously. The same validation mechanism, data block structure, and metadata system are reused across different ML engines, eliminating redundant validation processes and reducing overall resource consumption while maintaining engine-specific optimization capabilities.
Solution Approach 2:
The patent performs preliminary data validation, categorization, and block creation before data is consumed by ML engines. Data is pre-processed into validated blocks with metadata tags that can be efficiently reused by multiple engines. This preliminary action eliminates the need for each engine to independently extract and validate the same data, reducing redundant resource usage.
Data Source
AI summary
A method for preventing data leakage may include: identifying data that is generated by at least one framework application in response to a data request from a first machine learning (ML) engine of a plurality of ML engines; creating a plurality of data blocks based on the generated data, a category of the first ML engine, and a tag associated with the first ML engine and the at least one framework application; determining whether the plurality of data blocks are valid to share with the first ML engine using an activity block chain associated with each of the plurality of framework applications; based on the plurality of data blocks being valid, sharing the plurality of data blocks with the first ML engine, and otherwise discarding the plurality of data blocks not to share with the first ML engine.


