Spurious Data Injection for Malicious ML Training Detection
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing cybersecurity measures are inadequate in preventing stolen data from being used to train machine learning models, as malicious actors can still access and utilize encrypted data.
Innovation Solution
The introduction of spurious (fake) data into datasets, which are designed to confuse machine learning models, combined with the use of explainable artificial intelligence (XAI) to determine the impact of features on model performance, allowing for the optimal addition of spurious data to degrade model performance without overburdening systems.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If spurious data is added to a dataset to prevent malicious use, then data security is improved, but computational cost and storage requirements increase
Solution Approach 1:
The system implements a feedback loop where machine learning models are trained on the dataset with spurious data, and the model performance is monitored. Based on performance degradation metrics, the system dynamically adjusts the quantity and characteristics of spurious data added to the dataset, optimizing the balance between security and computational resources.
Solution Approach 2:
The system varies multiple parameters of the spurious data including quantity, distribution patterns, feature importance weights, and data characteristics. By changing these parameters adaptively, the system achieves effective malicious actor confusion while minimizing computational and storage overhead.
2Reliability
If spurious data is added to a dataset to prevent malicious use, then data security is improved, but storage requirements increase
Solution Approach 1:
The system uses feedback from model performance monitoring to dynamically adjust the quantity of spurious data stored. When performance degradation is sufficient to deter malicious use, the system stops adding more spurious data, optimizing storage utilization while maintaining security.
Solution Approach 2:
The system adds spurious data in controlled partial quantities rather than overwhelming amounts. By using partial action with strategic placement and characteristics of spurious data, the system achieves adequate security protection without excessive storage requirements.
3Reliability
If more spurious data is added to confuse malicious models, then model performance degradation is improved, but system resources are overburdened
Solution Approach 1:
The system continuously monitors machine learning model performance metrics and uses this feedback to determine when sufficient spurious data has been added to achieve the desired performance degradation, preventing resource overburdening while maintaining security effectiveness.
Solution Approach 2:
The system dynamically adjusts the characteristics and quantity of spurious data based on real-time model performance and system resource utilization. This dynamic adaptation ensures optimal balance between achieving model degradation and avoiding resource exhaustion.
4Reliability
If existing encryption and access control measures are used, then data protection is improved, but malicious actors can still access and use the data
Solution Approach 1:
The system converts the harmful effect of data theft into a beneficial security mechanism by adding spurious data that confuses malicious actors. The stolen data, which would normally be useful to attackers, is transformed into a tool for their confusion and frustration through the strategic inclusion of spurious samples.
Solution Approach 2:
The system performs preliminary action by pre-adding spurious data to the dataset before it is potentially stolen. This anticipatory measure ensures that even if data is stolen, it is already contaminated with confounding information that prevents effective malicious use.
Data Source
AI summary
In some aspects, a computing system obtain a first dataset including a set of original data samples and a first set of spurious data samples. Based on a time period expiring, the computing system may replace the first set of spurious data samples in the first dataset with a second set of spurious data samples. The computing system may obtain an indication that a second dataset is available via a third-party computing device. Based on a determination that a subset of samples of the second dataset correspond to the first set of spurious data samples, the computing system may determine a time window in which an incident occurred. As an example, the time window may be determined to correspond to a time before the first set of spurious data samples were replaced with the second set of spurious data samples.


