Data Mark Classification for Verifying ML Data Removal
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing machine learning models struggle to effectively remove personal data in response to 'right to be forgotten' requests, lacking reliable evaluation mechanisms for data removal effectiveness.
Innovation Solution
A method involving marking data elements with specific marks, training a machine learning model to recognize these marks, applying forgetting mechanisms, and evaluating the effectiveness of data removal through classification tests.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If data elements are removed from a trained machine learning model to fulfill 'right to be forgotten' requests, then data privacy compliance is improved, but the ability to verify complete data erasure and prove compliance to regulatory bodies deteriorates
Solution Approach 1:
The patent applies preliminary action by embedding unique identification marks into data elements before training the machine learning model. These marks are incorporated during the training phase, allowing subsequent verification that the model has genuinely forgotten specific data without requiring post-removal inspection. The marks are prepared in advance and integrated into the training dataset, enabling reliable compliance verification after data removal requests are fulfilled.
Solution Approach 2:
The patent uses identification marks as an intermediary mechanism between the data elements and the verification process. These marks serve as mediators that carry information about the original data's identity without containing the actual data content. After data removal, the presence or absence of these marks in the model's behavior provides verifiable evidence of compliance, bridging the gap between data erasure and proof of erasure.
2Ease of operation
If traditional data removal methods are used, then data deletion is simplified, but the ability to evaluate and rank different forgetting mechanisms deteriorates
Solution Approach 1:
The patent implements feedback by using the identification marks to measure and evaluate the effectiveness of different data removal mechanisms. After applying various forgetting mechanisms, the system queries the model with marked data elements and measures whether the marks are still detectable in the model's responses. This feedback loop enables quantitative comparison and ranking of different data removal approaches, transforming subjective evaluation into objective measurement.
Solution Approach 2:
The patent changes the parameter of data element representation by embedding identification marks that encode specific information about the original data. This parameter change allows the verification system to detect whether specific data elements have been forgotten by checking for the presence or absence of their unique marks, enabling precise evaluation of forgetting mechanism effectiveness without complicating the actual data deletion process.
3Measurement precision
If marked data elements are used for verification, then compliance proof is improved, but the complexity of the machine learning model training process deteriorates
Solution Approach 1:
The patent merges the verification capability directly into the training process by embedding identification marks within the training data elements themselves. Rather than adding separate verification infrastructure, the marks are integrated into the data preprocessing and model training workflow. This merging approach enables compliance verification using the same training and inference infrastructure, reducing overall system complexity despite the enhanced verification capability.
Data Source
AI summary
A method, computer system, and a computer program product for testing a data removal are provided. Data elements are marked with a respective mark per represented entity. The marked data elements, with labels indicating the respective marks, are input into a machine learning model to form a trained machine learning model. The trained machine learning model is configured to perform a dual task that includes a main task and a secondary task that includes a classification based on the labels. A forgetting mechanism is applied to the trained machine learning model to remove a data element including a test mark of the marked data elements. A test data element marked with the test mark is input into the revised machine learning model. The classification of the secondary task of an output of the revised machine learning model is determined for the input test data element.


