Machine Learning Model Data Removal for Compliance Recovery
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing technologies lack effective methods for managing and addressing issues with training data used in machine learning models, such as inappropriate, unethical, or legally problematic data, which can degrade model performance and compliance.
Innovation Solution
A model management device that identifies inappropriate training data and updates or retrain machine learning models using a new data set by deleting identified problematic data, incorporating features like data and model identifiers, and a processor for data management.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If machine learning models are trained using available training data, then model performance can be improved, but the model may be degraded by inappropriate, unethical, or legally problematic data
Solution Approach 1:
The system performs preliminary identification and management of training data before model training occurs. By establishing data identification information and usage information in advance, the system can preemptively detect and remove inappropriate data, preventing harmful factors from degrading model performance
Solution Approach 2:
The system implements a feedback mechanism where usage information of training data is continuously monitored and fed back to identify models trained with problematic data. This feedback loop enables the system to detect issues and execute corrective processes to maintain model reliability
2Reliability
If training data is managed and processed to remove inappropriate data, then data quality and compliance are improved, but additional processing steps and system complexity are introduced
Solution Approach 1:
The model management device performs multiple functions including data identification, model identification, and training data management within a single system. This multi-functionality reduces the need for separate specialized systems, thereby limiting the increase in overall device complexity while maintaining data compliance
Solution Approach 2:
The system automatically manages training data by utilizing data identification information and usage information to identify and process problematic data without requiring extensive external intervention. This self-service capability reduces operational complexity while ensuring data compliance
Data Source
AI summary
A technique for managing training data used to train a machine-learning model is disclosed. One aspect of the present disclosure relates to a model management device comprising: a data identifying unit that acquires data identifying information; a model identifying unit that, with reference to usage information of training data, identified a first machine-learning model which has been trained by using first training data identified by the data identifying information; and a processing unit that deletes the first training data from a first training data set used to train the first machine-learning model, and thereby creates a second training data set and executes processing the first machine-learning model.


