Generative AI Synthetic Data for AI Model Verification
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
When personal data is deleted due to a request for deletion, it becomes difficult to verify the behavior of artificial intelligence (AI) models trained on that data.
Innovation Solution
An information processing system that uses generative AI to generate artificial data for training AI models, ensuring that the artificial data does not include identifiable information, thus allowing for the verification of AI behavior even after personal data deletion.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If personal data is deleted in response to a deletion request, then compliance with data protection regulations is improved, but the ability to verify AI model behavior deteriorates
Solution Approach 1:
The patent creates a copy of the personal data (user data) and uses this copy to train the AI model. This copy can then be deleted without affecting the model's training, as the model has already learned from the data copy. This resolves the contradiction by allowing deletion of original personal data while preserving the ability to verify AI behavior through the trained model.
Solution Approach 2:
The patent performs preliminary training of the AI model using copies of personal data before the actual deletion request is processed. This preliminary action ensures that the model is already trained and verified before deletion occurs, allowing compliance with deletion regulations while maintaining verification capability through the pre-trained model.
2Loss of information
If personal data is used to train AI models, then the AI model can verify its behavior, but the ability to comply with data deletion requests deteriorates
Solution Approach 1:
The patent creates a copy of the personal data for training purposes and keeps this copy separate from the original data. When a deletion request is made, only the original personal data is deleted, while the copy remains available for model verification. This resolves the contradiction by separating the verification function from the original personal data.
Solution Approach 2:
The patent segments the data into original personal data and training data copies. This segmentation allows the original data to be deleted while the copies remain for verification purposes. The segmentation enables independent handling of data for different purposes (compliance vs. verification), resolving the contradiction between deletion compliance and verification capability.
3Object-affected harmful factors
If generative AI is used to create artificial data, then data privacy is improved, but the complexity of the data processing system increases
Solution Approach 1:
The patent introduces generative AI as an intermediary component that creates artificial data from personal data. This intermediary layer transforms personal data into synthetic training data, improving privacy while managing complexity through a dedicated transformation layer rather than direct manipulation of personal data.
Solution Approach 2:
The patent uses generative AI to create copies of personal data in the form of artificial data. These copies preserve the necessary patterns and characteristics for model training while eliminating direct references to actual individuals. This copying approach improves privacy by working with synthetic rather than real personal data, managing the complexity through standardized generation processes.
Data Source
AI summary
The information processing system includes a generator configured to generate, by generative AI trained using user data including information with which an individual is identifiable, a plurality of pieces of artificial data that is used to train a computational model different from the generative AI and does not include information with which an individual is identifiable.

