Synthetic Data Mirroring for Privacy-Safe ML Testing
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Data sharing, particularly cross-border data sharing for research and development, is hindered by privacy and security concerns, making it difficult to obtain high-quality testing data for machine learning applications.
Innovation Solution
A method and system that generate synthetic data using machine learning models to replicate the features and structural characteristics of real data, allowing secure and compliant data sharing without exposing sensitive information.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If real data is shared for research and development purposes, then the quality of testing data improves, but privacy and security concerns worsen
Solution Approach 1:
The patent creates synthetic data that copies the statistical properties, distributions, and relationships of real data without containing actual sensitive information. Machine learning models generate artificial datasets that replicate the structural characteristics and feature correlations of source data, enabling research and development while eliminating privacy and security risks associated with sharing real data.
2Reliability
If data localization regulations are implemented, then data security improves, but the ability to share data with external partners deteriorates
Solution Approach 1:
The patent introduces synthetic data as an intermediary that mediates between data security requirements and data sharing needs. Instead of sharing real data directly, organizations can share synthetic data with external partners, researchers, and developers. This intermediary form maintains security compliance while enabling collaboration and development activities that would otherwise be restricted by data localization regulations.
3Reliability
If synthetic data is generated using machine learning models, then data security improves, but the complexity of the data generation process worsens
Solution Approach 1:
The patent implements self-service capabilities where the synthetic data generation system automatically analyzes source data characteristics, selects appropriate generation methods, and produces synthetic datasets without requiring extensive manual configuration or expert intervention. The system autonomously handles data profiling, model selection, parameter optimization, and quality validation, reducing operational complexity while maintaining security benefits.
Data Source
AI summary
At least one processor may receive a sample data set, determine at least one feature of data in the sample data set, and determine at least one structural characteristic of the sample data set. The at least one processor may determine that at least a portion of the data is categorical data from the at least one feature and the at least one structural characteristic. By operating a machine learning (ML) model, the at least one processor may generate synthetic data having the same at least one feature as the categorical data. The at least one processor may package the synthetic data into a synthetic data set having the same at least one feature and at least one structural characteristic as the sample data set.


