Multi-tenant ML Experimentation via Synthetic Data Sandboxing
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Machine learning models in multi-tenant environments face limitations due to restricted access to real customer data for privacy and security reasons, leading to less accurate results when using artificial data for experimentation and improvements.
Innovation Solution
An experimentation framework that allows secure access to real tenant data for machine learning experiments, using an experimental interface to modify machine learning algorithms and feature engineering, while maintaining data isolation and security through tenant-specific virtual computing engines and access control systems.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If real customer data is used for machine learning experiments, then model accuracy is improved, but data privacy and security are compromised
Solution Approach 1:
The patent creates a synthetic copy of the data environment through a sandbox system that replicates production data characteristics without containing actual sensitive data. The sandbox generates synthetic datasets that preserve statistical properties and relationships needed for accurate model training while eliminating privacy and security risks associated with real customer data.
Solution Approach 2:
The sandbox acts as an intermediary layer between data scientists and the production data environment. It provides controlled access to data-like structures through synthetic copies, enabling experimentation without direct exposure to sensitive real data, thus mediating between the need for accurate modeling and data protection requirements.
2Reliability
If access to real tenant data is restricted for security reasons, then data security is improved, but experimentation quality deteriorates
Solution Approach 1:
Instead of restricting access entirely, the system creates high-fidelity synthetic copies of production data environments that preserve the statistical properties, relationships, and patterns necessary for quality experimentation. Data scientists can conduct experiments with these copies without compromising security, maintaining both security and experimentation quality.
Solution Approach 2:
The system segments the data access problem by separating the actual production data (kept secure) from the experimental data (synthetic copies provided to scientists). This segmentation allows simultaneous maintenance of security for real data and quality for experimentation through the use of synthetic data segments.
3Reliability
If synthetic data is used for machine learning experiments, then data security is maintained, but model accuracy deteriorates
Solution Approach 1:
The sandbox system dynamically adjusts parameters of the synthetic data generation process to match production data characteristics. By changing parameters such as data distributions, relationships, and statistical properties during sandbox initialization, the system creates synthetic data that preserves the accuracy needed for model training while maintaining security.
Solution Approach 2:
The sandbox performs preliminary actions by pre-configuring synthetic data environments with accurate representations of production data characteristics before experimentation begins. This preliminary setup ensures that synthetic data is optimized for model accuracy from the start, rather than being a simple placeholder.
Data Source
AI summary
The system and methods of the disclosed subject matter provide an experimentation framework to allow a user to perform machine learning experiments on tenant data within a multi-tenant database system. The system may provide an experimental interface to allow modification of machine learning algorithms, machine learning parameters, and tenant data fields. The user may be prohibited from viewing any of the tenant data or may be permitted to view only a portion of the tenant data. Upon generating an experimental model using the experimental interface, the user may view results comparing the performance of the experimental model with a current production model.


