Cloud ML Environment with Self-Service Data Cleansing
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current machine learning (ML) activities require significant effort and resources to set up and manage GPU farms, and existing solutions do not effectively address the need for self-service data selection and protection in cloud environments, compromising data integrity and confidentiality.
Innovation Solution
A self-service, auto-prep and cleanse cloud machine learning environment system that includes a data source interface, interactive user interface, and a processor for data provisioning, cleansing, and applying ML analytics, enabling users to create ML instances in a cloud services platform while maintaining data confidentiality and integrity through encryption and classification rules.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If GPU farms or clusters are used to perform machine learning activities, then machine learning performance is improved, but device complexity and setup effort increase significantly
Solution Approach 1:
The system enables users to independently provision and manage machine learning resources through automated interfaces. Users can self-service by selecting pre-configured GPU environments, data sources, and machine learning frameworks without requiring manual IT intervention for setup, thereby maintaining high ML performance while reducing setup complexity
Solution Approach 2:
The system performs preliminary configuration of GPU farms, data sources, and machine learning environments before users need them. Pre-configured templates and automated provisioning scripts prepare the infrastructure in advance, allowing users to quickly deploy machine learning activities without extensive setup effort
2Ease of operation
If data is transferred to cloud data storage for machine learning analytics, then machine learning accessibility is improved, but data confidentiality and integrity may be compromised
Solution Approach 1:
The system introduces an intermediary layer between the data sources and cloud storage that performs automated data classification and cleansing. This intermediary applies security policies, masks sensitive information, and validates data integrity before transfer, enabling cloud-based machine learning while protecting data confidentiality
Solution Approach 2:
The system creates processed copies of data for cloud storage rather than transferring original sensitive data. Data cleansing and classification processes generate sanitized versions suitable for machine learning analytics, maintaining confidentiality while enabling accessibility
3Productivity
If automated data provisioning and cleansing is implemented, then productivity is improved, but system complexity increases
Solution Approach 1:
The automated data provisioning and cleansing system operates autonomously using pre-defined classification rules and policies. The system self-manages data classification, cleansing, and provisioning workflows without requiring manual configuration or complex intervention, thereby improving productivity while keeping system complexity manageable through automation
Data Source
AI summary
An embodiment of the present invention is directed to leveraging GPU farms for machine learning where the selection of data is self-service. The data may be cleansed based on a classification and automatically transferred to a cloud services platform. This allows an entity to leverage the commoditization of the GPU farms in the public cloud without exposing data into that cloud. Also, an entire creation of a ML instance may be fully managed by a business analyst, data scientist and/or other users and teams.


