ML Compute Request Interface for Automated Data Labeling
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
The complexity of using public cloud services, particularly machine learning workloads, requires specialized knowledge and expertise, making it difficult for average users to migrate and integrate data effectively due to steep learning curves and high costs associated with processing customized data needs.
Innovation Solution
A user-friendly interface and software system that allows users to generate machine learning compute requests without requiring specialized knowledge, utilizing connectors to ingest, label, and process data, and automatically determine the necessary cloud services for processing, thereby simplifying data migration and integration.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Power
If public cloud services are used for machine learning workloads, then computing power and data processing capability are improved, but system complexity and learning curve increase significantly
Solution Approach 1:
The patent introduces an intermediary system that sits between the user and the public cloud services. This intermediary automatically generates machine learning compute requests by ingesting user data, labeling it appropriately, and translating it into cloud service-compatible formats. Users interact with simple file upload interfaces rather than complex cloud service configurations, while the intermediary handles the complexity of cloud service integration and data preparation.
2Productivity
If customized machine learning services are implemented, then data processing capability is improved, but cost increases due to data acquisition and transformation requirements
Solution Approach 1:
The system enables self-service machine learning by allowing users to upload raw data without needing to manually acquire, transform, or prepare data in specialized formats. The intermediary automatically performs data ingestion, labeling, and formatting, eliminating the need for users to invest time and resources in data preparation workflows that would otherwise require specialized expertise and tools.
Solution Approach 2:
The intermediary performs preliminary actions by pre-processing and labeling user data before it reaches the cloud service. This includes automatically generating labels based on file pointers and metadata, preparing data in the correct formats, and creating compute requests in advance. These preliminary actions eliminate the need for users to perform costly and time-consuming data transformation tasks.
3Measurement precision
If manual data labeling and compute request generation are performed, then data accuracy is improved, but time consumption and operational complexity increase
Solution Approach 1:
The patent replaces manual mechanical processes with automated computational processes. Instead of users manually labeling data and constructing compute requests, the system uses automated algorithms to ingest data, generate labels based on file pointers and metadata, and create cloud service requests. This substitution maintains data accuracy through systematic processing while dramatically reducing the time and effort required compared to manual operations.
Data Source
AI summary
The present technology can provide a simple to use interface for receiving a selected machine learning task and one or more file pointers indicating a network location where data to be input in the machine learning task is stored. The present technology can also provide a connector that can ingest the input data from the network location; and automatically label the input data to be suitable for the selected machine learning task. The connector can further generate a machine learning compute request comprising a control information specifying one or more parameters for the selected machine learning task and a machine learning dataset generated from the labeled sequences of input data.


