Local Data Preprocessing for AI Model Training
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Users face challenges in training machine-learning models for cloud-based AI services due to the need to transfer large amounts of raw data from user computing networks to service provider networks, which consumes significant bandwidth and may involve sensitive data that users want to keep private.
Innovation Solution
The service provider network provides software plug-ins that can be executed locally in user computing networks to preprocess and convert raw data into smaller, machine-consumable training data, which is then sent to the service provider network for use in training AI models, ensuring data privacy and reducing bandwidth usage.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If raw data is transferred from user computing networks to service provider networks for training ML models, then model training can be performed with user-specific data, but bandwidth consumption increases and data privacy is compromised
Solution Approach 1:
The patent extracts only the essential training data from raw data at the user's location, rather than transferring all raw data. The connector extracts features and converts raw data into training data locally, sending only the processed training data to the service provider network, thereby reducing bandwidth consumption while still enabling model training with user-specific data
Solution Approach 2:
The patent introduces a connector as an intermediary component that runs locally in the user computing network. This connector acts as a mediator between the raw data source and the service provider network, performing preprocessing, feature extraction, and training data generation locally before transmission, thus reducing the amount of data that needs to be transferred over the network
2Adaptability or versatility
If raw data is transferred from user computing networks to service provider networks for training ML models, then model training can be performed with user-specific data, but data privacy is compromised
Solution Approach 1:
The patent extracts only the necessary training information from raw data at the user's location, rather than exposing all raw data. The connector extracts features and converts raw data into training data locally, sending only the processed training data to the service provider network, thereby reducing data privacy exposure while still enabling model training with user-specific data
Solution Approach 2:
The patent performs preliminary processing of data at the user's location before data leaves the local network. The connector pre-processes raw data, extracts features, and converts it into training data format locally, so that when data is transmitted to the service provider network, it is already in a processed state that maintains privacy while being suitable for model training
3Loss of energy
If software plug-ins are provided for local pre-processing of raw data, then bandwidth consumption is reduced and data privacy is maintained, but system complexity increases
Solution Approach 1:
The patent creates a universal connector that can perform multiple functions: data extraction, feature extraction, training data generation, and model training. This multi-functional connector reduces the need for separate components for each task, thereby managing system complexity while providing comprehensive local pre-processing capabilities
Solution Approach 2:
The patent enables the connector to automatically perform pre-processing operations without requiring complex external orchestration. The connector self-manages the extraction of raw data, feature extraction, training data generation, and model training processes locally, reducing the complexity burden on the service provider network while maintaining reduced bandwidth consumption
Data Source
AI summary
Techniques for a service provider network to provide users with software components that pre-process raw data stored in user computing networks to generate training data that is usable by artificial-intelligence (AI) services. The AI services may utilize models to provide various functionality to users, and the users may desire to train the models with data sets that are specific to their data sets. The service provider network can develop software components that are configured to process raw data into training data for various AI services. The software components can be provided to the user computing networks and executed locally, rather than the raw data having to being moved from the user computing network and to the service provider network. The training data can then be sent to the service provider network and used by the AI services to train ML models for use by the user.


