Local Data Preprocessing for AI Model Training

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Users face challenges in training machine-learning models for cloud-based AI services due to the need to transfer large amounts of raw data from user computing networks to service provider networks, which consumes significant bandwidth and may involve sensitive data that users want to keep private.

Innovation Solution

The service provider network provides software plug-ins that can be executed locally in user computing networks to preprocess and convert raw data into smaller, machine-consumable training data, which is then sent to the service provider network for use in training AI models, ensuring data privacy and reducing bandwidth usage.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If raw data is transferred from user computing networks to service provider networks for training ML models, then model training can be performed with user-specific data, but bandwidth consumption increases and data privacy is compromised

Engineering Contradiction:
Improvemodel training with user-specific dataVSAvoidbandwidth consumption
Core Design Contradiction:
Adaptability or versatilityVSLoss of energy

Solution Approach 1:

The patent extracts only the essential training data from raw data at the user's location, rather than transferring all raw data. The connector extracts features and converts raw data into training data locally, sending only the processed training data to the service provider network, thereby reducing bandwidth consumption while still enabling model training with user-specific data

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The patent introduces a connector as an intermediary component that runs locally in the user computing network. This connector acts as a mediator between the raw data source and the service provider network, performing preprocessing, feature extraction, and training data generation locally before transmission, thus reducing the amount of data that needs to be transferred over the network

Inventive Principle:
Principle #24Intermediary (Mediator)

2Adaptability or versatility

If raw data is transferred from user computing networks to service provider networks for training ML models, then model training can be performed with user-specific data, but data privacy is compromised

Engineering Contradiction:
Improvemodel training with user-specific dataVSAvoiddata privacy exposure
Core Design Contradiction:
Adaptability or versatilityVSObject-affected harmful factors

Solution Approach 1:

The patent extracts only the necessary training information from raw data at the user's location, rather than exposing all raw data. The connector extracts features and converts raw data into training data locally, sending only the processed training data to the service provider network, thereby reducing data privacy exposure while still enabling model training with user-specific data

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The patent performs preliminary processing of data at the user's location before data leaves the local network. The connector pre-processes raw data, extracts features, and converts it into training data format locally, so that when data is transmitted to the service provider network, it is already in a processed state that maintains privacy while being suitable for model training

Inventive Principle:
Principle #10Preliminary action

3Loss of energy

If software plug-ins are provided for local pre-processing of raw data, then bandwidth consumption is reduced and data privacy is maintained, but system complexity increases

Engineering Contradiction:
Improvebandwidth consumptionVSAvoidsystem complexity
Core Design Contradiction:
Loss of energyVSDevice complexity

Solution Approach 1:

The patent creates a universal connector that can perform multiple functions: data extraction, feature extraction, training data generation, and model training. This multi-functional connector reduces the need for separate components for each task, thereby managing system complexity while providing comprehensive local pre-processing capabilities

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The patent enables the connector to automatically perform pre-processing operations without requiring complex external orchestration. The connector self-manages the extraction of raw data, feature extraction, training data generation, and model training processes locally, reducing the complexity burden on the service provider network while maintaining reduced bandwidth consumption

Inventive Principle:
Principle #25Self-service

Data Source

PatentUS11620473B1Pre-processing raw data in user networks to create training data for service provider networks
Publication Date: 2023.04.04 AMAZON TECH INC
  • US11620473B1 patent drawing
  • US11620473B1 patent drawing
  • US11620473B1 patent drawing

AI summary

Techniques for a service provider network to provide users with software components that pre-process raw data stored in user computing networks to generate training data that is usable by artificial-intelligence (AI) services. The AI services may utilize models to provide various functionality to users, and the users may desire to train the models with data sets that are specific to their data sets. The service provider network can develop software components that are configured to process raw data into training data for various AI services. The software components can be provided to the user computing networks and executed locally, rather than the raw data having to being moved from the user computing network and to the service provider network. The training data can then be sent to the service provider network and used by the AI services to train ML models for use by the user.