Edge Device Model Training with Distilled Data
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
The challenge is to efficiently store and process large amounts of data for training machine learning models in autonomous driving applications, where edge devices have limited storage resources and computing power, making it difficult to achieve high model accuracy with conventional solutions.
Innovation Solution
The method involves using an edge device to train models with distilled data, which represents historical data stored in a remote device, allowing for efficient storage and updating of models with fewer storage resources by using data distillation algorithms to reduce the data volume.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If all historical data is stored at the edge device for model training, then model accuracy is improved, but storage resource consumption increases significantly
Solution Approach 1:
The patent extracts only the essential characteristics and patterns from historical data to create distilled data representations. These distilled data contain the key information needed for model training while occupying minimal storage space at the edge device, thereby achieving high model accuracy without significant storage resource consumption.
Solution Approach 2:
Instead of storing complete historical data, the patent creates compressed distilled copies that capture the essential training information. These distilled data serve as efficient representations that can be stored locally at edge devices, enabling accurate model training with minimal storage requirements compared to storing all original historical data.
2Quantity of substance
If distilled data is used instead of historical data for model training, then storage resource consumption is reduced, but model accuracy may deteriorate
Solution Approach 1:
The patent transforms historical data into distilled data by changing its representation parameters - converting raw data into compressed forms that retain essential training information. This parameter transformation enables the distilled data to maintain sufficient accuracy for model training while dramatically reducing storage requirements.
Solution Approach 2:
The distilled data are designed to have localized quality characteristics that preserve the most critical training information while discarding redundant details. This selective preservation of essential features ensures that model training accuracy is maintained despite the reduced data volume.
3Loss of information
If new data is continuously collected and stored at the remote device, then data completeness is improved, but data processing time and resource consumption increase
Solution Approach 1:
The patent performs preliminary distillation of historical data at the remote device before it needs to be used for training. By pre-processing and compressing the data into distilled form, the system eliminates the need for time-consuming processing when the data is needed at edge devices, thereby reducing data processing time while maintaining data completeness.
Solution Approach 2:
The patent segments the data processing workflow into distinct stages: complete data collection and distillation at the remote device, and deployment of distilled data to edge devices. This segmentation allows comprehensive data processing to occur in advance at powerful remote infrastructure, while edge devices only need to handle the lightweight distilled data, significantly reducing their processing time and resource consumption.
Data Source
AI summary
A method in one embodiment includes receiving, at an edge device, new data for training a model, the edge device having stored distilled data used to represent historical data to train the model, the historical data being stored in a remote device, and the amount of the historical data being greater than the amount of the distilled data. The method further includes training the model based on the new data and the distilled data. With the data processing solution of this embodiment, the model can be trained at the edge device with fewer storage resources based on the distilled data, thereby achieving higher model accuracy.


