Edge Device Model Training with Distilled Data

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

The challenge is to efficiently store and process large amounts of data for training machine learning models in autonomous driving applications, where edge devices have limited storage resources and computing power, making it difficult to achieve high model accuracy with conventional solutions.

Innovation Solution

The method involves using an edge device to train models with distilled data, which represents historical data stored in a remote device, allowing for efficient storage and updating of models with fewer storage resources by using data distillation algorithms to reduce the data volume.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If all historical data is stored at the edge device for model training, then model accuracy is improved, but storage resource consumption increases significantly

Engineering Contradiction:
Improvemodel accuracyVSAvoidstorage resource consumption
Core Design Contradiction:
Measurement precisionVSQuantity of substance

Solution Approach 1:

The patent extracts only the essential characteristics and patterns from historical data to create distilled data representations. These distilled data contain the key information needed for model training while occupying minimal storage space at the edge device, thereby achieving high model accuracy without significant storage resource consumption.

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

Instead of storing complete historical data, the patent creates compressed distilled copies that capture the essential training information. These distilled data serve as efficient representations that can be stored locally at edge devices, enabling accurate model training with minimal storage requirements compared to storing all original historical data.

Inventive Principle:
Principle #26Copying

2Quantity of substance

If distilled data is used instead of historical data for model training, then storage resource consumption is reduced, but model accuracy may deteriorate

Engineering Contradiction:
Improvestorage resource consumptionVSAvoidmodel accuracy
Core Design Contradiction:
Quantity of substanceVSMeasurement precision

Solution Approach 1:

The patent transforms historical data into distilled data by changing its representation parameters - converting raw data into compressed forms that retain essential training information. This parameter transformation enables the distilled data to maintain sufficient accuracy for model training while dramatically reducing storage requirements.

Inventive Principle:
Principle #35Parameter changes

Solution Approach 2:

The distilled data are designed to have localized quality characteristics that preserve the most critical training information while discarding redundant details. This selective preservation of essential features ensures that model training accuracy is maintained despite the reduced data volume.

Inventive Principle:
Principle #3Local quality

3Loss of information

If new data is continuously collected and stored at the remote device, then data completeness is improved, but data processing time and resource consumption increase

Engineering Contradiction:
Improvedata completenessVSAvoiddata processing time
Core Design Contradiction:
Loss of informationVSLoss of time

Solution Approach 1:

The patent performs preliminary distillation of historical data at the remote device before it needs to be used for training. By pre-processing and compressing the data into distilled form, the system eliminates the need for time-consuming processing when the data is needed at edge devices, thereby reducing data processing time while maintaining data completeness.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent segments the data processing workflow into distinct stages: complete data collection and distillation at the remote device, and deployment of distilled data to edge devices. This segmentation allows comprehensive data processing to occur in advance at powerful remote infrastructure, while edge devices only need to handle the lightweight distilled data, significantly reducing their processing time and resource consumption.

Inventive Principle:
Principle #1Segmentation

Data Source

PatentUS11989263B2Method, electronic device, and computer program product for data processing
Publication Date: 2024.05.21 EMC IP HLDG CO LLC
  • US11989263B2 patent drawing
  • US11989263B2 patent drawing
  • US11989263B2 patent drawing

AI summary

A method in one embodiment includes receiving, at an edge device, new data for training a model, the edge device having stored distilled data used to represent historical data to train the model, the historical data being stored in a remote device, and the amount of the historical data being greater than the amount of the distilled data. The method further includes training the model based on the new data and the distilled data. With the data processing solution of this embodiment, the model can be trained at the edge device with fewer storage resources based on the distilled data, thereby achieving higher model accuracy.