Cloud Data Deduplication for Edge ML Storage Optimization
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Edge cloud environments face limitations in resources such as computing power and storage, making it impractical to maintain large data sets for machine learning, leading to reduced user experience and inaccurate model outputs due to the need for rerunning models multiple times.
Innovation Solution
A data materialization platform that migrates data between core and edge cloud environments, using boundary derivation to deduplicate records with similar attribute values, conserving resources and enabling more effective training of machine learning models.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If large data sets are maintained in edge cloud environment for machine learning, then model accuracy is improved, but storage resources are exhausted
Solution Approach 1:
The patent extracts and removes duplicate records from the data set before storing it in the edge cloud environment. By identifying and eliminating redundant data through deduplication processes, the system reduces the quantity of data that needs to be stored while maintaining the essential information needed for accurate machine learning model training.
Solution Approach 2:
The patent performs deduplication actions before data storage occurs. By pre-processing the data to remove duplicates prior to migration to the edge cloud, the system prevents storage resources from being consumed by redundant data, thereby enabling accurate model training without exhausting storage capacity.
2Measurement precision
If data migration and processing operations are performed frequently, then model training accuracy is improved, but processing time and resource consumption increase
Solution Approach 1:
The patent performs deduplication as a preliminary action before data migration to the edge cloud. By removing duplicate records in advance, the system reduces the amount of data that needs to be processed during model training, thereby decreasing processing time and resource consumption while maintaining training accuracy.
Solution Approach 2:
The patent extracts and removes duplicate data records before they enter the edge cloud environment. This extraction of redundant information reduces the overall data volume that requires processing, thereby reducing the time and computational resources needed for model training operations.
3Reliability
If duplicate records are retained in the data set, then data completeness is maintained, but storage efficiency deteriorates
Solution Approach 1:
The patent extracts and removes duplicate records from the data set while preserving unique information. Through deduplication processes that identify and eliminate redundant data, the system maintains data completeness in terms of unique information while significantly improving storage efficiency by reducing the total data volume.
Solution Approach 2:
The patent performs deduplication as a preliminary processing step before data storage. By removing duplicate records in advance, the system ensures that only essential unique data is stored, thereby maintaining data completeness regarding unique information while optimizing storage efficiency and reducing resource consumption.
Data Source
AI summary
In some implementations, a data materialization platform may perform a data migration process between a core cloud environment and an edge cloud environment. The data materialization platform may identify, in association with the data migration process, attribute values stored in a data repository of the core cloud environment, wherein the attribute values are to be used as inputs for a machine learning model that is to be executed by the edge cloud environment. The data materialization platform may analyze the attribute values to identify ranges, associated with a subset of the attribute values, for which outputs of the machine learning model are estimated to be approximately a same output. The data materialization platform may deduplicating, by the data materialization platform, the subset of the attribute values from the data repository of the core cloud environment to generate deduplicated attribute values that are associated with median range values.


