Distributed Machine Learning Engine for Siloed Data Privacy
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing predictive analytics solutions face challenges in handling distributed data sets due to assumptions of homogeneous data structures and the need for complex ETL processes, which are costly and infeasible in siloed data environments, especially when data is vertically fragmented or requires real-time analysis.
Innovation Solution
The Distributed DensiCube method and system enable machine learning rule set creation by preparing data identifiers, executing algorithms on distributed data silos, calculating quality control metrics, and combining them, allowing for parallel processing and real-time decision-making without data aggregation.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Ease of manufacture
If data is collected in a single host machine with homogeneous structure, then predictive analytics can be performed using existing tools, but complex and costly ETL processes are required to aggregate distributed data
Solution Approach 1:
The system segments the centralized predictive analytics process into distributed components that operate independently across multiple data silos. Each silo executes local machine learning algorithms on its own data partitions, eliminating the need for centralized ETL aggregation while maintaining analytical capabilities through coordinated rule set combination.
Solution Approach 2:
Instead of aggregating data from distributed silos to a central location (traditional ETL approach), the system inverts the workflow by sending data identifiers to silos and having them return trained rule sets. This reversal eliminates costly data movement and transformation operations while achieving the same predictive analytics objective.
2Reliability
If data is horizontally fragmented across multiple sites, then data privacy is preserved, but existing predictive analytics tools cannot process distributed data sets
Solution Approach 1:
The system introduces data identifiers as an intermediary mechanism that enables communication between distributed data silos and the central coordinating system. These identifiers allow the system to reference and combine results from distributed sources without requiring direct access to the actual data, thus preserving privacy while enabling analytics.
Solution Approach 2:
The system changes the operational parameters from processing raw data to processing data identifiers and rule sets. This parameter transformation allows existing predictive analytics algorithms to operate on distributed data by working with metadata representations rather than the actual data, maintaining compatibility while preserving data location and privacy.
3Productivity
If data is vertically fragmented with different features at different sites, then specialized data storage is achieved, but complex ETL processes are needed to integrate heterogeneous data structures
Solution Approach 1:
The system segments the data integration process by feature type, allowing each data silo to store and process vertically fragmented data locally according to its specialized structure. Machine learning algorithms are applied independently to each feature set, eliminating the need for complex pre-integration ETL processes while maintaining the ability to combine results through the rule set aggregation mechanism.
Data Source
AI summary
A novel distributed method for machine learning is described, where the algorithm operates on a plurality of data silos, such that the privacy of the data in each silo is maintained. In some embodiments, the attributes of the data and the features themselves are kept private within the data silos. The method includes a distributed learning algorithm whereby a plurality of data spaces are co-populated with artificial, evenly distributed data, and then the data spaces are carved into smaller portions whereupon the number of real and artificial data points are compared. Through an iterative process, clusters having less than evenly distributed real data are discarded. A plurality of final quality control measurements are used to merge clusters that are too similar to be meaningful. These distributed quality control measures are then combined from each of the data silos to derive an overall quality control metric.


