Distributed Machine Learning for PHI Data Without Movement
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
The manual and tedious process of gathering and analyzing high-resolution clinical data for research is hindered by legal restrictions on personal health information (PHI), making it difficult to apply existing distributed machine learning techniques to large-scale, cross-institutional medical data sets without violating medical regulations.
Innovation Solution
The SICKBAY platform enables seamless collection and analysis of high-resolution physiologic data, using a hierarchical distributed computing approach that allows researchers to develop and deploy predictive algorithms rapidly across multiple institutions while maintaining data privacy by preventing the movement of PHI data across institutional boundaries, utilizing the SICKBAY-SPARK system for distributed machine learning without an HDFS file system.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If distributed machine learning techniques are applied to large-scale medical data sets, then research productivity and analysis capability are improved, but legal restrictions on PHI data movement prevent direct application of existing techniques
Solution Approach 1:
The system segments the distributed machine learning process into two distinct components: (1) a centralized coordination layer that manages algorithm distribution and result aggregation, and (2) decentralized data processing nodes that execute computations locally on PHI data without transferring it. This segmentation allows the system to maintain data privacy regulations while achieving distributed processing capabilities.
Solution Approach 2:
The patent introduces an intermediary layer consisting of standardized interfaces and protocols that enable communication between the centralized coordination system and decentralized data nodes. This intermediary framework allows algorithms to be distributed and results to be aggregated without requiring direct data movement, thus bridging the gap between centralized control and distributed processing while maintaining PHI security.
2Reliability
If manual data gathering and analysis processes are used to comply with medical regulations, then data privacy is protected, but research time and processing duration increase substantially
Solution Approach 1:
The system performs preliminary actions by pre-configuring standardized data processing pipelines and compliance frameworks at each decentralized node before data processing begins. Legal agreements and data protection mechanisms are established in advance, allowing researchers to immediately begin distributed analysis without manual setup for each project, thus reducing overall research timeline while maintaining compliance.
Solution Approach 2:
The patent changes the operational parameters of data processing by shifting from centralized batch processing to distributed parallel processing across multiple nodes. This parameter change enables simultaneous analysis of multiple data sets while maintaining regulatory compliance through local processing, thereby reducing total research duration without sacrificing data protection standards.
3Power
If existing distributed machine learning techniques require data movement for processing, then computational efficiency is improved, but legal and administrative barriers prohibit data portability
Solution Approach 1:
The patent inverts the traditional distributed machine learning approach by keeping data stationary at decentralized nodes and moving only the computational algorithms and intermediate results. Instead of transporting data to centralized processing units, the system distributes processing capabilities to where the data resides, thereby maintaining computational efficiency while eliminating data portability requirements and associated legal barriers.
Data Source
AI summary
A system for distributed computing allows researchers performing analysis on data having legal or policy restrictions on movement of the data to perform the analysis on data at multiple sites without exporting the restricted data from each site. A primary system determines which sites contain the data to be executed by a compute job and sends the compute job to the associated site. Resulting calculated data can be exported to a primary collecting and recording system. The system allows the rapid development of analysis code for the compute jobs that can be executed on a local data set then passed to the primary system for distributed machine learning on the multiple sites without any changes to the code of the compute job.


