Zero-Trust Dataset Selection Using Sequestered Vector Matching
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
The challenge of sharing sensitive datasets, such as protected health information, between data stewards and algorithm developers is hindered by the risk of intellectual property leakage and compliance with regulations like HIPAA, which makes data transfer time-consuming and difficult, limiting the adoption of clinical AI applications.
Innovation Solution
A zero-trust computing environment is implemented using sequestered computing nodes and public-private key techniques to encrypt algorithms and datasets, allowing secure processing without exposing them to unauthorized parties, with systems for dataset selection, verification, and recommendation to ensure data quality and matching consumer needs.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If datasets are shared with algorithm developers for processing, then algorithm development and training can proceed, but data security and intellectual property protection are compromised
Solution Approach 1:
A secure computing environment acts as an intermediary between data stewards and algorithm developers. The environment receives encrypted datasets from data stewards, processes them with provided algorithms in an isolated manner, and returns results without allowing direct access to either the data or the algorithm code. This mediator approach enables collaboration while maintaining security boundaries.
Solution Approach 2:
Instead of sharing original sensitive datasets, the system creates and processes encrypted copies within the secure environment. The data is replicated in an isolated computational space where it can be processed without exposing the original sensitive information to external parties.
2Productivity
If large datasets are transferred from data stewards to algorithm developers, then comprehensive analysis can be performed, but transfer time and network resources are significantly consumed
Solution Approach 1:
The system extracts only the essential computational functionality from the data processing workflow and moves it to the data steward's environment through the secure computing platform. Instead of transferring terabytes of data, only the algorithm code and processing instructions are transmitted, while the actual data processing occurs in-place within the secure environment.
Solution Approach 2:
The system transitions from a traditional data-centric model (moving data across network) to a compute-centric model (moving processing capability). By changing the dimension of the solution from data transfer to algorithm deployment, the system avoids the bottlenecks of network transmission while maintaining analytical capabilities.
3Adaptability or versatility
If sensitive health information is shared outside controlled environments, then clinical AI models can be trained on diverse data, but compliance with regulations like HIPAA becomes difficult to maintain
Solution Approach 1:
The secure computing environment functions as an inert or isolated atmosphere where sensitive health information can be processed without exposure to external regulatory risks. The environment is designed with inherent security controls, access restrictions, and audit capabilities that automatically ensure HIPAA compliance, eliminating the need for complex external compliance management.
Solution Approach 2:
The secure computing environment provides a universal platform that handles multiple functions: data processing, security enforcement, compliance monitoring, and result verification. This multi-functional system replaces the need for separate compliance mechanisms and simplifies the overall regulatory adherence process.
Data Source
AI summary
Systems and methods for the selection, verification and recommendation of cohort sample sets is provided. In some embodiments, a dataset selection optimization includes first receiving at data stewards classes of data required by the data consumer. The data stewards process their data (or a subset of their data) into a vector set within a sequestered computing node. These vector sets are transferred to a core management system for minimizing a difference between a target vector and any combination of the data stewards' vector sets. A cost function may also be applied to the vector sets during this optimization. Once the data steward(s) that best match the target vector are identified, they may be placed in contact with the data consumer for access of their information.


