Zero-Trust Dataset Verification Without Sensitive Data Sharing
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
The challenge of sharing sensitive datasets, such as protected health information, between data stewards and algorithm developers is hindered by the risk of intellectual property leakage, large data transfer times, and regulatory compliance issues, particularly in healthcare settings where data access is restricted.
Innovation Solution
A zero-trust computing environment is implemented using sequestered computing nodes and public-private key techniques to encrypt algorithms and datasets, allowing secure processing without exposing them to unauthorized parties, with systems for dataset selection, verification, and recommendation.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If datasets are shared with algorithm developers for processing, then data analysis capability is improved, but data security and confidentiality deteriorate
Solution Approach 1:
The patent introduces a trusted third-party platform that acts as an intermediary between data stewards and algorithm developers. This platform enables secure data sharing through controlled access mechanisms, allowing algorithm developers to analyze datasets without direct possession of the data, thereby maintaining confidentiality while enabling analysis capability.
Solution Approach 2:
The patent segments the data access process into controlled portions through the intermediary platform. Instead of full data sharing, the platform provides algorithm developers with access only to the specific data subsets and processing capabilities needed for their algorithms, reducing security risks while maintaining analytical utility.
2Productivity
If large datasets are transferred to algorithm developers, then processing capability is improved, but transfer time deteriorates
Solution Approach 1:
The intermediary platform hosts the datasets and provides computational resources to algorithm developers. Instead of transferring large datasets across networks, the platform enables developers to access and process data remotely through secure connections, eliminating transfer time while maintaining processing capability.
Solution Approach 2:
The patent implements virtual copies or references to datasets through the intermediary platform. Algorithm developers work with data representations or accessed copies hosted on the platform rather than requiring physical data transfer, enabling immediate processing without transfer delays.
3Measurement precision
If sensitive data is shared broadly, then algorithm training quality is improved, but regulatory compliance deteriorates
Solution Approach 1:
The intermediary platform implements regulatory compliance mechanisms including access control, audit logging, and data protection measures. This enables algorithm developers to access diverse sensitive datasets for high-quality training while the platform ensures all access conforms to regulatory requirements such as HIPAA, maintaining both training quality and compliance.
Solution Approach 2:
The patent applies different access control and protection measures to different portions of sensitive data based on regulatory requirements. The intermediary platform enables selective access where developers can utilize data subsets appropriate for their algorithms while maintaining compliance through localized security measures applied to each data access instance.
Data Source
AI summary
Systems and methods for the verification of cohort sample sets is provided. In some embodiments, a sample dataset is received, and used to generate a sample vector set. The sample vector is computed by encoding the dataset according to a set of classes, generating a matrix of the encoded dataset (where the rows of the matrix correspond to patients and the columns to a class or subclass), and converting the matrix into a series of vector spaces. An example vector set is received and the difference between the sample vector set and the example vector set. Calculating the difference is by framing the distance as a p-value in a hypothesis test, compared against a threshold. When the p-value is above the threshold the sample dataset is rejected.


