Genomic Data Privacy via Trusted Execution and Policy Segmentation
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
The challenge lies in protecting personal genomic and bioinformatic information from reidentification attacks, particularly when genomic datasets are federated across multiple institutions, as existing systems fail to effectively manage access and privacy, risking unintended disclosure of sensitive information.
Innovation Solution
The implementation of policy-managed systems and methods that utilize disclosure accounting and computation-based approaches to securely handle genomic data, ensuring that only authorized access and a controlled amount of information are revealed, with mechanisms like encryption, secure processing units, and information metrics to enforce privacy and security policies.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If genomic data is shared across federated institutions to enable statistical analysis, then the value and utility of the data increases, but the risk of reidentification attacks and privacy breaches increases
Solution Approach 1:
The patent introduces a trusted execution environment (TEE) as an intermediary layer between the genomic data and the analysis processes. The TEE securely hosts the genomic data and computation resources, allowing federated institutions to perform statistical analysis without directly accessing the raw genomic information. This mediator enables data utility while preventing reidentification attacks by ensuring that even the institutions cannot access the plaintext genomic data outside the controlled TEE environment.
Solution Approach 2:
The patent segments the genomic data access control into multiple layers and components. It divides the data access permissions into fine-grained attributes (e.g., research purpose, institution type, data sensitivity level) and implements separate management for different data elements. This segmentation allows the system to provide targeted access to specific data subsets based on institutional credentials and research needs, maximizing data utility while minimizing reidentification risk by limiting exposure of any single individual's complete genomic profile.
2Reliability
If access control policies are enforced to protect individual privacy, then privacy protection improves, but the ability to conduct comprehensive statistical analysis deteriorates
Solution Approach 1:
The patent implements dynamic access control policies that automatically adjust permissions based on the current research context, institutional credentials, and sensitivity of the specific data elements being accessed. The system evaluates each analysis request against predefined privacy policies and computational disclosure thresholds, dynamically granting or denying access in real-time. This dynamic approach allows comprehensive statistical analysis to proceed when privacy risks are low, while automatically restricting access when privacy thresholds are approached, thus maintaining both privacy protection and analysis capability.
Solution Approach 2:
The patent changes the parameters of data access by transforming raw genomic data into derived statistical summaries and aggregated metrics that preserve analytical value while reducing reidentification risk. The system applies parameter transformations such as computing population-level statistics, frequency distributions, and risk scores that maintain scientific utility but eliminate direct links to individual identities. This parameter transformation enables comprehensive statistical analysis while inherently protecting individual privacy.
3Reliability
If derived resources are created through computation to preserve privacy, then privacy preservation improves, but the amount of original information available for analysis decreases
Solution Approach 1:
The patent applies partial action by selectively computing derived resources only for specific data elements and research questions where privacy preservation is critical, rather than transforming all genomic data uniformly. The system identifies which genomic variants and data elements require derivation based on their sensitivity and reidentification risk, applying computation-based privacy preservation only where necessary. This partial approach maintains full access to non-sensitive data for comprehensive analysis while protecting sensitive information through derived resources, thus minimizing information loss while preserving privacy.
Solution Approach 2:
The patent implements feedback mechanisms that monitor the information disclosure level during computational processes and adjust the derivation strategy in real-time. The system tracks the amount of information revealed through derived resources and compares it against predefined disclosure thresholds. When thresholds are approached, the system automatically adjusts computation parameters, selects alternative analysis methods, or requests additional privacy-preserving transformations. This feedback loop ensures that privacy preservation through computation does not unnecessarily eliminate valuable information, maintaining the optimal balance between privacy and analytical utility.
Data Source
Figure 1
Figure 2
Figure 3
AI summary
The present disclosure relates to systems and methods for facilitating trusted handling of genomic bioinformatics, and/or other sensitive information. Certain embodiments may facilitate policy-based governance of access to and/or use of information through enforced disclosure accounting processes. Among other things, embodiments of the disclosed systems and methods may mitigate the potential for various attacks, including reidentification attacks targeting particular individuals associated with information included in a genomic data set.