Genomic Data Anonymization Using SNP Re-Identification Risk Scoring
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing methods for genomic data anonymization either remove valuable research information or preserve identifying information, risking privacy breaches, without ensuring sufficient anonymization.
Innovation Solution
A method and system that calculate a re-identification risk score based on phenotypic probabilities and population proportions, masking SNPs if the score exceeds a threshold to ensure privacy while retaining relevant data.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If genomic data is shared and accessed by multiple researchers, then research collaboration and scientific progress are improved, but the risk of re-identification and privacy breaches increases
Solution Approach 1:
The patent introduces secure multi-party computation (SMPC) and homomorphic encryption as intermediary technologies that enable genomic data analysis without direct access to the raw data. These cryptographic protocols act as mediators between data owners and researchers, allowing collaborative analysis while maintaining privacy through mathematical transformations that preserve data utility without exposing sensitive information.
Solution Approach 2:
The patent creates cryptographic copies and transformed representations of genomic data that preserve analytical utility while removing identifying information. Through techniques like homomorphic encryption, the system works with encrypted copies of the data that can be processed and analyzed without revealing the underlying sensitive genomic information, enabling sharing without exposure.
2Object-affected harmful factors
If data anonymization techniques are applied to protect privacy, then privacy protection is improved, but data utility and research value may deteriorate
Solution Approach 1:
The patent transforms data parameters through cryptographic operations such as homomorphic encryption and secure multi-party computation. These parameter changes convert raw genomic data into mathematically transformed representations that maintain statistical and analytical properties necessary for research while fundamentally altering the data structure to prevent direct identification and inference of sensitive information.
Solution Approach 2:
The patent replaces traditional mechanical anonymization approaches (such as simple masking or aggregation) with advanced cryptographic mechanisms. Instead of physically or structurally modifying data through conventional means, the system uses mathematical transformations and cryptographic protocols that preserve data relationships and utility while providing stronger privacy guarantees through proven security mechanisms.
3Device complexity
If traditional anonymization methods are used, then implementation simplicity is maintained, but security against re-identification attacks is insufficient
Solution Approach 1:
The patent introduces cryptographic intermediaries (secure multi-party computation protocols and homomorphic encryption schemes) that mediate between simple data sharing needs and strong security requirements. These intermediary layers provide automated, standardized security mechanisms that are more reliable than manual anonymization while maintaining operational simplicity through established cryptographic frameworks.
Solution Approach 2:
The patent uses cryptographic copying mechanisms where encrypted representations of data are created and shared instead of raw data. These cryptographic copies maintain data relationships and analytical utility while providing built-in security against re-identification attacks, achieving both simplicity and reliability through proven cryptographic techniques.
Data Source
Figure 1
Figure 2
Figure 3
AI summary
Some embodiments are directed to a method for anonymizing a genomic data set. The method comprises receiving (410) the genomic data set and obtaining (420) a phenotypic probability for at least one phenotype informative single nucleotide polymorphism (SNP) of the genomic data set and a proportion of a population which exhibits a corresponding phenotypic trait. A re-identification risk score is computed (430) based on the genomic data set from the obtained phenotypic probability and the obtained proportion of the population which exhibits the phenotypic trait. If the re-identification risk score does not meet a threshold risk criterion, the genomic data set is anonymized by selecting (450) a phenotype informative SNP and masking (460) the selected phenotype informative SNP, and the re-identification risk score is re-computed. If the re- identification risk score meets the threshold risk criterion, the anonymized genomic data set is output (470).