Methods for anonymizing genomic data
Patent Information
- Application Number
- CN202180074039.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2020-10-29
- Filing Date
- 2021-10-22
- Publication Date
- 2026-09-15
- Estimated Expiration
- 2041-10-22
AI Technical Summary
此外,现有技术的方法涉及去除可能以某种至今未知的方式贡献于特定疾病的数据,这意味着有用信息可能会丢失
[0011] These examples help improve the privacy and security associated with genomic data, while also increasing, for example, the amount of information available to researchers. This article provides various examples and embodiments describing how to determine re-identification risk scores and how to anonymize genomic datasets.
Smart Images

Figure CN116438604B_ABST
Abstract
Description
Technical Field
[0001] The currently disclosed subject matter relates to methods and corresponding systems for anonymizing genomic datasets. The currently disclosed subject matter also relates to computer-readable media. Background Technology
[0002] Whole-genome sequencing is becoming increasingly affordable, with services like 23andMe and AncestryDNA offering sequencing of hundreds of thousands of SNPs for around $100. However, as more genomic information becomes available, concerns about privacy and security are also growing. Adversaries are increasingly able to deanonymize genomic databases by combining genotype and phenotype information in various ways. For example, an identification attack is an attack in which an adversary attempts to identify (among multiple genotypes) the genotype corresponding to a given phenotype. Another deanonymization attack is the perfect match attack, in which an adversary attempts to match multiple phenotypes with their corresponding genotypes. Based on whole-genome sequencing data, adversaries can also use statistical models to predict phenotypic traits. Due to current advances in genomics, the risks of using genomic data to identify individuals are rapidly increasing.
[0003] Quasi-identifiers, also known as indirect identifiers, are fields in a dataset that can be combined to identify an individual. Examples include gender, postal code, date of birth, occupation, and income. While many people may share the same gender, date of birth, or postal code, these combinations are likely unique to any one person, especially if that person lives in a sparsely populated rural area. Examples of indirect identifiers include phenotypic characteristics such as hair color and eye color.
[0004] Currently, whole genome sequences can be easily linked to phenotypic traits, such as eye color, hair color, skin color, and blood type, and subsequently, individuals can be identified. However, this problem will worsen as genomic research progresses. Typically, users and researchers will choose between two options: keeping all genomic information intact, risking privacy violations, or removing all potentially identifiable information from the dataset, which limits the data's usability.
[0005] Published U.S. patent application US 2020 / 0035332 A1 describes methods and systems for anonymizing genetic data. The methods and systems described therein identify ancestral identification marker (AIM) regions in genetic data. AIM regions in genetic data include single nucleotide polymorphism (SNP) alleles associated with a patient population belonging to a specific ancestral lineage. AIM regions that do not contain genetic variations associated with a specific disease may be masked or deleted from the genetic data.
[0006] One problem with existing techniques is that they cannot guarantee sufficient anonymity of the obtained genetic data. In some cases, simply masking or removing AIM regions that lack clinically relevant data may still produce genetic datasets that can re-identify individuals. Furthermore, existing methods involve removing data that may contribute to a particular disease in ways that were previously unknown, meaning that useful information may be lost.
[0007] Removing more data from a genetic dataset increases the risk of losing valuable and relevant information, thus reducing the data's usefulness. However, retaining more data in a genetic dataset increases the risk of re-identifying individuals from it. Therefore, it is beneficial to ensure that genetic datasets are sufficiently anonymized while retaining as much information as possible for applications such as research. Quantifying the risk of re-identification and ensuring that individuals can be re-identified from anonymized genomic datasets can improve patient privacy, security, and the amount of information available to researchers from anonymized genomic datasets. Summary of the Invention
[0008] It is advantageous to preserve as much genomic data as possible for researchers to access while protecting the privacy and security of the individuals whose data is used. Systems and computer-implemented methods for anonymizing genomic datasets are described and claimed herein. These systems and computer-implemented methods are designed to address these and other issues.
[0009] Existing methods for preparing genomic data either remove important research information from the genomic dataset, such as by deleting all genomic data related to visible phenotypic features, regardless of whether the genomic data is also related to the disease of interest, thus reducing the amount of knowledge that can be gained from its analysis, or retain too much individual identification information, thus facing the risk of security and privacy breaches.
[0010] The currently disclosed subject matter includes computer-implemented methods for anonymizing genomic datasets, systems for anonymizing genomic datasets, and computer-readable media. The method for anonymizing a genomic dataset may include receiving a genomic dataset. The genomic dataset may contain multiple alleles arranged with multiple single nucleotide polymorphisms (SNPs), said multiple SNPs including one or more phenotypic informational SNPs. Phenotypic informational SNPs may be SNPs associated with a phenotypic trait. The genomic dataset may correspond to the human genome. The method may further include obtaining a phenotypic probability for at least one phenotypic informational SNP. The phenotypic probability may be the probability that a phenotypic trait is expressed as at least one allele corresponding to at least one phenotypic informational SNP. For example, if the phenotypic trait is "blue eyes," then the phenotypic probability may be the probability that occupying an allele of a specific phenotypic informational SNP associated with eye color will result in blue eyes. The method may also include obtaining the proportion of a population exhibiting said phenotypic trait. For example, if the phenotypic trait is "blue eyes," then the proportion of the population will correspond to the proportion of people in the population who have blue eyes. The method also includes calculating a re-identification risk score based on the genomic dataset. A re-identification risk score indicates the risk of re-identifying individuals associated with a genomic dataset from the dataset. The re-identification risk score can be calculated based on the obtained phenotypic probability and the proportion of the population exhibiting said phenotypic trait. The re-identification risk score can then be compared to a threshold risk criterion. If the re-identification risk score does not meet the threshold risk criterion, the method may include anonymizing the genomic dataset by selecting phenotypic informative SNPs and masking the selected phenotypic informative SNPs. If the re-identification risk score meets the threshold risk criterion, the method may include outputting the anonymized genomic dataset.
[0011] These examples help improve the privacy and security associated with genomic data, while also increasing, for example, the amount of information available to researchers. This article provides various examples and embodiments describing how to determine re-identification risk scores and how to anonymize genomic datasets.
[0012] By using threshold risk criteria, acceptable risk levels can be considered, and genomic datasets can be anonymized accordingly. Furthermore, the amount of information retained in genomic datasets can be maximized by avoiding unnecessary removal of clinical relevance.
[0013] The currently disclosed aspects of the subject include corresponding systems for anonymizing genomic datasets.
[0014] The executable code for an embodiment of the method can be stored on a computer program product. Examples of computer program products include storage devices, optical storage devices, integrated circuits, servers, online software, etc. Preferably, the computer program product includes non-transient program code stored on a computer-readable medium for executing embodiments of the method when the program product is executed on a computer.
[0015] In an embodiment, the computer program includes computer program code adapted to perform all or part of the steps of one embodiment of the method when the computer program is run on a computer. Preferably, the computer program is embodied on a computer-readable medium.
[0016] Another aspect of the currently disclosed subject matter provides a method for making computer programs available for download. This aspect is used when a computer program is uploaded to, for example, Apple's App Store, Google's Play Store, or Microsoft's Windows Store, and when said computer program is available for download from such a store. Attached Figure Description
[0017] Further details, aspects, and embodiments will be described by way of example with reference to the accompanying drawings. Elements in the drawings are shown for simplicity and clarity and are not necessarily drawn to scale. In the drawings, elements corresponding to those already described may have the same reference numerals. In the drawings:
[0018] Figure 1 An example of an embodiment of a system for anonymizing genomic datasets is illustrated schematically.
[0019] Figure 2 An example of a genome dataset is illustrated schematically.
[0020] Figure 3 An example of an embodiment of a method for calculating a re-identification risk score is illustrated schematically.
[0021] Figure 4 An example of an embodiment of a method for anonymizing genomic datasets is illustrated schematically.
[0022] Figure 5 An example of a computer-readable medium having a writable portion including a computer program according to an embodiment is illustrated schematically, and
[0023] Figure 6 A representation of a processor system according to an embodiment is shown schematically.
[0024] List of reference numerals
[0025] 100 System
[0026] 110 Processor Subsystem
[0027] 120 External Network
[0028] 130 Input / Output Subsystem
[0029] 140 Memory
[0030] 142 Genome Dataset
[0031] Instruction 144
[0032] 150 Data Interfaces
[0033] 200 Genomes Dataset
[0034] 330 Database
[0035] 1000 computer-readable media
[0036] 1010 writable portion
[0037] 1020 Computer Program
[0038] 1110 (one or more) integrated circuits
[0039] 1120 Processing Unit
[0040] 1122 Memory
[0041] 1124 Application-Specific Integrated Circuit
[0042] 1126 Communication Components
[0043] 1130 Interconnection
[0044] 1140 Processor System
[0045] 1100 equipment Detailed Implementation
[0046] While the subject matter currently disclosed allows for many different forms of implementation, one or more specific embodiments are shown in the accompanying drawings and will be described in detail herein. It should be understood that this disclosure should be considered as illustrative of the principles of the subject matter and is not intended to limit it to the specific embodiments shown and described.
[0047] In the following description, for ease of understanding, the units of the embodiment are described in operation. However, it is clear that the various units are arranged to perform the functions described as being performed by them.
[0048] Furthermore, the subject matter disclosed herein is not limited to the embodiments, as features described herein or those recited in mutually different dependent claims may be combined.
[0049] Figure 1 An example embodiment of a system 100 for anonymizing genomic datasets is illustrated schematically. System 100 may include a processor subsystem 110 and an input / output subsystem 130. In some embodiments, system 100 may also include a memory 140 accessible via a data interface 150. Memory 140 may be local memory or remote memory. In some embodiments, system 100 may be communicatively coupled to an external network or external entity 120.
[0050] In an embodiment, the input / output (IO) subsystem 130 may include interfaces for receiving inputs and / or outputting results. For example, the IO subsystem 130 may be configured to receive a genomic dataset corresponding to a human genome.
[0051] A genomic dataset may contain multiple alleles arranged across multiple single nucleotide polymorphisms (SNPs). An SNP represents a location in the genome where gene variation typically occurs, and each allele is a variant of a given gene, genetic sequence, or SNP. That is, an SNP indicates a genomic location where at least a portion of the population has a different nucleotide. SNP data can therefore be considered mutation data, and this mutation data can be used to identify individuals from a genomic dataset. Most commonly, an SNP corresponds to a pair of alleles, which can be nucleobases (adenine (A), cytosine (C), thymine (T), or guanine (G)). For example, in an autosome, one allele is inherited from the mother, and the other from the father. For each SNP, it is generally known what the wild-type allele is and what the mutant allele is. The wild-type allele is the allele that typically produces the most commonly found phenotype in the population, while the mutant allele is the allele that produces a phenotype different from the wild-type phenotype. The allele occupying an SNP is called a genotype. Each SNP's constituent alleles may have associated genotype frequencies, which may indicate the frequency of the allele's occurrence at the SNP's location in a population (e.g., a region, country, continent, world, dataset, etc.), and associated phenotypic probabilities, which indicate the probability that the allele produces a particular phenotypic trait. Many SNPs may contribute to or correspond to one or more phenotypic traits. Such SNPs may be referred to as phenotypic informative SNPs. Examples of phenotypic traits include external phenotypic traits such as eye color, skin color, hair color, etc., and / or internal phenotypic traits such as blood type, disease predisposition, lactose intolerance, etc. Such phenotypic traits can be considered indirect identifiers because, although such traits alone may not identify an individual, the combination of many such traits reduces the number of potential populations to which a genome corresponding to a genomic dataset may belong, and may ultimately identify a specific individual. In some embodiments, the genomic dataset may include demographic data such as age information, address information, etc. Figure 2 The document provides simulated example fragments of the genome dataset, which will be further elucidated in their respective descriptions.
[0052] In an embodiment, the IO subsystem 130 may be configured to receive indications of a disease to be studied. A user, such as a researcher, may be interested in learning about or researching a specific disease. The disease is known to correspond to the selection of SNPs. For example, a user may be interested in researching prostate cancer. The IO subsystem 130 may receive user input indicating interest in prostate cancer and may provide the user or another subsystem or process of system 100 with a list of SNPs known to be associated with or contributing to prostate cancer. In some embodiments, a user may select or indicate a specific disease through the IO subsystem 130, and may retrieve known relevant SNPs, for example from an external source or from internal memory such as memory 140. In some embodiments, the disease of interest may be pre-determined or indicated in a genomic dataset, or obtained from it. In some cases, SNPs known to cause a specific disease may also affect phenotypic traits, such as eye color or blood type.
[0053] The IO subsystem 130 may also be configured to store the received genomic dataset in memory 140. In some embodiments, the IO subsystem 130 may be configured to receive user input, such as instructions for a specific disease to be studied, or selections of data in the genomic dataset to be prioritized. In some embodiments, the IO subsystem 130 may be configured to receive a target re-identification risk score to indicate the desired level of anonymization. For example, the target re-identification risk score may be used as a threshold risk criterion.
[0054] In some embodiments, the IO subsystem 130 may be configured to access an external network 120. The external network 120 may include a cloud-based network, servers, external databases, external devices, etc. In some embodiments, genomic datasets, threshold risk criteria, and / or information about at least one disease may be stored on the external network 120 and accessed through the IO subsystem 130.
[0055] In some embodiments, the IO subsystem 130 may include input devices configured to receive input from a user, such as a touchscreen, keyboard, mouse, touchpad, etc., or sensor inputs, such as a camera, microphone, proximity sensor, etc. In some embodiments, the IO subsystem 130 may include output devices such as a display, speaker, etc., to provide output to a user. In some embodiments, the IO subsystem 130 may be configured to process input and / or output from / to additional components, subsystems, or external entities. For example, the IO subsystem 130 may be configured to receive input from external devices, such as a cloud-based network, a server, or components of system 100.
[0056] In embodiments, memory 140 may be configured to store one or more genomic datasets, for example, in a database used to store genomic datasets of multiple individuals. Additionally or alternatively, memory 140 may be configured to store instructions or information for methods of anonymizing genomic datasets. Memory 140 may also store threshold risk criteria. Threshold risk criteria may be criteria used to ensure that anonymized genomic datasets are sufficiently anonymized before output or distribution, for example by ensuring that the re-identification risk score of the genomic dataset meets a specified risk level. (See reference...) Figure 3 The calculation of the re-identification risk score will be described in detail below. The threshold risk criterion can be, for example, a percentage indicating the risk of re-identifying a person from a genomic dataset, or an applicable population size indicating the number of people with the stated phenotypic trait. A large number of people sharing the same phenotypic trait corresponds to a low risk of re-identification. In some cases, such as when the starting population is small, using the applicable population size as the re-identification risk criterion may be particularly appropriate.
[0057] Memory 140 may include at least one database, such as genomic database 142 and / or SNP database 144. Genomic database 142 may be configured to store one or more genomic datasets corresponding to one or more individuals. SNP database 144 may be configured to store SNP information associated with one or more diseases, such as a list of SNPs corresponding to a specific disease. Memory 140 may be implemented as electronic memory, such as flash memory, or magnetic memory, such as hard disk, or optical memory, such as DVD. Memory 140 may include multiple discrete memories that together constitute memory 140. Memory 140 may include temporary memory, such as RAM. In the case of temporary memory 140, memory 140 may be associated with a retrieval device to obtain data before use and store the data in memory, for example, by means of an optional network connection (not shown).
[0058] In embodiments, memory 140 may include local memory and / or external (e.g., remote) memory. For example, genomic database 142 may be stored in local memory. SNP database 144 may be stored externally and, in some cases, may be accessible only by system 100. In another example, genomic database 142 may be stored externally, for example, in a central (e.g., government) database. Genomic database 142 and SNP database 144 may be stored in the same storage location or in different storage locations.
[0059] Processor subsystem 110 may include at least one processor and may be referred to as at least one processor circuit. In embodiments, processor subsystem 110 may be configured to determine a re-identification risk score for a genomic dataset and to anonymize the genomic dataset. In some embodiments, processor subsystem 110 may be configured to preprocess the genomic dataset, for example by masking (e.g., deleting) any direct identifiers therein. Direct identifiers may be SNPs whose data independently identify an individual without requiring additional information, such as data related to other SNPs. In some embodiments, processor subsystem 110 may preprocess the genomic dataset by obtaining a list of SNPs of interest to the user, for example by obtaining a list of SNPs related to or contributing to a specific specified disease directly from the user or from a database in memory 140 or via external network 120. In some embodiments, processor subsystem 110 may be configured to mask (e.g., delete) SNPs in the genomic dataset that are not related to a specified disease, for example by masking (e.g., deleting) SNPs not included in the obtained list of relevant SNPs. In some embodiments, the understanding of SNPs contributing to or related to a specific disease is incomplete, and researchers may not want to limit the genomic dataset based on an incomplete list of SNPs.
[0060] Processor subsystem 110 can be configured to calculate a re-identification risk score based on a genomic dataset. The calculation of the re-identification risk score can be based on one or more phenotypic traits. For each specified phenotypic trait, the genotype frequency of one or more phenotypic informative SNPs associated with said phenotypic trait, the phenotypic probability that said phenotypic informative SNPs produce said phenotypic trait, and the proportion of the population possessing said phenotypic trait can be used in the calculation of the re-identification risk score. (Refer to...) Figure 2 A more comprehensive description of these terms will refer to [reference needed]. Figure 3 Describe in detail the calculation of the re-identification risk score.
[0061] Processor subsystem 110 can be configured to compare a calculated re-identification risk score with a threshold risk criterion. The threshold risk criterion, also known as a threshold re-identification risk criterion, can be stored locally, for example in memory 140, received from a user as user input via I / O subsystem 130, or obtained from an external network 120, for example, via I / O subsystem 130. If the calculated re-identification risk score meets the threshold risk criterion, the genomic dataset is sufficiently anonymized and can be output to, for example, an external device or a user, or stored, for example, in local memory. If the calculated re-identification risk score does not meet the threshold risk criterion, the genomic dataset is not yet sufficiently anonymized, and processor subsystem 110 can be configured to anonymize the genomic dataset by masking (e.g., deleting) data corresponding to one or more SNPs.
[0062] System 100 may also include a data interface 150. The data interface 150 may include connectors, such as wired connectors (e.g., Ethernet connectors, optical connectors), or wireless connectors (e.g., antennas, such as Wi-Fi, 4G, or 5G antennas). The data interface 150 can provide access to the memory 140.
[0063] The various subsystems of System 100 can be located within a single device or communicate with each other via a computer network. The computer network can be the Internet, an intranet, a local area network (LAN), a wireless LAN, etc. The computer network can be entirely or partially wired, and / or entirely or partially wireless. For example, the computer network may include an Ethernet connection. For example, the computer network may include wireless connections such as Wi-Fi, ZigBee, etc. Subsystems may include connection interfaces arranged to communicate with other subsystems of System 100 as needed. For example, connection interfaces may include connectors, such as wired connectors (e.g., Ethernet connectors, optical connectors, etc.) or wireless connectors (e.g., antennas, such as Wi-Fi, 4G, or 5G antennas). The computer network may include additional components such as routers, hubs, etc.
[0064] Figure 2 An example of a genome dataset 200 is illustrated schematically. Genome dataset 200 may include multiple parameters for each SNP, such as... Figure 2The parameters described are as follows: SNP 210, indicating the location or other identifier of the SNP; genotype 220, indicating the allele at the SNP; allele frequency 1 230, indicating the frequency of the first allele among the multiple alleles occupying the SNP; allele frequency 2 240, indicating the frequency of the second allele among the multiple alleles occupying the SNP; genotype frequency 250, indicating the frequency of a genotype (e.g., an allele of the SNP) in a population, such as a regional or global population, or in some cases, a population of a dataset; phenotype 260 (also called phenotypic probability 260), indicating the probability of a particular phenotypic characteristic resulting from the genotype of the SNP and disease association 270, indicating whether the SNP has a known association with a particular disease of interest. It should be understood that genomic datasets may not contain all of the listed parameters, and genomic datasets may contain parameters other than those listed, and the listed parameters are illustrative only. For example, a genomic dataset might contain only SNP 210 and genotype 220, and any further information such as genotype frequency, phenotypic probability, and / or disease relevance could be obtained by querying the database or accessing data sources (e.g., via an external network) 120. For example, based on entries for SNP 210-a (SNP_E1) and associated genotypes ( Figure 2 In the table ("AA"), system 100 can be configured to look up associated genotype frequencies, in this example, the frequency of genotype AA at SNP_E1 in a population (64% in this simulated example), phenotypic probabilities—in this example, the probability that genotype AA at SNP_E1 will produce blue eyes is 40%, and whether SNP_E1 is known to have any association with prostate cancer (prostate cancer in...). Figure 2 In the table, these are represented as PCa. For example, genotype frequencies, phenotypic probabilities, and any related parameters can be obtained or accessed from the same data source (e.g., a single database) or from different data sources.
[0065] Figure 2 The simulated genome dataset 200 shown includes samples corresponding to multiple phenotypic informative SNPs. Genome dataset 200 corresponds to the simulated genomes of individuals with blue eyes, brown hair, and light skin. In this example, for simplicity, it is assumed that these three phenotypes form a set of indirect identifiers, based on which a re-identification risk score is calculated. However, the use of these phenotypic features is merely illustrative and not restrictive. More or fewer phenotypic features may be used. This example, which considers these three phenotypic features, will be used throughout this disclosure to illustrate the methods and apparatus described herein.
[0066] In this example, the first phenotypic trait is "blue eyes". SNPs contributing to or influencing eye color can be SNP_E1, SNP_E2, SNP_E3, SNP_E4, and SNP_E5. Based on the genomic dataset of a specific individual, SNP_E1 consists of genotype AA (e.g., a genotype containing two alleles, each adenine nucleotide). The frequency of genotype AA at the location corresponding to SNP_E1 in the population is 64% (based on this simulation example). This genotype AA at this location results in a 40% probability of blue eyes, and SNP_E1 is known to contribute to or be associated with prostate cancer.
[0067] Similarly:
[0068] • SNP_E2 consists of genotype AG, with a genotype frequency of 4.5%, an 80% probability of causing blue eyes, and has no known association or contribution to prostate cancer;
[0069] • SNP_E3 is filled with the GT genotype, which has a genotype frequency of 20%, a probability of producing blue eyes of 95%, and is known to be associated with prostate cancer;
[0070] SNP_E4 is composed of the CC genotype, with a genotype frequency of 81%, a probability of producing blue eyes of 50%, and has no known association with prostate cancer; and
[0071] • SNP_E5 is filled with genotype CT, with a genotype frequency of 17.5%, a probability of producing blue eyes of 70%, and no known association with prostate cancer.
[0072] Continuing the example, SNPs H1, H2, and H3 are SNPs related to hair color, while SNPs S1, S2, S3, and S4 are SNPs related to skin color. It is particularly important to note that SNPs E1 and H1 (denoted by 210-a, respectively) are identical SNPs—that is, they correspond to the same location in the genome and are filled by the same allele. This particular SNP affects both eye and skin color; the genotype AA filling this SNP has a 40% probability of producing blue eyes and a 55% probability of producing light skin.
[0073] It can be based on, for example Figure 2 The parameters shown in the table are used to calculate the re-identification risk score. (Refer to...) Figure 3 The calculation will be described in more detail.
[0074] Once calculated, the re-identified risk score can be compared to a threshold risk criterion, which could be an applicable group or a percentage. The applicable group could correspond to a proportion of a population in a specific region (e.g., a country, the world, etc.) or a group within a dataset, etc. (Refer to...) Figure 4 Further details regarding the threshold risk criterion are provided below. If the re-identification risk score does not meet the threshold risk criterion, data corresponding to one or more phenotypic informative SNPs can be masked, and the re-identification risk score can be recalculated without using data from the one or more masked phenotypic informative SNPs. Masking SNPs can include, for example, deleting data corresponding to the SNP, replacing data corresponding to the SNP with empty data, or any known masking method. The process of masking one or more SNPs and recalculating the re-identification risk score can be repeated until the re-identification risk score meets the threshold risk criterion.
[0075] Figure 3 An example of an embodiment of a method for calculating a re-identification risk score is illustrated schematically. A first phenotypic feature PT_current 310 can be selected. The first or current phenotypic feature PT_current 310 can be selected from a list of phenotypic features to be considered when calculating the re-identification risk score, or from a list of all known phenotypic features. In some embodiments, the list of phenotypic features from which the first phenotypic feature PT_current 310 can be included, such as external phenotypic features, such as eye color, hair color, etc., internal phenotypic features, such as blood type, etc., or a combination of external and internal phenotypic features. In some embodiments, the list of phenotypic features to be considered in the calculation of the re-identification risk score can be received as input from a user or from another subsystem or device. In some embodiments, the list of phenotypic features to be considered can be obtained from a database, or can be determined based on the phenotypic features exhibited by an individual whose genomic dataset is to be anonymized. For example, if the person has blue eyes, brown hair, and light skin, these phenotypic features may be included in the list of phenotypic features to be considered in the calculation of the re-identification risk score.
[0076] The first phenotypic trait PT_current 310 can be associated with one or more phenotypic informative SNPs present in the genomic dataset. That is, it is known that one or more locations on the genome contain mutations that cause or contribute to the expression of the first phenotypic trait. For illustration, these SNPs are shown as SNP_1 320-1, SNP_2 320-2, and SNP-n 320-n, but it should be understood that there may be more or fewer SNPs for a particular phenotypic trait, and the same SNP may contribute to multiple phenotypic traits to the same or different degrees.
[0077] For at least one of the identified SNPs, taking SNP_1 320-1 as an example, the genotype frequency Gfreq_1 340a-1 and phenotypic probability Pprob_1 340b-1 can be obtained, for example, from database DB 330. Database DB 330 can be stored in local storage such as memory 140 or in external storage such as cloud storage or external devices. For example, database DB 330 can be accessed via an external network such as external network 120. In some embodiments, database DB 330 can be a central database accessible to researchers from organizations, collaborators, etc. In some embodiments, the genomic dataset can include one or more of the genotype frequency Gfreq_1 340a-1 and phenotypic probability Pprob_1 340b-1, thereby allowing these values to be obtained without using a separate database. The genotype frequency Gfreq_1 340a-1 can indicate the frequency of SNP_1320-1 occupied by the allele indicated in the genomic dataset of SNP_1 320-1. The phenotypic probability Pprob_1340b-1 indicates the probability that those alleles produce or cause the first phenotypic trait. In some embodiments, genotype frequencies and phenotypic probabilities can be combined to obtain a risk term for each SNP used to calculate a re-identification risk score. The risk term can be an intermediate value. For example, the risk term PT_r_1350-1 for SNP_1320-1 can be determined by combining Gfreq_1340a-1 and Pprob_1340b-1. The risk term PT_r_1350-1 can be or can include the product of the genotype frequency Gfreq_1340a-1 and the phenotypic probability Pprob_1340b-1, or the sum of the logarithms of Gfreq_1340a-1 and Pprob_1340b-1, etc.
[0078] In some embodiments, these values can be obtained for each phenotypic informative SNP—for example, the genotype frequency Gfreq_2 340a-2 and phenotypic probability Pprob_2 340b-2 for SNP_2 320-2, and the genotype frequency Gfreq_n 340a-n and phenotypic probability Pprob_n 340b for SNP_n 320-n. Risk terms associated with each SNP can be calculated from these values, for example, to obtain risk terms PT_r_2 350-2 corresponding to SNP_2 320-2 and risk terms PT_r_n 350-n corresponding to SNP_n 320-n, and so on.
[0079] In some embodiments, the maximum risk term PT_r_max 360 can be determined, indicated by MAX 355. MAX 355 can return the maximum risk term PT_r_max 360 among the risk terms PT_r_1 350-1 to PT_r_n 350-n corresponding to SNPs SNP_1 320-1 to SNP_n 320-n. In some embodiments, MAX 355 can also return the SNP corresponding to the maximum risk term PT_r_max 360. For example, if the first phenotypic feature has three associated SNPs—SNP_1 320-1, SNP_2 320-2, and SNP_n 320-n—then PT_r_max 360 will be the largest of PT_r_1 350-1, PT_r_2 350-2, and PT_r_n 350-n. Although the use of maximum values has been mentioned above, it should be understood that this method is not limited to this. For example, in some embodiments, an average risk term (e.g., the average risk term of SNP 320 associated with a particular phenotypic trait) may be determined instead of a maximum risk term. For example, the choice between using an average risk term and a maximum risk term may be based at least in part on the risk that the de-identification effort is addressing or the type of attacker.
[0080] In some embodiments, a portion of the population PT_pop340c exhibiting the first phenotypic trait can be obtained, for example, from a database such as DB 330. The proportion of the population exhibiting the first phenotypic trait PT_pop 340c can be combined with the maximum risk term PT_r_max 360 to obtain a contribution term PT_cont 370 corresponding to the first phenotypic trait. For example, the contribution term PT_cont 370 can be the quotient of PT_pop 340c and PT_r_max 360, or the difference between the logarithmic terms of PT_pop 340c and PT_r_max 360. Once the contribution term PT_cont 370 is determined, the method may include selecting the next phenotypic trait PT_next 375 and repeating the process by setting PT_next 375 to PT_current 310, as... Figure 3 As shown by the arrows in the flowchart.
[0081] The contribution term PT_cont 370 can also be combined with other contribution terms, such as those corresponding to other phenotypic features, to obtain the total phenotypic feature term PT_tot 380. In the first iteration, for example for the first phenotypic feature, the total phenotypic feature term PT_tot 380 can be set only to the contribution term PT_cont 370 corresponding to the first phenotypic feature. In some embodiments, the total phenotypic feature term PT_tot 380 can be updated when determining the contribution term for each phenotypic feature. For example, the phenotypic feature contribution PT_cont 370 of the first phenotypic feature can be multiplied or logarithmically added, for example, by adding the logarithm of the contribution term. In some embodiments, the contribution term for each phenotypic feature of interest can be determined as described herein and combined after all said contribution terms have been computed, for example by finding the product of said contribution terms or by finding the sum of the logarithms of said contribution terms.
[0082] In some embodiments, the total phenotypic characteristic PT_tot 380 can be combined with the population size Pop 340d, which can be a regional or global population, or a proportion of an already determined population, to obtain the applicable population AP 390. For example, if a user is interested in studying prostate cancer in patients aged 50 to 75, the population might be the number of men aged 50 to 75 in a region of interest (e.g., Europe, the United States, or globally). The population size Pop 340d can be obtained from a database such as database DB 330, as input from the user, from storage, etc. For example, the applicable population AP 390 can be determined by multiplying the population size Pop 340d by the total phenotypic characteristic PT_tot 380, or by adding the population size Pop 340d to the logarithm of the total phenotypic characteristic PT_tot 380. The applicable population AP 390 can indicate the number of people whose genome is consistent with the genome dataset.
[0083] In some embodiments, the applicable population 390 can be used as the ReID Risk 395, for example, where a threshold risk criterion indicates the number of people. In some embodiments, the ReID Risk 395 can be determined based on the applicable population. For example, the ReID Risk 395 can be the risk that a particular individual can be identified from a genomic dataset and can be calculated as the reciprocal of the applicable population AP 390, such as 1 / AP. This may be appropriate if the threshold risk criterion is based on a risk level, for example, determined or prescribed by ethical or privacy requirements.
[0084] The following is a formula for calculating the re-identification risk score according to one embodiment:
[0085]
[0086]
[0087] in,
[0088] i∈II indicates the phenotypic features in the set II of phenotypic features.
[0089] s∈SNP i SNP i indicating phenotypic characteristics
[0090] Gfreq s,i Genotype frequencies of alleles that fill SNPs
[0091] Pprob s,i The alleles that fill the SNP will result in the phenotypic probability s of phenotypic trait i.
[0092] The set of SNPs indicates the SNPs that correspond to the following phenotypic feature i, for which Gfreq... s,i With Pprob s,i The product has a maximum value
[0093] Pop indicates group size
[0094] PT_Pop i Indicates the proportion of the population exhibiting phenotypic traits i
[0095] As mentioned above, the calculation of the applicable group is not limited to using the maximum term. In some embodiments, for example, the largest item It can be replaced with the average term of the SNP corresponding to phenotypic feature i in the set of SNPs, etc.
[0096] Furthermore, while the re-identification risk score is simply represented here as the reciprocal of the applicable group, it should be understood that the calculation is not limited to this. For example, other dimensions besides the applicable group can be used when calculating the risk score. Such dimensions may include one or more of the following: the attacker's capabilities (e.g., their ability to access various identity databases), the probability or chance of an attack (e.g., internal / external) based on existing context and thresholds, and / or weights corresponding to phenotypes and groups (e.g., the proportion of data subjects with re-identification risk above that threshold, the dependency between phenotypes that segment the applicable group, etc.).
[0097] In some embodiments, additional correction factors can be used to account for dependencies between multiple phenotypic traits. These additional correction factors can be based, at least in part, on available statistical data, such as those indicating an association between two traits or conditions. For example, consider the proportion of the population exhibiting both a high BMI (body mass index) and heart disease phenotype. Generally, obese individuals (medically defined as those with a BMI score exceeding a threshold) comprise approximately 20% of the total population. However, the obese population accounts for 40% of all heart disease patients. In this example, it is clear that a relationship exists between the high BMI phenotype and heart disease, and this relationship can serve as a correction factor. For example, the correction factor could be a 40% factor instead of a 20% factor when calculating the proportion of the population used in the term PT_Pop.
[0098] Figure 4 An example of an embodiment of a computer-implemented method for anonymizing genomic datasets is illustrated schematically.
[0099] The method may include, in an operation entitled “Receiving a Genome Dataset,” receiving a 410-genome dataset, for example via an external network 120, from a user or from a data source such as storage 150 or from an external source. In some embodiments, the genome dataset may be obtained after data corresponding to direct identifiers has been removed. That is, in some embodiments, the genome dataset may include data corresponding to indirect identifiers.
[0100] The method may include, in an operation entitled “Obtaining Parameters”, obtaining 420, the population proportion for at least one phenotypic trait, such as PT_pop, which indicates the proportion of the population exhibiting said phenotypic trait, and the genotype frequency, such as Gfreq, for at least one phenotypic SNP corresponding to said phenotypic trait, and the phenotypic probability, such as Pprob, as referenced. Figure 3 As described in [the document]. In some embodiments, obtaining parameter 220 may include obtaining the population size, such as Pop 340d. In some embodiments, obtaining parameter 420 may include obtaining a list of phenotypic features and / or a list of SNPs associated with each of at least one phenotypic feature.
[0101] The method may include, in an operation entitled “Calculating a Re-identification Risk Score,” calculating a re-identification risk score for the 430 genome dataset. Calculating the 430 re-identification risk score can be performed, for example, as referenced... Figure 3 As described.
[0102] The method may include, in an operation entitled “Comparison with a Threshold”, comparing a calculated re-identification risk score with a threshold risk criterion 440. If the re-identification risk score meets the threshold risk criterion, the method may proceed to an operation entitled “Output anonymized dataset”, wherein the genomic dataset is output 470, for example, to a user or another subsystem, function, or device. In some embodiments, outputting the anonymized dataset may include storing the anonymized dataset in, for example, memory 140. In some embodiments, the genomic dataset may be output to a database, such as database DB 330, or the genomic dataset may be output to an external device, such as a central database or central storage device, in the cloud or a remote device, for example, via an external network 120. In some embodiments, the genomic dataset may be encrypted before outputting (e.g., storing or transmitting) the genomic dataset. In some embodiments, the re-identification risk score may be output along with the anonymized dataset. Outputting the re-identification risk score and anonymizing the genomic dataset allows the anonymized dataset to be used, for example, in subsequent research or applications, where it may have different levels of acceptable re-identification risk (e.g., the threshold re-identification risk threshold may vary). By including a re-identification risk score that includes anonymizing genomic datasets, re-identification risk can be avoided or at least reduced if the re-identification risk score has reached a threshold for subsequent research or application.
[0103] In some embodiments, the re-identification risk score can be calculated as a percentage, for example by taking the reciprocal of the applicable population, as shown in Formula 2. In such embodiments, the threshold risk criterion can take the form of a percentage indicating the risk of re-identifying an individual from the genomic dataset. That is, the threshold risk criterion can indicate the probability that an individual can be identified from the genomic dataset. For example, if the genomic dataset provides a 0.05% risk of re-identification, then a 0.05% threshold risk criterion would indicate that acceptable anonymization has been achieved. Therefore, if the calculated re-identification risk score is below the threshold risk criterion, the threshold risk criterion may be met, and if the calculated re-identification risk score is above the threshold re-identification, the threshold risk criterion may not be met.
[0104] In some embodiments, the re-identification risk score can be calculated as an applicable population, for example, as shown in Formula 1. In such embodiments, the threshold risk criterion can be in the form of raw numbers, such as the raw population size. That is, the threshold risk criterion can indicate the number of people in the population whose genomic dataset can be identified. If the calculated re-identification risk score (e.g., the calculated applicable population) exceeds the threshold risk criterion, then the threshold re-identification risk score can be satisfied.
[0105] If comparing the re-identification risk score with the threshold risk criterion 440 indicates that the re-identification risk score does not meet the threshold risk criterion, then phenotypic informative SNPs present in the genomic data can be selected 450 and masked 460. Masking the selected phenotypic informative SNPs may include deleting data in the genomic dataset corresponding to the selected phenotypic informative SNP, replacing data corresponding to the selected phenotypic informative SNP with dummy data or empty data, or otherwise obscuring data corresponding to the selected phenotypic informative SNP.
[0106] Phenotypic informative SNPs can be selected by identifying the phenotypic informative SNP with the smallest contribution when calculating the applicable population. In other words, identify the phenotypic informative SNP that contributes the most when calculating the re-identification risk score, since the inverse of the applicable population can be used to calculate the re-identification risk score. Contribution items can be found in the reference... Figure 3 The determination is made that the contribution item corresponds to the phenotypic feature contribution item PT_cont 370. In some embodiments, the contribution item for each phenotypic feature is stored together with information identifying which phenotypic informative SNP corresponds to the contribution item, for example, in temporary storage. In operation 450, a phenotypic informative SNP can be selected whose corresponding risk item is the maximum risk item PT_r_max 360 that contributes to the minimum contribution item PT_cont 370.
[0107] In some embodiments, one or more phenotypic informative SNPs may have associated priority indicators. In some embodiments, data corresponding to phenotypic informative SNPs with associated priority indicators may be stored in a genomic dataset such that they are not selected and masked in operations 450 and 460, respectively. For example, selecting 450 phenotypic informative SNPs may include determining the smallest contribution associated with the phenotypic informative SNP (e.g., determining the phenotypic informative SNP that contributes the most to the re-identification risk score) without priority indicators. To illustrate this, consider the following example, where for a particular phenotypic trait, Figure 3The risk term PT_r_1 350-1 of SNP_1 320-1 is the highest risk term for the phenotypic feature (e.g., PT_r_1 350-1 = PT_r_max 355), and the resulting contribution term PT_cont 370 is the smallest contribution to the re-identification risk score, which does not meet the threshold risk criterion. If SNP_1 320-1 has an associated priority indication, the data corresponding to SNP_1 320-1 can be left unselected or masked even though it corresponds to the smallest contribution term. Instead, the next smallest contribution term and its associated phenotypic informative SNP can be determined. If the phenotypic informative SNP associated with the next smallest contribution term does not have a priority indication, the phenotypic informative SNP can be selected in operation 450, and its corresponding data can be masked in operation 460.
[0108] Priority indicators can be obtained through user input. In some embodiments, a user may be particularly interested in a subset of phenotypic informative SNPs and may wish to ensure that said subset exists in an anonymized dataset. In this case, the user can enter a list of SNPs of interest individually or as an additional field or flag in the genomic dataset to be anonymized. In some embodiments, priority indicators can be automatically assigned based on the proximity of SNPs known to be associated with a specific disease of interest. The proximity of one SNP to another can be determined by any known method, such as using the genomic pathway network described in patent application EP3479272 A1, which is incorporated herein by reference in its entirety, and in particular on page 4, lines 28–5, line 3 and page 6, lines 18–7, line 9. For example, SNPs within a predefined distance or proximity to the SNP of interest or SNPs known to contribute to a specific disease can be prioritized by assigning priority indicators to such SNPs.
[0109] For example, a user can indicate a specific SNP of interest. The distance between each phenotypic informative SNP in the genomic dataset and the indicated SNP can be determined, for example, using a genomic pathway network. If, for a given SNP, the distance between the indicated SNP and the SNP is below a threshold distance, the SNP can be added to a subset of phenotypic informative SNPs to which priority indication can be applied.
[0110] Once phenotypic informative SNPs are masked in operation 460, the re-identification risk score can be recalculated in operation 430 without using the data corresponding to the SNPs that provided masked phenotypic information. In other words, the phenotypic informative SNPs selected in operation 450 are effectively removed from the genomic dataset. Subsequent calculations of the re-identification risk score may not include the use of data corresponding to such phenotypic informative SNPs.
[0111] Phenotypic informative SNPs can be removed from the genomic dataset, for example, by masking SNPs that provide informative phenotypes, until the resulting re-identification risk score meets a threshold risk criterion. For example, if the threshold risk criterion represents an acceptable risk level, the selection and masking of phenotypic informative SNPs, as well as the recalculation of the re-identification risk score, may be repeated until the genomic dataset is sufficiently anonymized.
[0112] Figure 5 An example of a computer-readable medium having a writable portion including a computer program 1020 is schematically shown. The computer program 1020 includes functions for causing a processor system to perform actions such as... Figure 4 The computer program 1020 includes instructions for the method of providing diagnostic support to a user. The computer program 1020 may be embodied on the computer-readable medium 1000 as a physical marker or by magnetization of the computer-readable medium 1000. However, any other suitable embodiments are contemplated. Furthermore, it should be understood that although the computer-readable medium 1000 is shown herein as an optical disc, the computer-readable medium 1000 may be any suitable computer-readable medium, such as a hard disk, solid-state storage, flash memory, etc., and may be non-recordable or recordable. The computer program 1020 includes instructions for causing a processor system to perform the method of providing diagnostic support to a user.
[0113] Figure 6 A representation of a processor system 1140 according to an embodiment of a system 100 for anonymizing genomic datasets is schematically shown. The processor system includes one or more integrated circuits 1110. The architecture of the one or more integrated circuits 1110 is as follows: Figure 6 The circuit 1110 is schematically illustrated. It includes a processing unit 1120, such as a CPU, for running computer program components to perform methods and / or implement modules or units according to embodiments. The circuit 1110 includes a memory 1122 for storing programming code, data, etc. A portion of the memory 1122 may be read-only. The circuit 1110 may include a communication element 1126, such as an antenna, a connector, or both. The circuit 1110 may include an application-specific integrated circuit 1124 for performing some or all of the processing defined in the method. The processor 1120, memory 1122, application-specific IC 1124, and communication element 1126 may be interconnected via an interconnect 1130, which may be a bus. The processor system 1110 may be arranged for contact and / or contactless communication using antennas and / or connectors, respectively.
[0114] For example, in one embodiment, processor system 1140, such as a system for anonymizing genomic datasets, may include processor circuitry and memory circuitry, with the processor arranged to execute software stored in the memory circuitry. For example, the processor circuitry may be an Intel Core i7 processor, an ARM Cortex-R8, etc. In another embodiment, the processor circuitry may be an ARM Cortex M0. The memory circuitry may be ROM circuitry or non-volatile memory, such as flash memory. Alternatively, the memory circuitry may be volatile memory, such as SRAM memory. In the latter case, the system may include a non-volatile software interface, such as a hard disk drive, a network interface, etc., arranged to provide the software.
[0115] While system 100 for anonymizing genomic datasets is shown as including one of each described component, in various embodiments, there may be multiple components. For example, processor 1120 may include multiple microprocessors configured to independently execute the methods described herein, or configured to execute steps or subroutines of the methods described herein, such that the multiple processors cooperate to achieve the functionality described herein. Furthermore, in the case of implementing system 100 in a cloud computing system, the various hardware components may belong to separate physical systems. For example, processor 1120 may include a first processor in a first server and a second processor in a second server.
[0116] It should be noted that the above embodiments are illustrative and not limiting of the subject matter disclosed herein, and those skilled in the art will be able to devise many alternative embodiments.
[0117] In the claims, any reference numerals placed in parentheses shall not constitute a limitation on the claims. In the claims, the verb "comprising" and its conjunctions do not exclude the presence of elements or steps other than those stated in the claims. The words "a" or "an" preceding an element do not exclude the presence of a plurality of such elements. Expressions such as "at least one" preceding a list of elements indicate a selection from all or any subset of the elements in the list. For example, the expression "at least one of A, B, and C" should be understood to include only A, only B, only C, both A and B, both A and C, both B and C, or all of A, B, and C. The subject matter currently disclosed can be implemented by hardware comprising several different elements and by means of a suitably programmed computer. In device-type claims that enumerate several parts, several of these parts can be implemented by the same item of hardware. Although specific measures are recited in mutually different dependent claims, this does not indicate that combinations of these measures cannot be advantageously used.
[0118] In the claims, the reference numerals enclosed in parentheses refer to reference numerals in the drawings of exemplary embodiments or formulas of embodiments, thus increasing the comprehensibility of the claims. These reference numerals should not be construed as limiting the claims.
Claims
1. A computer-implemented method for anonymizing a genomic dataset, the genomic dataset comprising multiple alleles arranged in a plurality of single nucleotide polymorphisms (SNPs), the plurality of SNPs including one or more phenotypic informational SNPs, the phenotypic informational SNPs being SNPs associated with phenotypic traits, the genomic dataset corresponding to the human genome, the method comprising: Receive the genome dataset; Obtain the phenotypic probability for at least one phenotypic informative SNP and the proportion of the population exhibiting the phenotypic trait, where the phenotypic probability is the probability that the phenotypic trait is expressed as at least one allele corresponding to the at least one phenotypic informative SNP. A re-identification risk score is calculated based on the genomic dataset, the re-identification risk score indicating the risk of re-identifying individuals associated with the genomic dataset based on the genomic dataset, the re-identification risk score being calculated based on the obtained phenotypic probability and the proportion of the obtained population exhibiting the phenotypic trait; The re-identified risk score is compared with the threshold risk criterion; If the re-identification risk score does not meet the threshold risk criterion, then: The genome dataset was anonymized using the following method: Select phenotypic informative SNPs corresponding to the phenotypic features considered in the calculation of the re-identification risk score, and Masking the selected phenotypic informational SNPs; and Recalculate the re-identified risk score; If the re-identification risk score meets the threshold risk criterion, then: Output anonymized genome dataset.
2. The method according to claim 1, wherein, Repeat the following steps until the re-identified risk score meets the threshold risk criterion: The re-identified risk score is compared with the threshold risk criterion; The genome dataset was anonymized, and The re-identified risk score is recalculated.
3. The method according to claim 1 or 2 further includes encrypting the anonymized genomic dataset.
4. The method according to claim 1 or 2, wherein, Calculating the re-identification risk score includes: For each of at least one phenotypic trait: A risk term is calculated for a phenotypic informative SNP associated with a phenotypic trait. The risk term is calculated based on the genotype frequency of the phenotypic informative SNP and the phenotypic probability of the phenotypic trait associated with at least one allele of the phenotypic informative SNP, where the genotype frequency indicates the frequency of at least one allele of the phenotypic informative SNP in the population. Obtain the proportion of the population exhibiting the phenotypic trait; The re-identification risk score is calculated based on the calculated risk item for each of the at least one phenotypic features and the proportion of the population for each of the at least one phenotypic features.
5. The method according to claim 4, wherein, Calculating the re-identification risk score includes: For each of the multiple phenotypic traits: Obtain the proportion of the population exhibiting the phenotypic trait; Identify at least one phenotypic informative SNP associated with the phenotypic feature; Calculate the risk term for each phenotypic SNP in at least one identified phenotypic SNP; For the phenotypic characteristics, select the SNP with the highest risk term; and The contributor to the phenotypic trait is determined based on the proportion of the population exhibiting the phenotypic trait and the risk term of the selected SNP; and The applicable population value is determined based on the population and the contribution term for each of the multiple phenotypic traits; and The re-identification risk score is calculated based on the applicable group value.
6. The method according to claim 4, wherein, Selecting the phenotypic informative SNPs includes selecting SNPs whose risk terms are used to calculate the minimum contribution term.
7. The method according to claim 1 or 2, wherein, The one or more phenotypic informational SNPs comprise a subset of SNPs with priority indications, and wherein selecting the SNPs includes selecting phenotypic informational SNPs without priority indications.
8. The method according to claim 7, wherein, The subset of SNPs with the aforementioned priority indication is identified in the following manner: For each of the one or more phenotypic informative SNPs: Determine the distance between the SNP and the pre-specified SNP of interest; If the determined distance is within the threshold distance, then the SNP is added to the subset of SNPs with the priority indication.
9. The method according to claim 1 or 2, wherein, Masking the selected SNP involves deleting data entries from the genome dataset that represent the selected SNP.
10. The method according to claim 1 or 2, further comprising outputting the re-identification risk score.
11. The method according to claim 1 or 2, wherein, Calculating the re-identification risk score involves obtaining statistical information from a database about the dependencies between multiple phenotypic traits and applying a correction factor derived from the statistical information.
12. The method according to claim 1 or 2, further comprising: Identify at least one direct identifier, which is a SNP that independently identifies the person; and Mask at least one directly identified identifier in the genome dataset.
13. The method according to claim 1 or 2, wherein, The phenotypic features include external phenotypic features.
14. A computer-readable medium comprising transient or non-transient data representing instructions, which, when executed by a processor system, cause the processor system to perform a computer-implemented method according to any one of claims 1 to 13.
15. A system for anonymizing a genomic dataset comprising multiple alleles arranged in a plurality of single nucleotide polymorphisms (SNPs), the plurality of SNPs including one or more phenotypic informational SNPs, the phenotypic informational SNPs being SNPs associated with phenotypic traits, the genomic dataset corresponding to the human genome, the system comprising: The input / output subsystem is configured as follows: Receive the genome dataset; Obtain the phenotypic probability for at least one phenotypic informative SNP and the proportion of the population exhibiting the phenotypic trait, where the phenotypic probability is the probability that the phenotypic trait is expressed as at least one allele corresponding to the at least one phenotypic informative SNP. The processor subsystem is configured as follows: A re-identification risk score is calculated based on the genomic dataset, the re-identification risk score indicating the risk of re-identifying individuals associated with the genomic dataset based on the genomic dataset, the re-identification risk score being calculated based on the obtained phenotypic probability and the proportion of the obtained population exhibiting the phenotypic trait; The re-identified risk score is compared with the threshold risk criterion; If the re-identification risk score does not meet the threshold risk criterion, then: The genome dataset was anonymized using the following method: Select phenotypic informative SNPs corresponding to the phenotypic features considered in the calculation of the re-identification risk score, and Masking the selected phenotypic informational SNPs; and Recalculate the re-identified risk score; If the re-identification risk score meets the threshold risk criterion, then: Anonymous genomic datasets are output via the input / output subsystem.
Citation Information
Patent Citations
Disease-oriented genomic anonymization
EP3479272A1
Method and apparatus for masking clinically irrelevant ancestry information in genetic data
US20200035332A1
Disease-oriented genomic anonymization
CN109416932A
Disease-oriented genomic anonymization
US20190333607A1