Incremental anonymisation of dynamic datasets
The method addresses the challenge of integrating new data into anonymized datasets by determining compatibility with anonymization schemes, enabling efficient and privacy-focused data sharing through techniques like binning and encryption, thus enhancing dataset usability.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2025-09-02
- Publication Date
- 2026-03-12
AI Technical Summary
Existing methods fail to efficiently manage the integration of new data into anonymized datasets, particularly in dynamic environments, leading to suboptimal usability and privacy concerns.
A computer-implemented method for determining compatibility of a batch of data with an anonymization scheme, allowing incremental addition to a shared anonymized dataset by creating a share enable message to anonymize and publish a subset of the updated dataset, using techniques like binning, encryption, and pseudonymization.
Enables efficient and privacy-conscious handling of dynamic datasets by ensuring compliance with anonymization schemes, facilitating the sharing of new data while maintaining privacy and minimizing access to original data.
Smart Images

Figure EP2025074985_12032026_PF_FP_ABST
Abstract
Description
[0001] METHOD AND COMPUTER NETWORK
[0002] TECHNICAL FIELD
[0003] The present disclosure relates to a computer-implemented method and a computer network.
[0004] BACKGROUND
[0005] The use of personal or sensitive data raises concerns about privacy and data protection. Such concerns constitute a significant obstacle for innovation based on such data, especially in the medical field where data based innovation is crucial for helping patients and for relieving the healthcare systems. To address these concerns, the concept of anonymized data has emerged as a way to ensure the privacy of individuals whilst still enable data analysis and processing.
[0006] Anonymized data refers to information which has been modified or stripped of identifying details, such as names, addresses, or other personally identifiable information (PI I). This process helps to protect the privacy of individuals by making it difficult or impossible to link the data back to specific individuals.
[0007] The lifecycle planning of anonymized data involves the careful consideration of the entire data journey, from collection to disposal. It includes the development of strategies and protocols to ensure that the data remains anonymized throughout its lifespan and is handled in compliance with relevant privacy regulations.
[0008] Batch processing is a method for analyzing and manipulating large volumes of data. Batch processing offers advantages such as scalability, efficiency, and the ability to perform complex computations on large datasets.
[0009] The combination of anonymized data lifecycle planning and batch processing provides a framework to handle and analyze data, whilst maintaining privacy and minimizing access to original data.
[0010] Overall, the integration of anonymized data lifecycle planning and batch processing serves as a foundation for responsible and privacy-conscious data practices.
[0011] However, the inventors have noted the absence of an efficient method to maintain maximum usability of anonymized data poses significant challenges when handling dynamic datasets, e.g., datasets increasing over time with new data batches. Existing solutions do not adequately address the constant influx of new information, leading to suboptimal usability.
[0012] The present disclosure was arrived at in light of the above considerations. SUMMARY
[0013] In a first aspect, embodiments of the invention provide a computer-implemented method for determining if adding a batch of data to an extant dataset, thereby creating an updated extant dataset, would result in additional data being shareable under an anonymization scheme, whereby the data in the batch of data and the extant dataset contain or represent personally identifiable information linked to one or more values, the method comprising: receiving the batch of data or a representation thereof; determining whether at least a subset of the updated extant dataset including at least a portion of the received batch of data would be compatible with the anonymization scheme; and responsive to determining that it would be, transmitting to a recipient a share enable message so as to cause the creation of an anonymization of the at least a subset of the updated extant dataset in accordance with the anonymization scheme and the publication of this anonymized data.
[0014] Such a method may allow incremental addition to a shared anonymized dataset (i.e. the at least a subset), and may facilitate the sharing readiness assessment of new anonymization scenarios (e.g., with higher granularity).
[0015] The values in the batch of data may be clinical or medical values, for example the height and age of a person and / or a medical (or biological) test result. The personally identifiable information may be a name, patient ID, age indication (e.g., a birthdate, a birth year, an age at the time a test was taken) or similar. The batch of data may be received from a medical data source. The values may be characters, e.g., ‘M’ or Male or ‘F’ for female, “70 mg / dL” to indicate the result of a medical test, e.g., for determining a blood glucose concentration. Similarly, the values may be imaging data or other clinical data.
[0016] The batch of data may be a message, the message containing or representing personally identifiable information linked to one or more values.
[0017] The share enable message may cause the creation of the anonymization of the at least a subset of the updated extant dataset by being configured to prompt a data sharing platform to create the anonymization of the at least a subset of the updated extant dataset in accordance with the anonymization scheme and the publication of this anonymized data. Creating the anonymization may comprise utilizing previously anonymized data of the extant dataset, e.g., by combining the previously anonymized data with newly anonymized data (e.g., of the batch of data).
[0018] The subset of the updated extant dataset may be a true subset of the updated extant dataset (i.e., not the entire updated extant dataset) or may be the full updated extant dataset.
[0019] The method may further include, upon receipt of the share enable message, one of:
[0020] (a) anonymizing the batch of data or portion thereof and integrating the anonymized version to an anonymized version of the extant dataset which has been previously anonymized according to the anonymization scheme;
[0021] (b) anonymizing the batch of data or portion thereof together with data, e.g., previously withheld data, in the extant dataset according to the anonymization scheme; and / or
[0022] (c) anonymizing the updated extant dataset or a shareable portion thereof according to an anonymization scheme different to an anonymization scheme according to which the extant dataset has previously been anonymized.
[0023] The method may further include publishing the anonymized updated extant dataset or shareable portion thereof.
[0024] The extant dataset may be an initial dataset, containing non-anonymised or only semi-anonymised records (e.g., where some obfuscation and / or generalization of personally identifiable information has already taken place), or may be a live dataset which corresponds to the initial dataset having been augmented by prior batches of data. The updated extant dataset is therefore the initial data plus any previously received batches of data. The updated extant dataset then takes the role of the extant dataset for the next iteration of the method. The published dataset corresponds to the anonymized version of the shareable portion (i.e. the at least a subset) of the updated extant dataset. The shareable portion of a dataset, e.g., the updated extant dataset, is the portion of the dataset which is currently compatible with the anonymization scheme. The shareable portion may comprise all records of the dataset (where all data can be shared under the anonymization scheme) or can be a true subset of the dataset (i.e. at least some records are currently being withheld). Publishing may mean making available to at least a third party, e.g., a third party that does not have access the nonanonymized updated extant dataset. Publishing may not necessarily mean making it available to the public, but may e.g., mean making available to authorised subscribers.
[0025] The anonymization of the at least a subset of the updated extant dataset in accordance with the anonymization scheme may include only anonymizing the subset of the updated extant dataset that had not be anonymized previously (e.g., after a previous assessment using the method) and adding the newly anonymized data entries to the previously anonymized data entries.
[0026] In some cases, only a portion of the received batch may be anonymized, e.g., where the extant dataset (for this iteration) does not comprise previously withheld data that now became shareable (for example, where all data of the extant data set already had been shared).
[0027] In some cases, the addition of the batch of data to the extant data set leads to previously withheld data in the extant dataset (for this iteration) newly becoming shareable. In an example, the anonymization scheme comprises binning the ages into decades and a re-identification risk threshold for a bin is satisfied if a bin comprises 20 or more subjects. The extant database prior to adding the batch of data comprises 8 subjects in the bin “20-29”, the batch of data comprises 13 subjects in this bin so that the updated extant dataset now comprises 21 subjects in that bin, so that the bin now becomes shareable, and now all of these 21 subjects’ record will be anonymized. Of course, it is also possible in that case that the previously withheld 8 subject’s records were already anonymized before receiving the new bath of data and the anonymized new 13 subject’s records are simply added to the previously anonymized 8 subjects’ records. Examples of types of bins can include bins defined e.g., by age, weight, height, location, and / or combinations thereof.
[0028] Option (c) may be performed when it is determined that the updated extant dataset is compatible with the further anonymization scheme, i.e. an anonymization scheme different to an anonymization scheme according to which the extant dataset has previously been anonymized. The further anonymization scheme may, for example, provide a more granular anonymized representation of the data than the previously used anonymization scheme.
[0029] For example, the extant dataset comprised at least a subset of the extant dataset that was compatible with a first anonymization scheme and this at least a subset was shared under that first anonymization scheme (i.e. using the rules defined thereby) as a first published dataset and the adding of the batch of data to the extant dataset may lead to a compatibility of at least a subset of the updated extant dataset with a second anonymization scheme (e.g., a more granular anonymization scheme). In such cases, it may be the case that the all the records of the at least a subset of the updated extant are newly anonymized according to the second anonymization scheme and shared as a second published dataset. The first published dataset may be retained (i.e., the publication of the second published dataset does not supplant the first).
[0030] Anonymizing data, e.g., of the at least a subset of the updated extant dataset, may include data obfuscation and / or generalization of personally identifiable information in the extant dataset. The anonymization may for example include binning the data according to personally identifiable information. For example, the data records may be binned by grouping personally identifiable information values by ranges of age (e.g., ages 0-9, 10-19, 20-29). In this example, the anonymization scheme may be the scheme by which data, e.g., of the at least a subset of the updated extant dataset, is binned into categories. An anonymization scheme may comprise one or more re-identification risk thresholds, e.g., an identification threshold based on the number of subjects in a bin. Anonymization comprising binning the data may include removing a specific value and replacing it with an indication of the bin into which the data is to be placed (e.g., the age “25” is replaced by the bin label “20-29”).
[0031] The anonymization may include redaction of at least some of the data. Redaction of the data may include deleting or writing over at least some of the data, e.g., with dummy characters. This may be useful if the respective data is not needed for the research, e.g., where all initials of individuals are replaced by “N. N.”.
[0032] The anonymization may include encryption. Encryption may include encrypting at least some of the data e.g., using an encryption key. In particular, in this example, the anonymization may include format-preserving wherein the encrypted part of the data has a same format as the non-encrypted version of that part.
[0033] The anonymization may include scrambling. Scrambling may include permanently modifying at least some of the data e.g., by rearranging or changing one or more values of the data. For example, “Joe Doe” may be modified to “Ejo Ode” or to “Jxx Dxx”.
[0034] The anonymization may include pseudonymization. Pseudonymization may include replacing at least some of the data with a pseudonym, e.g., “Abraham Lincoln” with “Joe Doe”.
[0035] The anonymization may include statistical data replacement. Statistical data replacement may include replacing at least some of the data based on statistical learning of the data and then using these statistics to replace the existing data in a realistic manner, thereby obfuscating the data (e.g., in an undetectable way). Statistical data replacement may allow for maintaining statistical information of the modified data, which may allow improved maintaining of its suitability for later research.
[0036] The anonymization may include machine learning based methods, such as replacing data using a trained machine learning algorithm. The machine learning algorithm may for example be trained using a representative sample of a population.
[0037] The anonymization may include applying a transformation model to the data. For example, the anonymization may include a global transformation scheme, wherein a same transformation is applied to each subset of the data. In this example, each value of an attribute's domain may be transformed to a same generalization level resulting in full-domain generalization. A global transformation process can, therefore, usefully provide attribute suppression in the anonymized data. In this context, attributes may refer to indirect identifiers (or quasi-identifiers, or keys) that do not directly identify an individual in the data but may together with other indirect identifiers form an identifier that can be used for linkage attacks.
[0038] The anonymization may include applying a local transformation scheme, wherein a number (e.g., a plurality) of different transformations are applied to different proper subsets of the data. A predetermined maximum number of transformations to be applied to the data may be specified. In contrast to the global transformation scheme, in this example, different generalization levels may be used for a same attribute value in different records of the data. The local transformation scheme can, therefore, usefully provide cell suppression (e.g., of individual datums) in the anonymized data.
[0039] Generalization hierarchies may be used to directly reduce the uniqueness of attribute values of the data or to form clusters of the data that may be transformed using further methods, such as microaggregation as discussed below. Random sampling of at least some of the data may be performed so as to reduce privacy risks.
[0040] The anonymization may include removing individual attributes, attribute values and / or, one or more complete records from the data. This may be controlled by defining appropriate hierarchies, by performing local or global transformations as discussed above, and / or by specifying a limit for a maximal number of data records which may be removed from the data.
[0041] The anonymization may include aggregation. In these examples, sets of numeric attribute values in the data may be transformed into a common value by user-specified aggregation functions. Prior to the aggregation, clustering may be performed based on value generalization hierarchies. An example of an aggregation method for anonymizing data is data binning.
[0042] Transformation rules that are represented as functions may be defined, which can be used to perform on-the-fly categorization (e.g., binning) of continuous variables in the data during the anonymization.
[0043] The anonymization of the data may include combining multiple anonymization techniques such as one or more of the above described anonymization methods.
[0044] The anonymization scheme may comprise a diversity criterion, such as ’-diversity. The diversity criterion may for example include: distinct t -diversity: at least t distinct values Pll values exist in a class; entropy t -diversity: the entropy of a class is larger than an t -dependent value, e.g. Iog( / ); recursive (c- ^-diversity: the most common value does not appear too often, while less common values are ensured to not appear too infrequently.
[0045] Such diversity criterion may for example be utilized for ensuring that the bins are sufficiently diverse for achieving an acceptable re-identification risk.
[0046] A representation of a data set (e.g., a batch of data, the extant dataset, and / or the updated extant dataset) may comprise meta information, e.g., meta information indicating a plurality of categories into which personally identifiable information can be sorted or generalized. The representation of the personally identifiable information in the batch of data may be meta information indicating the personally identifiable information bin or bins in which the relevant information would sit. The representation may comprise encrypted data, e.g., tokenized meta information. The representation of a dataset (e.g., a batch of data, the extant dataset, and / or the updated extant dataset) may comprise encrypted meta information indicating a plurality of categories into which personally identifiable information can be sorted. In some examples, the meta information may relate to a predefined maximum number of transformations to be applied to the data, user-specified generalization hierarchies, and / or information on types of attribute values to suppress in the data. Compatibility with the anonymization scheme may be based on, e.g., determined by, comparing a calculated re-identification risk to a re-identification risk threshold. Re-identification may be computed, in some examples, as an inverse to the number of individuals falling into the same bin (or combination of bins). For example, a re-identification risk of 20 individuals in a dataset with age 10-19 and living in Basel equals 1 / 20 = 5% and, in this case, sharing the data relating to people in the age-category of 10-19 and replacing the actual age with this age bin may cause at most a 5% risk of de-identification of an individual, which may be deemed acceptable in a given scenario.
[0047] The meta information may comprise a diversity index, such as an index indicating a measure related to an t -diversity. The re-identification risk threshold may comprise a diversity related threshold, such as ^-diversity index being above a defined value. Determining the compatibility of a subset with an anonymization scheme may utilize a diversity index of the subset and / or data categories (e.g. bins) therein, e.g., a comparison of a diversity index to a defined value.
[0048] The method may further comprise obtaining information of the anonymization scheme, e.g., bins according to which the anonymization of the dataset, e.g., the updated extant, shall be done and / or the respective re-identification risk threshold.
[0049] A dataset, e.g., a subset of the updated extant dataset, may be compatible (e.g., compliant) with an anonymization scheme if the re-identification risk of that subset after being anonymized according to rules of the anonymization scheme is below a threshold defined by the anonymization scheme.
[0050] The method may further comprise: obtaining meta information relating the extant dataset’s compatibility with the anonymization scheme; creating, (e.g., determining and / or calculating) meta information relating to the updated extant dataset’s compatibility with the anonymization scheme; wherein the step of determining whether the at least a subset of the updated extant dataset would be compatible with the anonymization scheme is based on the meta information relating to the updated extant dataset’s compatibility. “Based on” here may mean that the decision is determined by the meta information, e.g., by a computation using the meta information (or part thereof) as input. In an example, the determining depends on the number of subjects in a bin and the meta information comprises information that allows determining this number, e.g., an encrypted representation of each bin and the number of subjects for each bin.
[0051] Storing of meta information in a watchdog module and decision making based on this meta information may allow for the secure retention of an initial version of the extant dataset (i.e., an nonanonymized or only semi-anonymized version) in the control of a first party, e.g., the data owner, and a decision making (on what data to be shared) by a watchdog module controlled by a second party. In some cases, the watchdog module may receive plain personally identifiable information comprised in a batch of data, but store only meta information (e.g., counts for respective bins) long term. In other cases, the watchdog may operate purely by receiving and processing only meta information, thus limiting the risks caused by an attack on the watchdog module.
[0052] In some examples, the watchdog module receives only a representation of the batch of data. The representation may contain neither the plain personally identifiable information nor the one or more values, but merely a representation of them intelligible to a watchdog module and usable to determine compatibility with an anonymization scheme. For example, receiving the information that a batch of test result data comprises two entries in the age range 10-19 and seventeen entries in the age range of 20-29 may be sufficient for the watchdog module to decide which of these group of entries may be shared by a data sharing platform.
[0053] A representation of the batch of data may comprise meta information or information that allow deducting, e.g., computing, the meta information. This may allow that a watchdog module can operate by receiving and processing such a representation of the batch of data.
[0054] The meta information may be indicative of the size of the bins of a dataset (e.g., the updated extant dataset), e.g., where the compliance with the anonymization scheme depends on the size of the bins.
[0055] The meta information for an anonymization scheme may be related to, in particular represent, the personally identifiable information being anonymized according to that anonymization scheme. For example, where the anonymization scheme foresees the binning of age indicators into decades, the meta information may comprise a, possibly encrypted, naming of the respective decades and a count of how many subjects to which the values of the data sets are links fall into the respective decade. The method may be triggered by the receipt of the batch of data, and may be repeated for a plurality of subsequently received batches of data. In this case, the updated extant database then becomes the extant database for the next iteration.
[0056] The meta information may be encoded meta information, for example through application of a oneway function such as a hash (for example MD5, SHA-1 , or SHA-2). The meta information may be tokenized. Encrypting at least some of the meta information, e.g., the age range of a bin, may improve the security of such personally identifiable information comprised in, represented by, and / or deductible from the meta information. For example, the watchdog module may long-term store only the tokenized version of each bin name of an anonymization scheme and the count of subjects in the extant data set for each bin so that the stores information may only be of limited use for an attacker.
[0057] The plurality of subsequently received batches may be received from a plurality of different data sources. According to some examples, the initial extant dataset is distributed between multiple data controllers, and, only upon receiving the share enable messages, the respective portions are anonymized, these anonymized data portions are shared and amalgamated, and this amalgamated set of anonymized data is published. This can allow the aggregation of datasets which, whilst individually would not be shareable, in aggregation would be shareable by the data sharing platform. This can allow for providing anonymized datasets that otherwise would not be practically possible and thereby enable research that otherwise would not be possible.
[0058] The method may include an initialisation process, the initialisation process including steps of: receiving an initial dataset, the initial dataset comprising personally identifiable information linked to one or more values; anonymizing at least a subset of the initial dataset according to the anonymization scheme (e.g., after a determination that the at least a subset of initial dataset including would be compatible with the anonymization scheme); and publishing this anonymized data
[0059] In analogy to the later iteration steps, the initialisation process may also include creating meta information and / or receiving a representation of the initial dataset that comprises and / or allows creating meta information. The meta information may be information adapted to and / or specific to the anonymization scheme(s), e.g., describing the initial dataset to a degree that allows determining its compatibility with the anonymization scheme(s).
[0060] The initial dataset may be the amalgamation of data received from a plurality of different data sources. For example, the method may include receiving portions of an initial dataset from each of a plurality of different data sources, and combining the received portions to arrive at the initial dataset. Similarly, later batches of data (or representations thereof) may be received from different data sources.
[0061] Determining whether the updated extant dataset is compatible with the anonymization scheme can comprise determining whether previously withheld data (that is, data in the extant dataset which could not be shared during an earlier iteration including the initialisation process) is now shareable according to the anonymization scheme due to the addition of the batch of data.
[0062] The determination step may include determining, for each of a plurality of different anonymization schemes, which of the plurality of anonymization schemes are compatible with at least a subset of the updated extant dataset. This may include determining a largest subset of the updated extant dataset (possibly the full updated extant dataset) that is compatible with a given anonymization scheme. The share enable message may comprise: an indication of which of the plurality of anonymization schemes are found to be compatible with at least a subset of the updated extant dataset, and an indication of the at least a subset of the extant dataset found to be compatible with the respective anonymization scheme. The plurality of different anonymization schemes may represent different levels of granularity, e.g., a first anonymization scheme may use decade-age bins (0-9, 10-19, etc.) and a second anonymization may use half-decade-age bins (0-4, 5-9, 10-14, 14-19, etc.). Where the bins are indicative of location, the bins may e.g., be more or less specific (e.g., country, county, city, neighbourhood, etc.).
[0063] Where there is a plurality of bin based anonymization schemes, the initialisation process may include create a representation (e.g., a token) for each bin of each anonymization scheme. This representation may be stored by the watchdog module and the respective counts may be adjusted with each batch of data. This may allow the watchdog module to keep track of the bin sizes of the extant data set by only storing the counts for each representation of a bin. The representation may be uniquely determined, e.g., a hash of the bin’s name, which can allow for an exact reproduction of the representation (e.g., token) during each iteration.
[0064] The determination step may be performed by a watchdog module on a computer (sub-)network separate from a computer (sub-)network of the recipient of the share enable message. The two computer (sub-)networks may be controlled by different entities. The two computers (sub-)network may be subject to different administrations. The watchdog module may be separated from the recipient’s computer (sub-)network by a firewall. In some examples, no (non-anonymized) personally identifiable information may be received or stored by the watchdog module, further, in some examples, no values (e.g., medical test results) linked to the personally identifiable information may be received or stored by the watchdog module.
[0065] The recipient of the share enable message may be a data sharing platform. The recipient may be an entity being in control of the extant dataset, e.g., a data owner.
[0066] In a second aspect, embodiments of the invention provide a computer network, comprising: a data sharing platform, configured to: store an extant dataset, the extant dataset containing or representing personally identifiable information linked to one or more values; obtain a batch of data, the batch of data containing or representing personally identifiable information linked to one or more values; and transmit, to a watchdog module, the batch of data or a representation thereof; and the watchdog module being configured to: receive the batch of data or the representation thereof; determine whether at least a subset of an updated extant dataset, created by adding the batch of data to the extant dataset, is compatible with an anonymization scheme; and responsive to determining that it would be, transmitting to the data sharing platform a share enable message so as to cause the creation of an anonymization of the at least a subset of the updated extant dataset in accordance with the anonymization scheme and the publication of this anonymized data. The computer network may comprise at least two separate computer (sub-) networks, a first computer (sub-)network on which the watchdog module operates and a second computer (sub-)network on which the data sharing platform module operates.
[0067] In a third aspect, embodiments of the present invention provide a computer, comprising a processor and memory, the memory containing machine executable instructions which, when executed on the processor, cause the processor to perform the method of the first or fourth aspects and optionally including any one, or any combination insofar as they are compatible, of the optional features set out with reference thereto.
[0068] In a fourth aspect, embodiments of the present invention provide a computer-implemented method for determining if adding a batch of data to an extant dataset, thereby creating an updated extant dataset, would result in additional data being shareable under an anonymization scheme, whereby the data in the batch of data and the extant dataset contain or represent personally identifiable information linked to one or more values, the method comprising: receiving the batch of data or the representation thereof; obtaining meta information relating the extant dataset’s compatibility with one or more anonymization schemes; creating meta information relating to the updated extant dataset’s compatibility with each of the anonymization schemes; determining whether at least a subset of the updated extant dataset including at least a portion of the received batch of data would be compatible with at least one of the anonymization schemes based on the meta information relating to the updated extant dataset’s compatibility; and responsive to determining that it would be, transmitting to a recipient a share enable message so as to cause the creation of an anonymization of the at least a subset of the updated extant dataset in accordance with the at least one of the anonymization schemes and the publication of this anonymized data.
[0069] The invention includes the combination of the aspects and preferred features described except where such a combination is clearly impermissible or expressly avoided.
[0070] Further aspects of the present invention provide: a computer program comprising code which, when run on a computer, causes the computer to perform the method of the first aspect; a computer readable medium storing a computer program comprising code which, when run on a computer, causes the computer to perform the method of the first aspect; and a computer system programmed to perform the method of the first aspect.
[0071] BRIEF DESCRIPTION OF THE DRAWINGS
[0072] Fig. 1 A - 1 F are flow diagrams of example methods;
[0073] Fig. 2A and 2B are schematics illustrating levels of data in datasets;
[0074] Fig. 3A and 3B show dependencies for example methods;
[0075] Fig. 4 is a plot of scenario delay in time against record rate (1 / time);
[0076] Fig. 5 is a plot of batch number against ratio of records above a re-identification risk threshold; and
[0077] Figs. 6 - 8 are diagrams showing computer networks. DETAILED DESCRIPTION
[0078] Aspects and embodiments of the present invention will now be discussed with reference to the accompanying figures. Further aspects and embodiments will be apparent to those skilled in the art.
[0079] Fig. 1 A shows an example of the disclosed method as a flow diagram. The method 100 begins with Step S102, in which a batch of data is received. The batch of data includes at least one, but more commonly a plurality of, data records. Each record includes personally identifiable information (e.g., name, date of birth, height, patient ID etc.) and one or more values (e.g., medical history, medical test result, medical image, medical diagnosis etc.). The values may include a medical test result, such as blood glucose value (e.g., “83 mg / dL”, or simply “83” when the unit is standardised or the units are provided separately), a Troponin level (e.g., “0.04 ng / mL”, or simply “0.04” where the units are standardised or the units are provided separately), or a positive COVID-19 test result (e.g., “yes” or “no”). The batch of data may be received from a medical data source, e.g., a hospital information system, laboratory information system, healthcare information system, and / or a medical data platform. In a step, not shown but which may be carried out subsequent to S102, the batch of data may be appended to a live dataset corresponding to a non-anonymized or only semi-anonymized version of the extant dataset. This live dataset corresponds to an initial dataset which has had previously received batches of data appended to it. In this sense, the batch of data is and / or represents an update to this live dataset. The aim of the method is to decide which additional data of the updated live dataset is shareable under a given anonymization scheme. The additional data may be at least a portion of the batch of data and / or at least a portion of previously withheld data of the extant dataset. The updating of the live dataset and the anonymization may be performed by a data sharing platform which is separate to the entity performing Step S104 below (i.e. , the watchdog module) In some examples, the batch of data includes a representation of the personally identifiable information as discussed above (e.g., a one-way function applied to the personally identifiable information or bin in which it would sit, and / or the value in question).
[0080] After having received the batch of data, or the representation thereof, it is determined in Step S104 whether at least a subset of the live dataset (i.e. the updated extant dataset), the subset including at least a portion of the batch of data, would be compatible with the anonymization scheme, i.e. would be shareable if anonymized in accordance with the anonymization scheme. The term subset here may refer to a true subset of the updated extant dataset (i.e. not all data of the updated extant data set) or the full updated extant dataset. The determination in step S104 may include determining whether the risk of re-identification of any person associated with the values in the respective record or the updated live dataset exceeds a, e. predefined, re-identification risk threshold. For example, if the data record indicated the patient to have a height and weight which placed them in category by themselves, then the risk of re-identification would be high.
[0081] Where it is determined that the at least a subset would be compatible with the anonymization scheme, a share enable message is transmitting to a recipient in step S106. The share enable message comprises an indication of the shareable at least a subset. The share enable message may also include an indication of the anonymization scheme, especially where the anonymization scheme is not canonical, e.g., where the method checks for compatibility with a plurality of anonymization scheme (e.g., iteratively). The share enable message may cause the creation of an anonymization of the at least a subset of the updated extant dataset in accordance with the anonymization scheme and the publication of this anonymized data as will be explained below.
[0082] Figure 1 A also includes two optional steps (shown in dashed boxes), which may be performed prior to Step S104. Step S110 includes obtaining meta information relating to the extant (i.e., live) dataset. In Step S112, meta information relating to the received batch of data is created. Whilst shown in series, Steps S1 10 and S1 12 can actually be performed in parallel or in either order. Step S1 12 may be performed before step S102. The determining whether the at least a subset of the updated extant dataset would be compatible with the anonymization scheme may be based on the meta information relating to the updated extant dataset, e.g., relating to the updated extant dataset’s compatibility with the anonymization scheme. The meta information relating to the updated extant dataset may be created from the meta information relating to the extant (i.e., live) dataset the meta information relating to the received batch of data, e.g., by amalgamation of the two meta information.
[0083] According to some examples, the watchdog module deletes information on the batch of data (e.g., the batch of data, any personally identifiable information therein, and / or any values therein) other than this meta information after having performed an increment of the method. The watchdog module’s long-term memory might solely be based on such meta-information, whereby the risks of sharing data with the watchdog module is limited. The meta-information may be encrypted, e.g., by an one-way- function, which may further increase security.
[0084] Instead of receiving the batch of data as such, the watchdog module may receive a representation of the batch of data that allows creating meta information which enables the determination of compatibility with the anonymization scheme. For example, the representation may comprise the meta information. The meta information may be created using an uniform algorithm, such that the meta information on a dataset, e.g., the batch of data, may be identical independent on when or by which entity the meta information is created. According to an example, the anonymization scheme is based on bins of data representing categories of personally identifiable information and the meta information comprises a hashed name of each of the bins as a token and its respective count. Knowing the possible bins of the anonymization scheme and the hash-function will allow calculating for each record to which token it is associated and increasing the respective count by one. This way, the watchdog module can keep long-term track of the counts for each bin without storing the actual personally identifiable information for each record. In case compliance with an anonymization scheme can be determined by the counts of the bins, the watchdog module can decide on the shareability of bins of data in the live database solely based on such tokens and the respective counts. The risks caused by a possible attack on the watchdog module is limited: an attacker would only see the tokens and the counts.
[0085] Figure 1 B shows a detailed example of the method as may be performed by the watchdog module. In Step S202, a batch of data is received. In Step S204, for each bin of each anonymization scheme, the current token and current bin count is obtained, e.g., from a memory connected to the watchdog module. In Step S206, for each record in the batch of data, and for each bin of each anonymization scheme, the respective token is calculated and the counter incremented if appropriate (if the record would fall within the bin in question).
[0086] For example, a record of the batch of data indicates the subject’s age to be 22 years and their height to be 172 cm and the first anonymization scheme in question bins according to categories of decades and decimetres. The watchdog module calculates that this subject falls in the bin with the name “20- 29 170-179” in accordance with the first anonymization scheme and calculates the hash value of this name as “7e7a2”. The watchdog module then searches the token “7e7a2” in its memory and increases that token’s counter by one. In this example, the watchdog module further evaluates the compatibility with a second anonymization scheme that bins according to half-decades and decimetres and further calculates that this subject falls in the bin with the name “20-24 170-179” in accordance with the second anonymization scheme and calculates the hash value of this name as “0ag7cts” and increases that token’s counter by one. After the watchdog module has calculated the token that that subject falls into for each anonymization scheme, the watchdog module may delete the record so that only the increased counter remains as information with the watchdog module.
[0087] In Step S208, for each anonymization scheme of the plurality of anonymization schemes, the shareability of the respective bins are checked (e.g., against a reidentification risk threshold). If a shareable bin’s size was increased or if a new bin became sharable a transmit share enable message is transmitted, in Step S210.
[0088] Figure 1 C shows an example of the method as may be performed by the data sharing platform. In Step 302, the receive share enable message referred to previously is received. In Step S304, the data sharing platform integrates the batch of data into the live dataset to arrive at an updated live dataset (i.e. , the updated extant dataset). Step 304 can be performed before, after, and / or in parallel to Step 302. In Step S306, in accordance with the share enable message, the data sharing platform anonymizes the at least a subset of the updated live dataset according to at least one anonymization scheme as indicated in the share enable message. The anonymized data is then published in Step S308.
[0089] In some examples, some of the records to be shared may have already been anonymized earlier (e.g., during an initial process or an earlier iteration) and only the records that have not been anonymized before are now anonymized and added to the earlier anonymized record.
[0090] Figure 1 D shows a detailed example of a method performed as may be by the data sharing platform. The dashed box indicates that the Step S402 - S408 take the place of Steps S304 and S306 discussed previously. Again, the share enable message is received in Step S302, and in this instance indicates that only some of the data of the batch of data is newly shareable. In Step S404, the shareable portion of the batch of data is anonymized and this newly anonymized data is added to the previously anonymized data in Step S408. Finally, this (updated) anonymized data is published in Step S410.
[0091] Figure 1 E shows a detailed example of a method as may performed by the data sharing platform. The dashed box indicates that the steps S502 - S508 take the place of Steps S304 and S306. Again, the share enable message is received in Step S302. This time, the share enable message denotes that the data in the batch of data and some previously withheld data in the extant / live dataset are newly shareable. In Step 504, the shareable portion of the batch of data (which may only be a subset, or may be all of the batch of data) is anonymized. In Step S506, a previously withheld but now shareable portion of the extant / live dataset is also anonymized. Steps S504 and S506 may be performed in parallel, or in series in any order. In Step S508, the newly anonymized data (including both the portion of the batch of data and the previously withheld data from the extant / live dataset) is added to previously anonymized data. Next, in Step S510, the update anonymized data is published. Figure 1 F shows a detailed example of a method as may be performed by the data sharing platform. The dashed box indicates that steps S602 - S606 take the place of Steps S304 and S306. Again, the share enable message is received in Step S302. This time, the share enable message denotes (which may be in addition to other share information) that at least a subset of the data in the live / extant dataset, once updated, can be shared according to a new anonymization scheme. In Step S604, the batch of data is integrated into the live / extant dataset to arrive at the updated extant dataset (which could be done before, after, and / or in parallel to Steps 302 and 602). In Step S606, the shareable (under the new anonymization scheme) portion of the updated extant dataset is anonymized in accordance with the new anonymization scheme. This newly anonymized data is then published in Step S608.
[0092] Figures 2A and 2B illustrate in graphical form an example of determination of compatibility with a anonymization scheme using bins. Figure 2A shows the situation for the extant / live dataset before the integration of a new batch of data. Of the four depicted bins of data (e.g., four classifications into which personally identifiable information can be sorted, such as age-decades), only one has a sufficient counter (e.g., number of entrants) to pass the threshold size and therefore be shareable with an acceptable risk of identification. By threshold size, it is meant the size of the bin (i.e., the value of the count) required for the risk of reidentification to be below a reidentification risk threshold. This is discussed in more detail below. This means that at this stage only the data in or corresponding to Bin 1 is shared in anonymized form, as anonymized data. Figure 2B shows the situation when a new batch of data is received and integrated into the live / extant dataset to arrive at the updated live / extant dataset. This is indicated by the newly indicated (grey) portions of the respective bins, which have grown larger as a result of the new batch of data being received. This results in the data in or corresponding to Bins 2 and 3 newly being shareable after the extant dataset is updated, whilst data in or corresponding to Bin 4 still cannot be shared. Also, the data newly added to Bin 1 can also be shared (in addition to the already shared data of Bin 1). The watchdog module may only store indications of the bins and their respective sizes and transmit, in the depicted case, that the updated Bins 1 , 2, and 3 may be shareable.
[0093] Figures 3A and 3B show possible dependencies for an example of the disclosed methods. The possible dependencies may include yet further examples and / or features which are not depicted. Figure 3A shows an example for an initial data workflow. The initial data is stored in the extant data storage (e.g., unmodified, processed, and / or anonymized). The initial data is tokenized according to all anonymization schemes by creating every token according to every defined anonymization scheme and allocating the respective counter to each of the tokens. This tokenized meta information is stored in the watchdog module. Based on the tokenization, one or more chosen subsets of the initial dataset are anonymized in accordance with respective anonymization schemes, and this anonymized data is stored in the anonymization data storage, from where it can be published, e.g., shared with a selection of entities.
[0094] Figure 3B shows an example for a workflow for a later update of the extant data set. The update is performed by adding a batch of data to the extant data in the extant data storage. The batch of data is tokenized according to all anonymization schemes by calculating every token for the batch of data according to every anonymization scheme and increasing the respective counters. This updated tokenized meta information is again stored in the watchdog module. Based on the tokenization, one or more chosen subsets of the updated extant dataset are anonymized in accordance with respective anonymization schemes. The thereby anonymized data is used to update the anonymization data storage. Based on the tokenization, one or more chosen subsets of the initial dataset anonymization schemes are anonymized in accordance with respective anonymization schemes, and this anonymized data is stored in the anonymization data storage. The updated anonymized dataset in the anonymization data storage can be published, e.g., shared with a selection of entities.
[0095] The tokenization may be performed by the watchdog module or by another entity that then sends the tokens and the respective counters or counter updates to the watchdog module. The watchdog module takes the decision which subsets of the (updated) extant database shall be shared under which anonymization schemes. If data could be shared under multiple anonymization schemes, the watchdog module may be configured to prioritize anonymization schemes e.g., based in its granularity and / or the amount of data being shareable under the respective anonymization scheme. The watchdog module may be configured to share different data subsets under different anonymization schemes, at least in suited situations.
[0096] Table 1 is an example of the tokenization to be utilized by a watchdog module. The entries of columns 1 and 2 comprising personally identifiable information (in this case age and height) that are used to compute the tokens, and the columns ‘token’ and ‘count’ may be stored in the watchdog module and used for assessing compatibility with the anonymization scheme in accordance with which these tokens were created. The tokens in this example correspond to the entry in column 1 concatenated with the entry in column 2 (so e.g., 20-30 joined to 110-120) and then hashed. This allows that, for each new batch of data, the tokens to be made in the same manner and so the watchdog module in each case knows which token should be increased. Based on each counter, and as discussed above, the watchdog module can choose the appropriate anonymization scheme (e.g., the one currently used and / or a new, e.g., more granular, scheme). Storing only the count related meta information in the watchdog module (in a uniform manner) can allow the watchdog module to store sufficient information to decide on whether and how to share the data. That is, it allows the watchdog module to decide on the anonymization scheme(s) to be used for the overall data (which comprises the existing and new data) and which parts of the overall data to be shared according to that scheme, without storing any sensitive information (such as the medical values and / or the personally identifiable information). In some examples, the watchdog module may cease counting for a specific bin after the bin relating to this token has been deemed shareable (according to the reidentification risk) and instead replace the count value with an indicator that the threshold has been reached. For this purpose, the watchdog module may e.g., store the column “Shared (present in anonymized storage)”, which in this case indicates which bins of the present extant database have been shared. The column “Candidate (count > threshold)” in this example indicates that with the newly added batch of data the respective bin has sufficient count to be shared. By comparing it to the prior column, one can deduct which bins are newly sharable, which in Table 1 is reflected by an entry of “1” in the column “Refresh benefit”.
[0097]
[0098] Table 1
[0099] In the example of Table 1 , prior to adding the new batch of data, only the bin represented by the first row of Scheme A was shared (i.e. the bin of age “20-29” and height “170-179”) using the anonymization scheme A. After adding the batch of data to create the updated extant database, the bin represented by the first and - newly - the bin represented by the fourth row (i.e. the bin of age “30-39” and height “180-189”) are shareable under Scheme A. In addition, under Scheme B, the bin represented by the respective first row (i.e. the bin of age “20-24” and height “170-179”) and the bin represented by the respective seventh row (i.e. the bin of age “35-39” and height “170-179”) are sharable. That means that in this case, a first subset of the data could be shared under Scheme A and a second subset could be shared under Scheme B. In the depicted case, Scheme B is more granular than Scheme A. Sharing more granularly anonymized datasets may allow for improving research on this data.
[0100] The watchdog module may be configured to decide which of these subsets may be shared in this case, for instance prioritizing a more granular scheme under certain conditions. In some cases, the watchdog module may be configured to allow publishing multiple subsets under multiple schemes.
[0101] Figure 4 is a plot of scenario delay in time against record rate (1 / time). With use of existing data, it is possible to estimate when a new anonymization scenario (i.e. a data set under a given anonymization scheme) will be released based on the ‘record rate’, or number of new records per unit of time. Such estimate may be used for planning research activities using the respective dataset.
[0102] Figure 5 is a plot of batch number against ratio of records above a re-identification risk threshold for a given anonymization scheme. The figure shows the developments for different datasets A to I, which are gradually updated by new batches of data. In the depicted example, a subset of each of datasets A to C are shared starting from the initial step (batch 0), while a subset of each of datasets D to H are shared at later stages. The graphs occasionally show declines, which indicate that new data is added without being shared. As shown by the graphs, this method typically allows for larger and larger portions of the live dataset being shared (the exception in this example is dataset I for which no data is shared even after the 49thupdate). The method described herein therefore is shown to allow for the creation of larger and larger datasets and therefore supporting or even enabling data related research.
[0103] Figure 6 is a diagram showing a computer network. A data sharing platform 602 is connected to a data source 605 and configured for receiving data, e.g., batches of data, therefrom. The data sharing platform 602 is further connected to one or more client devices 606a - 606n and configured for publishing / sharing anonymized data to / with them. The data sharing platform 602 is also connected to a watchdog module 605 and configured for receiving share enable messages therefrom. The watchdog module 603 is configured for performing parts of the methods discussed above, e.g., the methods exemplified in Figs. 1A-1 B. The data sharing platform 602 is configured for performing parts of the methods discussed above, e.g., the methods exemplified in Figs. 1C-1 F. In the depicted example, it is configured to receive a batch of data and transmit it or a representation (e.g., meta information) of it to a watchdog module 603. As depicted by the dotted arrow, the batch of data or its representation could instead be provided from the data source 605 directly to the watchdog module 603. The watchdog 603 uses the batch of data or its representation for determining which parts of the updated extant dataset, i.e. the extant dataset plus the batch of data, can be published under one or more given anonymization scenarios. The watchdog module 603 thereby serves as a watchdog for controlling which data can be shared in which manner by the data sharing platform 602. The watchdog module 603 itself does not need to, and typically will not, store personally identifiable information (e.g., age, address) or sensitive data (e.g., test results). Rather, it utilizes meta information (e.g., bin counters) of the same as its long term memory to determine when and what data in the extant data storage is shareable under various anonymization scenarios.
[0104] The (updated) extant dataset (or a representation) may be stored in the data sharing platform 602. After receiving the share enable message, the data sharing platform 602, which serves as an anonymized data storage, publishes an anonymized version of the indicated subset(s) of the updated extant dataset to the client devices 606a - 606n. The anonymization may, at least in part, have been done before receiving the latest share enable message. It may even be the case that the data sharing platform 602 comprises only an anonymized version (or another representation) of the updated extant dataset, e.g., an anonymization of each of the records according to each of the anonymization schemes, and the original version of the extant dataset remains which the data source 605.
[0105] The data sharing platform 602 and the data source 605 may be on the same network, or may even be part of (or connected to) a same system, e.g., a hospital information system, a laboratory information system, or a laboratory middleware. Each of the depicted entities, in particular the data sharing platform 602 and watchdog module 603, may be a virtual computer present on a cloud computing platform, or a physical computer. The data sharing platform 602, the watchdog module 603, and the client devices 606a - 606n will typically be separated from another, e.g., separated by a firewall, and / or subject to different administration.
[0106] Figure 7 is a diagram showing a computer network which differs from Fig.6 in that there are multiple data sources 605a - 605m. The data of these data sources may each not be shareable as is, but shareable if unified to a larger dataset. The data sources 605a - 605m provide the initial data and the batches of data (or representations thereof, e.g., meta information) to the data sharing platform 602, which then transmits the data (or representations thereof) to the watchdog module 603. Of course, the batch of data or its representation (e.g., meta information thereon) could instead be provided from the data source 605 directly to the watchdog module 603 in this case as well. Figure 8 is a diagram showing a computer network. In the depicted example, a live / extant dataset is distributed across each of a plurality of data sources 605a - 605m, and each data source contains only a portion of the overall dataset. The data sources directly transmit the batches of data (or representation thereof) to the watchdog module 603, which creates a long term memory on the compatibility with one or more anonymization schemes based on meta information deducted from the batches of data (or representation thereof). The watchdog module 603 determines the subset(s) of this distributed (updated) extant datasets and the anonymization scheme(s) under which they can by published, and sends respective share enable messages to the data sources 605a - 605m. The watchdog module may only send share enable messages to those data sources that are affected by the current iteration. The data sources 605a - 605m transmit anonymized data in accordance with the received share enable messages to the data sharing platform 602 which unites them to an anonymized dataset that the data sharing platform 602 shares with the client devices 606a - 606n. In this example the data sharing platform does not receive non-anonymized data and only acts as the anonymized data storage. The extant data storage in this case is the group of data sources 605a - 605m, each comprising only a portion of the extant dataset. The watchdog 603 module may only receive an anonymized representation of the batches of data and / or delete the batches of data after updating the meta information, and, thus, in this example, there is no single entity that has access to the complete original data of the extant dataset at any point. Thereby, the method allows for sharing anonymized datasets that otherwise could not be created, thereby enabling research that otherwise may not be possible. In the example of Figure 8, all of the components (data sources 605, data sharing platform 602, watchdog module 603, and client devices 606) may be separated from one another, e.g., separated by a firewall, and / or subject to different administration.
[0107] In a variation of this example, the data sources 605a - 605m could even share the anonymized data directly the data shares with the client devices 606a - 606n, e.g., where the data sources 605a - 605m and the client devices 606a - 606n overlap (or are identical), e.g., form a federated network for data exchange.
[0108] The systems and methods of the above embodiments may be implemented in a computer system (in particular in computer hardware or in computer software) in addition to the structural components and user interactions described.
[0109] The term “computer network” includes the hardware, software and data storage devices for embodying a system or carrying out a method according to the above described embodiments. For example, a computer system may comprise a central processing unit (CPU), input means, output means and data storage. The computer system may have a monitor to provide a visual output display. The data storage may comprise RAM, disk drives or other computer readable media. The computer system may include a plurality of computing devices connected by a network and able to communicate with each other over that network.
[0110] The methods of the above embodiments may be provided as computer programs or as computer program products or computer readable media carrying a computer program which is arranged, when run on a computer, to perform the method(s) described above.
[0111] The term “computer readable media” includes, without limitation, any non-transitory medium or media which can be read and accessed directly by a computer or computer system. The media can include, but are not limited to, magnetic storage media such as floppy discs, hard disc storage media and magnetic tape; optical storage media such as optical discs or CD-ROMs; electrical storage media such as memory, including RAM, ROM and flash memory; and hybrids and combinations of the above such as magnetic / optical storage media.
[0112] While the disclosure has been described in conjunction with the exemplary embodiments described above, many equivalent modifications and variations will be apparent to those skilled in the art when given this disclosure. Accordingly, the exemplary embodiments of the disclosure set forth above are considered to be illustrative and not limiting. Various changes to the described embodiments may be made without departing from the spirit and scope of the disclosure.
[0113] In particular, although the methods of the above embodiments have been described as being implemented on the systems of the embodiments described, the methods and systems of the present disclosure need not be implemented in conjunction with each other, but can be implemented on alternative systems or using alternative methods respectively.
[0114] The features disclosed in the description, or in the following claims, or in the accompanying drawings, expressed in their specific forms or in terms of a means for performing the disclosed function, or a method or process for obtaining the disclosed results, as appropriate, may, separately, or in any combination of such features, be utilised for realising the disclosure in diverse forms thereof.
[0115] While the disclosure has been described in conjunction with the exemplary embodiments described above, many equivalent modifications and variations will be apparent to those skilled in the art when given this disclosure. Accordingly, the exemplary embodiments of the disclosure set forth above are considered to be illustrative and not limiting. Various changes to the described embodiments may be made without departing from the spirit and scope of the disclosure. For the avoidance of any doubt, any theoretical explanations provided herein are provided for the purposes of improving the understanding of a reader. The inventors do not wish to be bound by any of these theoretical explanations.
[0116] Any section headings used herein are for organizational purposes only and are not to be construed as limiting the subject matter described.
[0117] Throughout this specification, including the claims which follow, unless the context requires otherwise, the word “comprise” and “include”, and variations such as “comprises”, “comprising”, and “including” will be understood to imply the inclusion of a stated integer or step or group of integers or steps but not the exclusion of any other integer or step or group of integers or steps.
[0118] It must be noted that, as used in the specification and the appended claims, the singular forms “a,” “an,” and “the” include plural referents unless the context clearly dictates otherwise. Ranges may be expressed herein as from “about” one particular value, and / or to “about” another particular value. When such a range is expressed, another embodiment includes from the one particular value and / or to the other particular value. Similarly, when values are expressed as approximations, by the use of the antecedent “about,” it will be understood that the particular value forms another embodiment. The term “about” in relation to a numerical value is optional and means for example + / - 10%.
Claims
CLAIMS1 . A computer-implemented method for determining if adding a batch of data to an extant dataset, thereby creating an updated extant dataset, would result in additional data being shareable under an anonymization scheme, whereby the data in the batch of data and the extant dataset contain or represent personally identifiable information linked to one or more values, the method comprising: receiving the batch of data or a representation thereof; determining whether at least a subset of the updated extant dataset including at least a portion of the received batch of data would be compatible with the anonymization scheme; and responsive to determining that it would be, transmitting to a recipient a share enable message so as to cause the creation of an anonymization of the at least a subset of the updated extant dataset in accordance with the anonymization scheme and the publication of this anonymized data.
2. The computer-implemented method of claim 1 , further comprising: obtaining meta information relating the extant dataset’s compatibility with the anonymization scheme; creating meta information relating to the updated extant dataset’s compatibility with the anonymization scheme; wherein the step of determining whether the at least a subset of the updated extant dataset would be compatible with the anonymization scheme is based on the meta information relating to the updated extant dataset’s compatibility.
3. The computer-implemented method of claim 1 or 2, wherein the method includes, upon receipt of the share enable message, one of:(a) anonymizing the batch of data or portion thereof and integrating the anonymized version to an anonymized version of the extant dataset which has been previously anonymized according to the anonymization scheme;(b) anonymizing the batch of data or portion thereof, together with data in the extant dataset according to the anonymization scheme; and / or(c) anonymizing the updated extant dataset or a shareable portion thereof according to an anonymization scheme different to an anonymization scheme according to which the extant dataset has previously been anonymized.
4. The computer-implemented method of claim 3, wherein step (c) is performed when it is determined that the updated extant dataset is compatible with the different anonymization scheme, wherein the further anonymization scheme provides a more granular anonymized representation of the data than the first anonymization scheme.
5. The computer-implemented method of any of claims 3 to 4, wherein anonymizing the received batch of data, portion thereof, or the updated extant dataset includes data obfuscation and / or generalization of personally identifiable information.
6. The computer-implemented method of any preceding claim, wherein compatibility with the anonymization scheme is determined by comparing a calculated re-identification risk to a reidentification risk threshold.
7. The computer-implemented method of any claim 2 to 6, wherein the representation of the extant dataset comprises encrypted meta information indicating a plurality of categories into which personally identifiable information can be sorted.
8. The computer-implemented method of any preceding claim, wherein the method is triggered by the receipt of the batch of data, and repeated for a plurality of subsequentially received batches of data.
9. The computer-implemented method of any preceding claim, including an initialisation process, the initialisation process including steps of: receiving an initial dataset, the initial dataset comprising personally identifiable information linked to one or more values; anonymizing at least a subset of the initial dataset according to the anonymization scheme; and publishing this anonymized data.
10. The computer-implemented method of any preceding claim, wherein the determination step is performed by a watchdog module which is on a computer network separate to the recipient of the share enable message.
11. The computer-implemented method of claim 10, wherein no non-anonymized personally identifiable information is received or stored by the watchdog module.
12. The computer-implemented method of any preceding claim, wherein the determination step includes determining, for each of a plurality of different anonymization schemes, which of the plurality of anonymization schemes are compatible with at least a subset of the updated extant dataset, and including in the share enable message: an indication of which of the plurality of anonymization schemes are found to be compatible with at least a subset of the updated extant dataset, and an indication of the at least a subset of the updated extant data found to be compatible with the respective anonymization scheme.
13. The computer-implemented method of any preceding claim, the method comprising: after receiving the batch of data or the representation thereof, obtaining meta information relating the extant dataset’s compatibility with one or more anonymization schemes; creating meta information relating to the updated extant dataset’s compatibility with each of the anonymization schemes;determining whether the at least a subset of the updated extant dataset including at least a portion of the received batch of data would be compatible with at least one of the anonymization schemes based on the meta information relating to the updated extant dataset’s compatibility; and responsive to determining that it would be, transmitting to the recipient the share enable message so as to cause the creation of an anonymization of the at least a subset of the updated extant dataset in accordance with the at least one of the anonymization schemes and the publication of this anonymized data.
14. The computer-implemented method of the previous claim, further comprising storing the created meta information relating to the to the updated extant dataset’s compatibility with the one or more anonymization schemes.
15. A computer network, comprising: a data sharing platform, configured to: store an extant dataset, the extant dataset containing or representing personally identifiable information linked to one or more values; obtain a batch of data, the batch of data containing or representing personally identifiable information linked to one or more values; and transmit, to a watchdog module, the batch of data or a representation thereof; the watchdog module being configured to: receive the batch of data or the representation thereof; determine whether at least a subset of the updated extant dataset, created by adding the batch of data to the extant dataset, is compatible with an anonymization scheme; and responsive to determining that it would be, transmitting to the data sharing platform a share enable message so as to cause the creation of an anonymization of the at least a subset of the updated extant dataset in accordance with the anonymization scheme and the publication of this anonymized data.
Citation Information
Patent Citations
Computer-implemented privacy engineering system and method
US20230359770A1
Anonymizing time-series data using matrix profile
WO2024059538A1