ANONYMIZE DATA

DE502023002224D1Active Publication Date: 2025-12-11SIEMENS HEALTHINEERS AG
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
DE502023002224
Authority / Receiving Office
DE · DE
Patent Type
Patents
Current Assignee / Owner
Filing Date
2023-01-17
Publication Date
2025-12-11
Estimated Expiration
2043-01-17

AI Technical Summary

Technical Problem

Existing anonymization methods require all data to be collected before anonymization can be applied, leading to prolonged unusability of the data for secondary purposes.

Method used

A method and device for anonymizing data using generalization, allowing anonymization of initial data records with a first set of mapping ranges, and subsequent data records with increasingly more comprehensive mapping ranges, ensuring k-anonymity and l-diversity without waiting for complete data collection.

Benefits of technology

Enables continuous anonymization and processing of data as it is collected, maintaining anonymity and enhancing data usability over time, while adhering to data protection regulations.

✦ Generated by Eureka AI based on patent content.
Patent Text Reader
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL AREA

[0001] The present invention relates to methods and devices for anonymizing data. In particular, the present invention relates to a computer-implemented method for anonymizing data by means of generalization, as well as a corresponding device. BACKGROUND OF THE INVENTION

[0002] Anonymization involves altering personal data in such a way that this data can no longer be attributed to a specific or identifiable natural person, or only with a disproportionate amount of time, cost and manpower.

[0003] This description includes persons of male, female or other gender identities, regardless of the grammatical gender of a particular term.

[0004] For example, an institution or group of institutions may collect personal data from individuals for a primary purpose and wish to use the data for other purposes, known as secondary purposes. For instance, they might want to share the information with other institutions. For example, a group of healthcare providers (e.g., hospitals) collects personal data from patients to treat them. This data may include personal information such as name, date of birth, address, and the like, as well as health data such as the type of illness, blood test results, X-rays, etc. The healthcare providers then wish to use the data to promote clinical research, for example, for clinical trials. According to the General Data Protection Regulation (GDPR) and other regulations, this is only possible under very strict conditions.In this context, anonymization techniques are very helpful, as anonymized data cannot be associated with an identifiable person (or "data subject") and are therefore no longer protected by data protection regulations.

[0005] A fundamental principle of these data protection regulations is data minimization, which states that whenever personal data is required for a specific primary purpose, the amount of data collected or processed must be limited to what is necessary in relation to those purposes. This applies particularly to personal health data, which is subject to very strict regulations for the protection of privacy and data protection, such as the GDPR (General Data Protection Regulation) in the EU or the HIPAA (Health Insurance Portability and Accountability Act) in the USA.

[0006] One technique that contributes to compliance with the principle of data minimization is anonymization, where the persons responsible for data processing do not need access to the actual identity details or the actual values ​​of the data attributes, but the minimized data is still useful for the specified purposes, in particular for the secondary purposes mentioned above.

[0007] In most cases, much personal data is captured and stored in structured formats as datasets with attribute-value pairs. Some examples in the field of medical data are structured formats such as DICOM and the HL7 (Health Level 7) standards, for example, the FHIR (Fast Healthcare Interoperability Resources) standard. These can be used to store vital signs measurements, results of medical examinations, and especially information from imaging devices (such as CT, MRI, and ultrasound).

[0008] According to Recital 26 of the EU General Data Protection Regulation (GDPR), the principles of data protection should not apply to anonymous information, i.e., information that does not relate to an identified or identifiable natural person, or personal data that has been anonymized in such a way that the data subject cannot be identified or is no longer identifiable. This regulation then also no longer applies to the processing of such anonymous data, for example, for statistical or research purposes.

[0009] A database is anonymized when the identity of the data subjects is unknown, unidentifiable, inaccessible, and untraceable (untrackable). Anonymization refers to a technique for removing information that could be used to identify or otherwise link the data subjects to them, or to obtain sensitive information about them. Anonymized data should not relate to any specific or identifiable natural person, nor to other anonymized personal data, even if the data subject is not, or is no longer, identifiable. In other words, anonymous data should not be able to be linked, either directly or indirectly, to the individuals from whom it originates.The idea is that the resulting anonymized data requires no further privacy protection and, in particular, does not require a special mechanism for managing information and relationships after release. It must be (virtually) impossible to derive information, i.e., to deduce the value of one attribute from the values ​​of a set of other attributes with a high degree of probability.

[0010] A very well-known and widely used method is k-anonymization and related techniques, which are connected to concepts such as t-closeness, l-diversity, and differential privacy. k-anonymity is a simple and easy-to-use technique. The basis is that the data records of the various subjects (e.g., people) are modified using one of the techniques "suppression" or "generalization" so that each data record corresponding to a subject (i.e., a person) can no longer be distinguished from a relatively large set of other subjects. More precisely, the anonymity set is the set of subjects that share the same attributes, such that they cannot be distinguished from one another within the given context. "Relative anonymization" refers to this indistinguishability from the perspective of a specific group of observers to whom the anonymized database is disclosed.

[0011] Therefore, if the anonymity groups are large enough – in the case of k-anonymization they contain at least k subjects – the subjects are considered anonymous, or more precisely, k-anonymous.

[0012] However, if the entire database is not known in advance, existing methods cannot be applied. In other words, to use existing methods, one must wait until all data from all participants has been entered into the database to analyze the required anonymization levels for specific degrees of suppression or generalization. In some cases, the data is therefore unusable for a relatively long period, for example, until all data has been collected and recorded.

[0013] The publication "AN APPLICATION OF INTEGRATED CLUSTERING TO MRI SEGMENTATION" by Vito di Gesu et al. (pattern recognition letters, Elsevier, Amsterdam, NL, vol. 15, no. 7, July 1, 1994 (1994-07-01), pages 731-738, XP000453039, ISSN: 0167-8655, DOI: 10.1016 / 0167-8655(94)90078-7) discusses the use of magnetic resonance imaging (MRI) in computer-aided diagnostic procedures. MRI images are multidimensional, so radiologists must combine different sources of information to classify tissues for diagnostic purposes. Four automatic clustering methods are described and their integration into an "Information Fusion Clustering" approach, which supports the preliminary segmentation and classification of body tissue.

[0014] The article "Secure Anonymization for Incremental Datasets" by JI-WON BYUN et al. (January 1, 2006, SECURE DATA MANAGEMENT LECTURE NOTE IN COMPUTER SCIENCE; LNCS, SPRINGER, BERLIN, DE, PAGES 48-63, XP019040079, ISBN: 978-3-540-38984-2) addresses how to ensure k-anonymity and l-diversity for incremental, continuously growing datasets. The article proposes a method that enables the anonymization of growing datasets without causing redundant computations. SUMMARY

[0015] One object of the present invention is therefore to provide an automated method for anonymizing databases, which can also be used when not all entries of the database are yet known, but the data are entered by different subjects only over time.

[0016] According to the present invention, this problem is solved by a computer-implemented method for anonymizing data and a device for anonymizing data, as defined in the independent claims. The dependent claims define embodiments of the invention.

[0017] Regardless of the grammatical gender of a particular term, persons with male, female or other gender identities are included.

[0018] According to the present invention, a computer-implemented method for anonymizing data by means of generalization is provided. The data comprise, at a first time point, a first set of first data records and, at a second time point, a second set of second data records. The second time point is, for example, temporally after the first time point, i.e., the second time point is later than the first time point. The first data records are a subset of the second data records; in particular, the first data records are a true subset of the second data records. In other words, less data is available at the first time point than at the second time point. At the first time point, only the first data records are available, whereas at the second time point, the first data records and further data records are available, the first data records and the further data records together constituting the second data records.The interval between the first and second data points can be several days or weeks, and further processing of the data for secondary purposes may be required at or shortly after the first point in time. Therefore, anonymization is necessary at the first point in time. The procedure involves generating an initial generalization for the first data records that fulfills the required anonymization criteria. These criteria could be, for example, k-anonymity, t-closeness, or l-diversity. The initial generalization comprises a first set of mapping ranges used to generalize the values ​​of a quasi-identifier within the data. These mapping ranges could, for example, be value intervals used to abstract the values ​​of the quasi-identifier.The mapping areas can also be multidimensional, with value ranges assigned to multiple quasi-identifiers, where each dimension of the mapping area is assigned one of the multiple quasi-identifiers. Based on this initial generalization, the first data records can already be anonymized and thus processed anonymously. The procedure further includes generating a second generalization for the second data records, which fulfills the required anonymization. For example, the second generalization can be generated when the second data records are available at the second time point. The second generalization comprises a second group of mapping areas, with which values ​​of the quasi-identifier are generalized. The second group of mapping areas includes more mapping areas than the first group.Each generalization thus achieves the desired anonymization, allowing the initial data sets collected to be processed in a sufficiently anonymized form right from the start. As soon as further data sets are collected, they can be anonymized using the second generalization and thus successively made available for further processing.

[0019] It is clear that the procedure is not limited to a first and second generalization at a first and second time point, but that the procedure can be used for further generalizations, e.g., a third and a fourth generalization, etc., to anonymize the entirety of data records available at the respective time points and make them available for further processing. Since the second or subsequent group of assignment domains of the second or subsequent generalization has increasingly more assignment domains than the first or preceding group of assignment domains of the first or preceding generalization, the benefit gained from the anonymized data, for example, in an evaluation for scientific research, can increase continuously with increasing data volume.

[0020] For example, the data at a later point in time might comprise a further number of additional records. These additional records are a superset of previous records at a previous point in time, preferably a true superset. According to the method, a further generalization for these additional records is generated based on a previous generalization. This further generalization includes another group of mapping domains for the quasi-identifier. This further group includes more mapping domains than a previous group of mapping domains in the preceding generalization. The preceding generalization could be, for example, the second generalization or a further generalization iteratively based on the second generalization.

[0021] The assignment ranges can be chosen, for example, such that for each value of the quasi-identifier, the value is assigned to at most one of the assignment ranges of the first group, and for each value of the quasi-identifier, the value is assigned to at most one of the assignment ranges of the second group. By choosing the assignment ranges so that each value of the quasi-identifier is assigned to only one of the assignment ranges of the respective group, a desired k-anonymity can be achieved in a simple way.

[0022] Furthermore, a set of values ​​assigned to the entirety of the assignment areas of the first group can be identical to a set of values ​​assigned to the entirety of the assignment areas of the second group. Even if, at the outset, due to the smaller number of initial data records and the thus potentially limited set of values, it is not necessary to make the first group of assignment areas so comprehensive that it extends far beyond the set of values ​​of the initial data records, this can be useful in order to minimize changes during subsequent processing of the anonymized data and to prevent inferences about sensitive data that would be possible by increasing the totality of assignment areas.

[0023] In further examples, for each assignment domain of the second group, a set of values ​​of that assignment domain is a subset of a set of values ​​of exactly one assignment domain of the first group. It is important to note that these need not be proper subsets; that is, some assignment domains of the second group may be identical to a corresponding assignment domain of the first group. However, some assignment domains of the second group are proper subsets of exactly one assignment domain of the first group. In other words, an assignment domain of the second group never extends over two assignment domains of the first group. Consequently, an assignment domain of the second group either corresponds exactly to a corresponding assignment domain of the first group, or an assignment domain of the first group is partitioned into two or more assignment domains of the second group.This ensures that the desired anonymity is maintained even when considering the allocation of data jointly, first into the allocation areas of the first group and later into the allocation areas of the second group.

[0024] For example, each allocation area of ​​the first group is assigned a specific interval of values ​​(e.g., age or postal codes). The intervals of the first group have independent interval lengths. The length of each interval in the first group depends on the number of records from the first data set assigned to that interval. This allows for a simple way to achieve the desired generalization, such as K-anonymity, at least with respect to an attribute of the first data set assigned to the respective allocation area.

[0025] In other exemplary procedures, a statistical distribution of values ​​of the quasi-identifier is determined. This determination is based, for example, on additional data comprising datasets that include the quasi-identifier. This additional data may have been collected independently of the data mentioned above, for example, in a different clinical study or in a different context. However, this additional data includes the quasi-identifier as an attribute, allowing a statistical distribution of values ​​for this quasi-identifier to be determined. Assuming that this statistical distribution also applies, at least in the long term, to the data to be processed using the present procedure, a target generalization is generated based on this static distribution.The target generalization comprises a target group of mapping domains used to generalize the values ​​of a quasi-identifier in the data. The first and / or second generalizations are additionally generated depending on the target generalization.

[0026] In further examples, each mapping area of ​​the first group comprises at least a two-dimensional mapping area, with which values ​​of the quasi-identifier and values ​​of at least one other quasi-identifier in the data are generalized. The generalization thus applies to at least two quasi-identifiers. Generalization can also apply to more than two quasi-identifiers. One dimension of the mapping area increases accordingly. With three quasi-identifiers, three-dimensional mapping areas result; with four quasi-identifiers, correspondingly four-dimensional mapping areas, and so on.

[0027] Each assignment domain of the second group comprises a two- or higher-dimensional assignment domain used to generalize values ​​of the quasi-identifier and values ​​of at least one other quasi-identifier of the data. The same principles apply to the two- or higher-dimensional assignment domains as to the intervals mentioned above, which correspond to a one-dimensional case. For example, in the two-dimensional case, the assignment domains of the second group can be designed such that each is completely contained within an assignment domain of the first group. In other words, assignment domains of the second group never overlap two assignment domains of the first group. Therefore, the assignment domains of the second group are either smaller than or the same size as corresponding assignment domains of the first group.Conversely, an allocation area of ​​the first group is either assigned to exactly one allocation area of ​​the second group, or it is divided into two or more allocation areas of the second group, resulting in a greater number of allocation areas of the second group than the number of allocation areas of the first group.

[0028] As described above, information on the quasi-identifiers from other studies can also be used when employing two-dimensional or higher-dimensional mapping domains. For example, statistical distributions of values ​​of the quasi-identifier and values ​​of at least one other quasi-identifier can be determined, perhaps based on additional data from other studies that include datasets containing the quasi-identifier and the at least one other quasi-identifier. A target generalization can then be generated based on these statistical distributions. The target generalization comprises a set of two-dimensional or higher-dimensional mapping domains used to generalize the values ​​of the quasi-identifier and the at least one other quasi-identifier. The first and / or second generalizations can also be generated depending on the target generalization.Especially when very few data records are available when generating the first set of assignment domains, this approach can prevent the creation of boundaries between assignment domains that are later, for example, at least partially unsuitable for the second or third set of assignment domains. This can occur particularly with two-dimensional or higher-dimensional assignment domains during further refinement. Using this "prior knowledge" about the value distributions of the quasi-identifiers from other studies can be advantageous in avoiding such later shortcomings.

[0029] Another aspect of the present invention relates to a device for anonymizing data by means of generalization. The data can be collected and recorded over a longer period of time, for example, days, weeks, or months. The data therefore comprises, at a first time point, a first set of initial data records and, at a second, later time point, a second set of subsequent data records. The initial data records are a subset, preferably a true subset, of the subsequent data records. The device comprises a processing unit configured to generate a first generalization for the initial data records that fulfills the required anonymization criteria. The first generalization comprises a first set of mapping domains with which values ​​of a quasi-identifier of the data are generalized. The processing unit is further configured to generate a second generalization for the subsequent data records.The second generalization comprises a second group of mapping domains used to generalize values ​​of the quasi-identifier. Like the first generalization, the second generalization fulfills the required anonymization. The second group of mapping domains includes more mapping domains than the first group.

[0030] The processing device comprises, for example, a microprocessor controller with memory and input / output devices, such as a computer system or a server. The data to be anonymized can be stored in the processing device's memory. Parameters for generalization, such as the desired degree of anonymization (e.g., the desired value for k in the case of k-anonymity), as well as a categorization of the data attributes into identifiers, quasi-identifiers, and sensitive attributes, can be set, for example, via the input / output devices by an operator or by means of a corresponding configuration file.

[0031] The device is thus designed to carry out the previously described procedure and therefore also includes the advantages described above.

[0032] The present invention further relates to a computer program product comprising a computer-readable program code configured to cause a processing device to perform the steps of the method described above.

[0033] The present invention also relates to a computer-readable storage medium configured to store a computer program product comprising a computer-readable program code configured to cause a processing device to perform the steps of the previously described method.

[0034] It is clear that the features mentioned above and those explained below can be used not only in the combinations specified, but also in other combinations or in isolation from one another, without deviating from the scope of the invention. BRIEF DESCRIPTION OF THE DRAWINGS

[0035] The present invention will below be described in detail with reference to the figures, which illustrate embodiments of the present invention. FIG. 1 schematically shows a relationship between refinements of generalizations. FIG. 2 schematically shows two generalizations at two different times. FIG. 3 schematically shows values ​​of two quasi-identifiers of a set of data records. FIG. 4 schematically shows three different generalizations for the set of data records. Fig. 3FIG. 5 schematically shows generalizations with different interval widths. FIG. 6 schematically shows two generalizations at two different times. FIG. 7 schematically shows two generalizations at two different times. FIG. 8 schematically shows possible generalizations based on an existing generalization, as well as a target generalization. FIG. 9 schematically shows possible generalizations based on an existing generalization, as well as a target generalization. FIG. 10 schematically shows a device for anonymizing data by means of generalization. FIG. 11 schematically shows the procedural steps of a method for anonymizing data by means of generalization. DETAILED DESCRIPTION OF THE EXECUTION FORMS

[0036] Some examples in this disclosure generally relate to one or more circuits, control devices, or other electrical devices. All references to the circuits, control devices, and other electrical devices and the functionality they provide are not intended to be limited to what is shown and described herein. Even if certain designations may be assigned to the various circuits, control devices, and other electrical devices, these designations are not intended to limit the functionality of the circuits, control devices, and other electrical devices. Such circuits, control devices, and other electrical devices may be combined with one another and / or separated in any manner, depending on the type of electrical implementation desired.Naturally, any circuit or other electrical device disclosed herein may comprise any number of microcontrollers, graphics processing units (GPUs), integrated circuits, memory devices (e.g., FLASH, random access memory (RAM), read-only memory (ROM), electrically programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), or other suitable variants thereof), and software, which interact to perform the operations, processes, and / or procedural steps disclosed herein. Furthermore, one or more of the electrical devices may be configured to execute program code contained in a non-transient, computer-readable medium, programmed to perform any number of the disclosed functions.

[0037] The following section describes a general method for anonymizing a given dataset. The data can, for example, be viewed as a large matrix, as shown in the table below: Identifiers Quasi-identifiers Sensitive Attribute name address Postcode birth date Gender BMI diagnosis Mr Müller Lonestr 1 33210 23.12.1945 masculine 27 Cancer Mr. Abel Homestr 5 335021 01.06.1979 female 24 Covid ... ... ... ... ... ... ...

[0038] The rows of the matrix correspond to the data subjects, for example, one row for each person. The columns correspond to the attributes, for example, name, address, postal code, date of birth or age, gender, BMI, and diagnosis. Each row of the table represents a data record relating to a specific subject, and the values ​​in the various columns are the values ​​of the attributes associated with those subjects. Each attribute is further classified into one of the following three categories: Identifiers: These are attributes that, on their own, typically identify individuals, such as name, address, phone number, and ID number. Quasi-identifiers: These are attributes that, in combination and when linked to external information, can be used to identify individuals. That is, even if a single attribute does not identify an individual, combining several of these attributes can be problematic. For example, given a person's date of birth, postal code, and weight or BMI, the probability of identifying that person can be very high. Sensitive attributes: These are attributes containing personal information that should not be publicly linked to a person / user / identifier, such as diagnosis, salary, etc.

[0039] To achieve k-anonymity, the values ​​of the identifier attributes are eliminated (for example, deleted or replaced with random values), and the values ​​of the quasi-identifier attributes (and, in some cases, the values ​​of sensitive attributes) are generalized by replacing the data value with a less precise, semantically consistent value, or suppressed by eliminating the information. Replacing a data value with a less precise value is also referred to as "abstracting" the data value. For example, a person's age in years, e.g., "52," can be abstracted to a less precise age range, e.g., "50 to 60."

[0040] Assuming that the possible data values ​​for the age of patients in a medical study can range from 0 to 120 years, for example, only an age between 40 and 100 years is of interest for the purposes of a study.

[0041] A suppression (English: suppression This information is equivalent to stating that the age lies in the range [0.120], i.e., no information about the age is given. A generalization (English: generalization ) can be done, for example, by using age intervals.

[0042] Regarding the setting of intervals, the following definition is used in this description: [away[ is a half-open interval that includes the first value a and the last value b does not include. Although this description mainly concerns half-open intervals of the form [away[ The methods described herein are not limited to the use of the first value a. Half-open intervals ]a,b] can also be used, which do not include the first value a and include the last value. bInclude the values. Furthermore, alternating open and closed intervals can be used, for example. It is important to note that when choosing intervals, care should be taken to avoid a value being included in two or more intervals, or relevant values ​​not being included in any interval.

[0043] One possible generalization is, for example, to use only the interval [40,100[ and to suppress (delete) all records that are not in this interval.

[0044] Using an example with intervals related to age, the following describes how generalizations can be "refined" and what conditions must be met for a generalization to be "finer than" another generalization. The terms "refine" and "finer than" will be used throughout this description according to the following definitions.

[0045] In the example described below, with reference to FIG. 1 Different generalizations P0 to P3 were used, which have different divisions of the age interval [40,100[.

[0046] The generalization P0 denotes a set which contains only one interval: P0={[40,100[}.

[0047] The generalization P1 contains a set of six intervals: P 1 = 40 50 , 50 60 , 60 70 , 70 80 , 80 90 , 90 100 .

[0048] The generalization P2 contains a set of four intervals: P 2 = 40 55 , 55 70 , 70 85 , 85 100 .

[0049] It is important to note that the intervals of a generalization are non-overlapping; that is, no value can be assigned to two intervals of a generalization. Therefore, each value of a quasi-identifier can be assigned to exactly one interval of a generalization.

[0050] Each of the generalizations P0, P1, and P2 represents a generalization from the specific age values. Specific age values ​​are, for example, precise ages in years, such as 36 years or 52 years. When generalizing specific age values, the age value of each data point is assigned to one of the value intervals of the generalization. For example, when using generalization P0, the age value 52 is assigned to the single interval [40, 100]. When using generalization P1, the age value 52 is assigned to the interval [50, 60], and when using generalization P2, the age value 52 is assigned to the interval [40, 55].

[0051] For the purposes of this description, a generalization X is a generalization X if and only if finer thanA different generalization Y exists if every interval in X is contained in exactly one interval in Y. Generalization P1 is therefore finer than generalization P0. Likewise, generalization P2 is finer than generalization P0. However, although generalization P1 has smaller intervals than generalization P2, generalization P1 is not finer than generalization P2 because, for example, the interval [50, 60] of generalization P1 is not completely contained in exactly one interval of generalization P2. Likewise, generalization P2 is not finer than generalization P1.

[0052] The generalization P3 contains a set of twelve intervals:

[0053] Generalization P3 is finer than generalization P2 and also finer than generalization P1. Each interval in P3 is contained in an interval of P1, and each interval of P3 is also contained in an interval of P2. FIG. 1This relationship is illustrated graphically. An arrow signifies that the starting point of the arrow is associated with a generalization that is finer than a generalization associated with the endpoint of the arrow.

[0054] Furthermore, it states in the FIG. 1 An arrow between the interval sets P3 and P1, for example, indicates that anonymization using the intervals of P3 provides more accurate information for a secondary purpose than anonymization using P1.

[0055] Each combination of quasi-identifiers appearing in the table above defines an anonymization set. Subjects (e.g., persons) with identical quasi-identifiers after anonymization—that is, with the same values ​​of the generalized quasi-identifiers—are indistinguishable. These anonymization sets are also called equivalence classes.

[0056] For example, in k-anonymization, generalization and / or suppression is selected to ensure that all anonymization sets contain at least k subjects.

[0057] If an attribute is suppressed, it can be deleted from data that is passed on for further processing, such as a study or other secondary purposes, or the values ​​of this attribute can be set to a predefined value in all data records, indicating that this attribute contains no usable information. If an attribute is generalized, the concrete value of the attribute is replaced by an abstracted value, for example, a value that indicates the interval in which the actual concrete value of the data record lies. If, for example, the twelve intervals of generalization P3 defined above are labeled A, B, C, ... L, then a concrete age of 52 years in a data record can be replaced by the letter C, which denotes the interval [50,55[. Intervals represent one method for generalization.Generalization involves an abstraction of the values. A mapping is established between the values ​​of the interval and a label for the interval, which represents the abstraction value for the values ​​within that interval. If several quantities are generalized simultaneously, for example, n quantities, a corresponding abstraction can be performed for a corresponding n-dimensional domain. For two quantities, these domains can be viewed as areas, preferably rectangles. More generally, in this description, such an n-dimensional domain is referred to as a mapping domain, which, for example, corresponds to an interval in the one-dimensional case and to an area in the two-dimensional case.

[0058] Besides k-anonymization, there are other similar methods such as t-closeness and l-diversity. The techniques proposed here are also directly applicable to these. However, the three methods mentioned (k-anonymization, t-closeness, and l-diversity) are usually only used when the database to be anonymized is fully known before anonymization. Therefore, to apply these methods when data is collected over a long period, it is generally necessary to wait until all data has been collected before proceeding with suppression and generalization.

[0059] The following describes anonymization techniques, particularly k-anonymization, that enable the automatic anonymization of databases, specifically achieving k-anonymization, even when not all database entries are known a priori, but rather entries for different subjects are entered over time. In other words, these techniques allow the anonymization of data even if not all data is initially available.

[0060] In a first example procedure, a database is anonymized at various points in time. At each point in time, respective "snapshots" of the database are created and anonymized accordingly. Between each snapshot, the database grows, i.e., new data records are added. Each snapshot is anonymized across all non-anonymized data records available at that time.

[0061] For example, at a first point in time, the database contains an initial set of first records, and at a later point in time, a second set of second records. The set of first records is a proper subset of the set of second records. A first generalization is performed on the first records, which, for example, satisfies a required k-anonymity. To do this, as described above, values ​​of quasi-identifiers in the database are generalized, for example, by replacing concrete values ​​with a range of values ​​from a group of value ranges. In other words, for this first generalization, a group of assignment ranges, such as value intervals for age, is defined, and in each record, instead of the concrete value itself, only the assignment range associated with the concrete value is recorded.The generalized database can fulfill the required anonymization criteria and therefore be used for other secondary purposes, such as clinical trials, without violating data protection requirements. The original database, i.e., the database before generalization, can be expanded with additional datasets over time. For example, numerous datasets can be added over the course of a few days, weeks, or months. At a later point in time, a second generalization can be performed on the then-available datasets, which also fulfills the required anonymization criteria, such as the aforementioned k-anonymization. Again, values ​​of quasi-identifiers in the database are generalized at this point. Since more datasets are now available, the second generalization can encompass more mapping ranges than the first generalization.This allows the anonymized database data to have greater significance at the second point in time while still meeting data protection requirements. At later points in time, further snapshots of the dataset can be created, and further generalizations with even more assignment domains can be performed, thereby further improving the usability of the anonymized data in subsequent investigations or studies.

[0062] FIG. 2 shows two different snapshots of a database with generalizations of a quasi-identifier with real, concrete values ​​between 0 and 12.

[0063] At an initial time point T1, for example, there are 14 data records, each of which is represented by a circle in the Figure 2are shown. A first generalization at time T1 contains three assignment ranges, which in the present example are three intervals of width 4, dh The values ​​of the quasi-identifier are generalized or abstracted to the three intervals [0,4[, [4,8[, [8,12[]. At a later time T2, there are 18 data records. A second generalization at time T2 contains four assignment ranges, which in this example are four intervals of width 3. dh the four intervals [0,3[, [3,6[, [6,9[, [9,12[. Each of these intervals is k-anonymous on its own, e.g. 4-anonymous. However, it should be noted that changing the generalization may unexpectedly reveal sensitive information, as demonstrated in [reference to relevant section]. FIG. 2 The example shown illustrates this.

[0064] An observer with access to both snapshots T1 and T2 can determine that, for example, at time T1, one of the records (marked with an arrow) is in the interval [0,4[ and at time T2, it is in the second interval [3,6[). To do this, the observer can, for instance, analyze the values ​​of the sensitive attributes of the records and recognize that the two records correspond to the same original record, since the sensitive attributes for the two records are identical in the anonymized databases. The record with these identical sensitive attributes is in [0,4[ at T1 and in [3,6[ at T2). If the record concerns a person, for example, the observer can therefore conclude for that person that the generalized attribute (i.e., the value of the quasi-identifier of this record) is less than 4 but greater than or equal to 3.He is therefore not protected by 4-anonymity if the observer has access to both snapshots of the database.

[0065] This can be avoided, for example, by ensuring that a generalization applied later is more refined than a generalization applied previously. For instance, if a subdivision is chosen at a specific time, e.g., the interval [0,4[ at time T1 in the example above, then only finer subdivisions of [0,4[ should be chosen later, but not subdivisions that extend beyond an interval boundary. In other words, the generalization should be refined each time.

[0066] Another exemplary procedure therefore involves anonymizing snapshots of the database at different times, with each new choice of generalization of quasi-identifiers being finer than the previous one.

[0067] This also applies to databases where multiple quasi-identifiers are generalized, such as age or date of birth and postal code in the table above. In these cases, generalization is a multidimensional problem, with the dimension corresponding to the number of quasi-identifiers.

[0068] Especially in multidimensional problems with many quasi-identifiers, it can happen that at an initial point in time one decides on a division of the assignment domains, which may also be multidimensional, but at that point it is not yet clear which decision is suitable or optimal in the long term, especially considering that subsequent generalizations should represent refinements of the previous generalizations.

[0069] For example, given the data records already in the database, it might initially seem sensible to generalize a particular attribute in a certain way, but in the long run, this could lead to a suboptimal solution. This will be illustrated below using an example from FIG. 3 and FIG. 4 The example shown illustrates this.

[0070] Given are two quasi-identifiers q1 and q2, each with real values ​​between 0 and 12. In FIG. 3 Each circle displays the values ​​of q1 and q2 for a corresponding data set. To achieve a desired anonymization, for example 4-anonymity, it is possible to generalize the values ​​to intervals of width 3 (e.g., [0,3[, [3,6[, [6,9[, [9,12[])) or width 4 (e.g., [0,4[, [4,8[, [8,12[]). FIG. 4 shows corresponding generalizations.

[0071] In the example of the FIG. 3 However, it is at the time when the in FIG. 3Given the existence of the datasets shown, it is not possible to generalize both quasi-identifiers to intervals of length 3. At least one of the two quasi-identifiers must be generalized to intervals of length 4 to achieve 4-anonymity. The three possibilities, including the option of generalizing both to intervals of length 4, are shown in FIG. 4 shown.

[0072] The three in FIG. 4 However, the demonstrated methods of grouping the data sets by dividing the values ​​of the quasi-identifiers into intervals of length 3 or 4 result in none of these generalizations being finer than another.

[0073] If one of the three options is selected at a first time point, it can be difficult to optimally divide these intervals by refining the selection at a later time point. Depending on the values ​​for the additional data subjects that are available at the second time point, one or the other of the generalizations may prove to be ineffective in retrospect. FIG. 4 This would be optimal. However, since there is no information about the expected number of data records still to be collected and their content, deciding on one of the three options is difficult at this initial stage.

[0074] In many cases, however, additional general information is available, such as distributions of quasi-identifier values, which can be considered during generalization at an early stage, particularly during initial generalization. For example, in clinical studies, it may be known that young patients tend to suffer from certain clinical conditions less frequently than older patients, or vice versa for other clinical conditions. Therefore, in optimal classification, the intervals in the areas where a disease occurs less frequently are chosen to be larger than in the areas where the disease occurs more frequently. This can result in intervals of varying lengths, for example, during initial generalization. FIG. 5 shows a corresponding generalization for two quasi-identifiers q1 and q2.

[0075] In the center, where there are more data records, the intervals and the two-dimensional areas are smaller, while they are larger at the edges. Ideally, the number of data records in each rectangle has the same number of elements, with the minimum of this number being used as the value k for k-anonymity.

[0076] In this exemplary procedure, the database is anonymized at various points in time (snapshots), and each new generalization is more refined than the previously used generalization. Furthermore, the mapping areas, such as intervals or rectangles, are chosen to be of different sizes depending on the density of the quasi-identifier values. For example, intervals within a generalization can have different lengths.

[0077] The advantage is that more information can be made available for secondary purposes without violating the required k-anonymity, and subsequent generalizations can be chosen more optimally at later times.

[0078] In the example above, the FIG. 3-5 Each combination of quasi-identifiers, e.g., age, number of hospital visits, and BMI, is abstracted independently. Age is divided into intervals (possibly with intervals of varying lengths), the number of hospital visits is divided into intervals with a specific granularity, and BMI is also divided into a specific number of sets. However, this approach did not take into account that these values ​​may correlate with each other, and therefore some combinations may occur much more frequently than others. For example, the combinations in FIG. 5The resulting "long and narrow rectangles" are not optimal in at least some areas, such as in the FIG. 3 Above, where no records fall into the narrow rectangle at the top center. A subdivision that takes into account that some combinations occur more frequently than others can be achieved using rectangular allocation areas as in the FIG. 6 The representation must be taken into account.

[0079] In this FIG. 6The diagram shows the allocation ranges of two generalizations for a database at two points in time, T1 and T2. Two quasi-identifiers, q1 and q2, are anonymized. One required anonymization is, for example, 4-anonymity. The left side shows the state at the first point in time, T1. At time T2, shown on the right, more data records are available. Accordingly, a finer generalization, i.e., a finer subdivision into rectangles, can be chosen at time T2. As can be seen from the FIG. 6 As can be seen, it also applies here that an assignment domain of the generalization at time T2 is assigned to only exactly one assignment domain of the generalization at time T1, i.e., each assignment domain of the generalization at time T2 is a subset of only one assignment domain of the generalization at time T1.

[0080] In this example, the rectangles can be adjusted to accommodate distributions that are not products of distributions of the individual quasi-identifiers. For instance, the distribution of values ​​in the "bottom right corner" and the "top left corner" might be higher than in the "top right corner." Therefore, the rectangle in the "top right corner" can be wider in both directions, as shown in the... FIG. 6 This is necessary if the quasi-identifiers are not independent of each other.

[0081] In this example, the database is anonymized using various "snapshots" over time. Each new choice of abstraction is more refined than the previous one. The quasi-identifiers are not abstracted independently into intervals or mapping spaces, but rather combinations of them are abstracted into multidimensional mapping spaces, for example, in the two-dimensional case, in the form of the... FIG. 6 shown rectangles.

[0082] This allows more information to be made available for secondary purposes, such as clinical trials, without compromising k-anonymity.

[0083] Especially in situations with many dimensions, long-term optimization—that is, optimization over multiple time points with increasing database size—is quite difficult to achieve, and it is difficult or even impossible to predict which choice of individual generalization steps will lead to optimal results in the future. On the other hand, the increasingly refined generalization necessitates that an initial mapping must be chosen at the very first time point during the initial generalization, which influences all subsequent mappings. Boundaries of the initial mapping should not be violated in later mappings in order to avoid the limitations imposed by the generalization process. FIG. 2 to avoid the described problem.

[0084] In another example of anonymizing database data, statistical data from one or more additional sources is analyzed. In the case of medical data, these sources could include data from hospitals, a public clinical registry, the Robert Koch Institute, the Federal Statistical Office, preprocessing of the available input data, and so on. Similar sources exist for traffic or mobility data. For medical data, the incidence of the disease or the clinical diagnosis across patients of different ages and depending on body weight, gender, etc., is an important starting point. Similarly, for mobility data, data on the use of specific transport routes—highways or trains—is available.Using these statistical data and an expected number of records that will likely arrive throughout the entire data collection period (or over a long period if data collection has no specific end date), an abstraction target is defined. This abstraction target is a partitioning of the quasi-identifiers into mapping areas using the k-anonymity method. If multiple quasi-identifiers are subjected to anonymization, the mapping areas have a corresponding dimension. For example, with two quasi-identifiers, the mapping areas are rectangles, as shown in [reference]. FIG. 6 This statistical analysis can specify not only the (multidimensional) allocation areas (for example, rectangles) but also the relative sizes of the allocation areas, which are given by the proportion of expected cases in the respective allocation area.

[0085] Since these sources generally do not provide concrete datasets, but only an expected number of datasets, this step can be performed with synthetic (simulated) data using the known distributions. An abstraction goal is a set of relatively small mapping regions (for example, rectangles in the two-dimensional case) that are likely to be reached and, in some cases, actually are. In some cases, however, these can only be approximated because some deviations are necessary, as will be explained later. The mapping regions can be products of intervals of quasi-identifiers, with the same or different widths, as in FIG. 5 , or allocation areas as in FIG. 6However, the assignment ranges should not be too small, as the synthetic data are only a first approximation of the expected data and there will likely be deviations when the real data are used.

[0086] The abstraction goal is therefore an abstraction that should be approximately achieved as the size of the database increases, and which can be achieved if the future data exhibits the expected distribution and the deviation of the future obtained data from the expected data is not too large.

[0087] The abstraction target is used to decide which quasi-identifier (or combination of quasi-identifiers) should be refined. Particularly in multidimensional generalizations with multiple quasi-identifiers, it is possible to perform a refinement on either one quasi-identifier or another during a generalization.

[0088] Without an abstraction goal, however, it is unclear at the time the decision is made which decision will be optimal in the long run. Since subsequent generalizations always represent refinements of the preceding generalizations, an optimal distribution of assignment ranges for a subsequent generalization can be difficult, depending on the values ​​for the next data subjects, if the boundaries of the assignment ranges in a previous generalization are unfavorable. Based on statistical data about the expected number and distribution of future data records and the resulting abstraction goal, generalizations can be appropriately designed, even if the database contains only a few records. For example, the abstraction goal makes it easy to see whether an interval, rectangle, or multidimensional assignment range should be subdivided in one way or another.The abstraction goal should be more refined than the generalization in the next step.

[0089] In one example, an abstraction target is created, as just explained, and then used to make the correct selection from available options for generalizations.

[0090] The abstraction goal guides the selection as follows. The best choice is the one where the abstraction goal is finer than the chosen generalization (or any generalization that satisfies this condition if there are multiple options). There is a simple criterion to determine whether a generalization X is finer than a generalization Y, as shown in FIG. 7As illustrated, a generalization X is finer than a generalization Y if and only if every mapping domain (e.g., every rectangle in the two-dimensional case for two quasi-identifiers q1 and q2) of generalization X is contained in a mapping domain of generalization Y. Fig. 7 1a ⊆ 1, 1b ⊆ 1, 2a ⊆ 2, 2b ⊆ 2, 2c ⊆ 2, 3a ⊆ 3, 3b ⊆ 3, 4a ⊆ 4, 4b ⊆ 4, 5 ⊆ 5, 6a ⊆ 6, 6b ⊆ 6, 7a ⊆ 7, 7b ⊆ 7 (to the left of the subset symbol, the range of generalization X is given, and to the right, the range of generalization Y). Generalization X is therefore finer than generalization Y. If generalization X is finer than generalization Y, then conversely, generalization Y is coarser than generalization X.

[0091] FIG. 8Figure 1 shows an example with two quasi-identifiers. Starting with a generalization Y, the area 2 shown therein is to be subdivided in a further generalization step to provide more precise, yet anonymized, data for a secondary purpose. For example, 4-anonymity is required. Area 2 contains eight data records, each represented by a circle in generalization Y. To maintain 4-anonymity, the area can be further subdivided in various ways, for example, by a vertical division, as shown in generalization V, or by a horizontal division at different heights, as in generalizations Hu, H, and Hl. FIG. 8This is illustrated. Since, as previously discussed, subsequent generalizations should always be chosen more precisely than the current generalization, each of the generalizations V, Hu, H, and Hl represents a constraint for a subsequent generalization. At the point in time when one of the generalizations V, Hu, H, and Hl must be selected, it is often not yet clear which is more or less favorable for subsequent generalizations. By considering an expected distribution of data sets from statistical data from other sources, a likely favorable generalization can be selected. In the example of the FIG. 8 A typical statistical distribution of data sets of the relevant quasi-identifiers is known from other sources, for example, from medical studies from sources of the Robert Koch Institute, from which a target generalization X can be derived, which places area 2 into the FIG. 8The allocation areas shown are subdivided into 2a, 2b and 2c to achieve 4-anonymity. From FIG. 8 It is evident that only via generalization H is a refinement towards the target generalization X possible. Therefore, generalization H should be chosen in the next step.

[0092] If none of the available possible generalizations has the property that the corresponding target generalization is finer than any possible generalization, one can either wait until more data has been added to the database, or choose the generalization that comes closest to the abstraction target, depending on how much deviation from the abstraction target can be tolerated. This deviation from the abstraction target can be a fixed, predefined value or a parameter of the procedure, which can be set, for example, depending on how many data records are expected.

[0093] For example, if, as in FIG. 9 As shown, given a current abstraction Y, an abstraction goal G, and several abstractions C1, C2 that refine Y, and provided none of the several abstractions C1, C2 is coarser than G, the procedure consists of choosing the abstraction such that the deviation from the goal G is minimal. In the example in FIG. 9 Only two possible abstractions C1, C2 are shown. However, the method can easily be extended to any number of possible abstractions C1, C2, ...Cn.

[0094] If the abstraction target G is not finer than an abstraction C of the possible abstractions, the deviation of abstraction C from the abstraction target G can be determined as follows. First, the deviation of each mapping domain (e.g., each rectangle) from C from G is defined as follows.

[0095] For each mapping area RG defined in G, those mapping areas RC defined in C that intersect mapping area RG are considered. If only one mapping area RC intersects mapping area RG, then RG is contained within RC, and the deviation of RG is defined as zero. If multiple mapping areas RC intersect mapping area RG, the magnitudes of these intersections are considered, and the mapping area RC with the largest intersection with RG is selected. The deviation of mapping area RG is the sum of the magnitudes of these intersections, excluding the largest intersection. In other words, the deviation of mapping area RG is the smallest sum of all intersections except one, namely the largest intersection.

[0096] The deviation of C from G is the sum of all deviations of all assignment domains in C from G. Therefore: C is coarser than G (G is finer than C) if and only if the deviation of C from G is zero.

[0097] In another example, the following considerations are taken into account. Typically, the interval sets for a quasi-identifier are a partition of the set of values ​​of interest. As previously discussed in connection with Fig. 1 As described, P1, P2, and P3 are each partitions of the interval [40,100], i.e., the values ​​relevant to age. A partition of a set A is a set of subsets that are mutually disjoint and cover (or exhaust) the set A. In some cases, however, it is possible to use interval sets that do not represent a partition.

[0098] For example, disjoint sets can be used that do not cover all possible values. For instance, a particular disease might occur either in young people (e.g., 16-22 years) or in older people (over 50, but not over 80). Instead of choosing a partition of [16,80[, a set of interesting age groups could be chosen, for example:

[0099] Any data record whose corresponding quasi-identifier contains a value outside these intervals is eliminated (suppressed). This can be particularly useful if the proportion of suppressed data records is not too high or if the attribute does not represent an important value for the secondary purposes.

[0100] In summary, k-anonymity is refined for a set of data records (i.e., for a database, as described above). However, in many situations, the data records are not available from the outset, but rather more are added over time. Since it is of great practical use to be able to use the values ​​in anonymized form even if not all values ​​are yet available, the previously described procedure involves gradually refining abstractions of quasi-identifiers as the amount of data increases, so that k-anonymity is maintained and the data is released in a more precise anonymized form over time, thereby increasing its usability while permanently preserving privacy.

[0101] During refinement, an abstraction target can be considered, based on a statistical analysis of expected quasi-identifier values ​​from other data sources. Stepwise refinements can be applied to approximate the abstraction target, always choosing a k-anonymization that deviates as little as possible from the abstraction target.

[0102] The procedures described above can be carried out automatically by a device, for example a computer system. Fig. 10This illustrates aspects relating to a corresponding device 1000 for anonymizing data by means of generalization. The device 1000 comprises a processing device 1002, e.g., a processor, and a memory 1004. The device 1000 also comprises an interface 1006. Via the interface 1006, the device 1000 can access a database 1050 in which data to be anonymized is stored and anonymized data can be stored. The device 1000 can further comprise a computer-readable storage medium 1008, for example, a hard disk, a read / write memory, or a read-only memory, in which a computer program product, for example, software, is stored. This software product comprises computer-readable program code designed to cause the processing device 1002 to execute the processing steps described below.

[0103] Initially, the data to be anonymized comprises an initial set of first data records, and later, a second set of second data records. The first data records are a subset of the second data records.

[0104] Combined with Fig. 11 The following describes the processing steps of a process 1100, which are carried out by the device 1000, in particular by the processing device 1002.

[0105] The processing device 1002 is configured to generate an initial generalization for the first data records, fulfilling the required anonymization (step 1102). This initial generalization comprises a first group of mapping areas used to generalize the values ​​of a quasi-identifier of the data. The processing device 1002 is further configured to generate a second generalization for the second data records (step 1104), fulfilling the required anonymization. This second generalization comprises a second group of mapping areas used to generalize the values ​​of the quasi-identifier. The second group includes more mapping areas than the first group. Further generalizations can be performed at subsequent times when additional data records become available.The data sets thus generalized can be stored in database 1050 as anonymized data sets or transmitted in another way to another database for secondary purposes (step 1106).

[0106] As previously described, the data can include, for example, patient data. A healthcare provider (such as a hospital) might collect personal data from patients in order to treat them. This data can include personal information such as name, age, address, and the like, as well as health data such as the type of illness, blood test results, X-rays, etc. The healthcare provider can then use the data for secondary purposes, for example, to promote clinical research in anonymized form, such as for clinical trials, or to train a diagnostic system based on artificial intelligence or machine learning. In another example, a healthcare provider in the transportation sector might collect data to, for example, bill parking fees, road tolls, or fares.This data can, in turn, include personal information such as name, address, vehicle make and model, routes traveled, and journey times. In anonymized form, this data can be used for traffic planning or management as a secondary purpose. Another example involves the electronic collection of data in buildings, such as in elevators or at doors. For access control, this data can include personal information such as name, company affiliation, a facial image, and typical movement patterns, e.g., movement paths and associated times. In anonymized form, this data can be used, for example, to optimize elevator controls or to plan traffic flow within buildings.

Claims

1. Computer-implemented method for anonymising data by means of generalisation, wherein the data comprises a first number of first datasets at a first time point and a second number of second datasets at a second time point, wherein the first datasets are a subset of the second datasets, wherein the method comprises: - generating (1102) a first generalisation for the first datasets that fulfils a required anonymisation, wherein the first generalisation comprises a first group of assignment ranges by means of which values of a quasi-identifier of the data are generalised, and - generating (1104) a second generalisation for the second datasets that fulfils the required anonymisation, wherein the second generalisation comprises a second group of assignment ranges by means of which values of the quasi-identifier are generalised, wherein the second group comprises more assignment ranges than the first group, characterised by - determining a statistical distribution of values of the quasi-identifier, - generating a target generalisation on the basis of the statistical distribution, wherein the target generalisation comprises a target group of assignment ranges by means of which values of a quasi-identifier of the data are generalised, - generating the first generalisation and / or the second generalisation additionally as a function of the target generalisation, so that the target generalisation is taken into account when generating the group of the first or the second assignment ranges, in order to avoid that boundaries between assignment ranges are set up which are at least to some extent unsuitable for the second or a third group of assignment ranges, wherein for refinement an abstraction target is taken into account, which is based on statistical data from a statistical analysis of expected values of the quasi-identifiers from other data sources, wherein the abstraction target is specified using the statistical data and an expected number of datasets that will likely arrive during an entire data acquisition period, wherein the abstraction target is a subdivision of the quasi-identifiers into assignment ranges using the k-anonymity method.

2. Computer-implemented method according to claim 1, wherein, when generating the assignment ranges of the first group, for each value of the quasi-identifier it applies that the value is assigned at most to one of the assignment ranges of the first group, wherein the values assigned to the assignment ranges of the first group are the quasi-identifiers of the first datasets, and wherein, when generating the assignment ranges of the second group, for each value of the quasi-identifier it applies that the value is assigned at most to one of the assignment ranges of the second group, wherein the values assigned to the assignment ranges of the second group are the quasi-identifiers of the second datasets.

3. Computer-implemented method according to claim 1 or claim 2, wherein a value set which is assigned to the totality of assignment ranges of the first group is identical to a value set which is assigned to the totality of assignment ranges of the second group.

4. Computer-implemented method according to one of the preceding claims, wherein for each assignment range of the second group it applies that a value set of the respective assignment range is a subset of a value set of precisely one assignment range of the first group.

5. Computer-implemented method according to one of the preceding claims, wherein a respective interval is assigned to a respective assignment range of the first group, wherein the intervals of the first group have interval lengths independent of one another, wherein a respective interval length of a respective interval of the first group is dependent on a number of datasets of the first datasets that are assigned to said interval.

6. Computer-implemented method according to one of the preceding claims, wherein the data at a further time point comprises a further number of further datasets, wherein the further datasets are a superset of preceding datasets at a preceding time point, wherein the method further comprises: generating a further generalisation for the further datasets on the basis of a preceding generalisation, wherein the further generalisation for generalising the quasi-identifier comprises a further group of assignment ranges for the quasi-identifier, wherein the further group comprises more assignment ranges than a preceding group of assignment ranges of the preceding generalisation, wherein the preceding generalisation is the second generalisation or a further generalisation iteratively based on the second generalisation.

7. Computer-implemented method according to one of the preceding claims, wherein each assignment range of the first group comprises an at least two-dimensional assignment range by means of which values of the quasi-identifier and values of at least one further quasi-identifier of the data are generalised, and wherein each assignment range of the second group comprises an at least two-dimensional assignment range by means of which values of the quasi-identifier and values of at least one further quasi-identifier of the data are generalised.

8. Computer-implemented method according to claim 7, further comprising: - determining statistical distributions of values of the quasi-identifier and of values of the at least one further quasi-identifier, - generating a target generalisation on the basis of the statistical distributions, wherein the target generalisation comprises a group of at least two-dimensional assignment ranges by means of which values of the quasi-identifier and of the at least one further quasi-identifier are generalised, - generating the first generalisation and / or the second generalisation additionally as a function of the target generalisation.

9. Device for anonymising data by means of generalisation, wherein the data comprises a first number of first datasets at a first time point and a second number of second datasets at a second (later) time point, wherein the first datasets are a subset of the second datasets, wherein the device (1000) comprises a processing device (1002) which is embodied to perform the steps of the method according to claim 1.

10. Device according to claim 9, wherein the device (1000) is embodied to perform the method according to one of claims 2-8.

11. Computer program product comprising a computer-readable program code which is embodied to cause a processing device (1000) to perform the steps of the method according to one of claims 1-8.

12. Computer-readable storage medium which is embodied to store therein a computer program product comprising a computer-readable program code which is embodied to cause a processing device (1002) to perform the steps of the method according to one of claims 1-8.