Method for pseudonymizing patient data
By distributing the TTP using SMPC and employing symmetric/asymmetric encryption, the method ensures robust security and unique key assignment, addressing the challenges of confidentiality, integrity, and availability in PPRL systems.
Patent Information
- Application Number
- PCT/AT2025/060214
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-06-05
- Filing Date
- 2025-05-28
- Publication Date
- 2025-12-11
AI Technical Summary
Existing Privacy-Preserving Record Linkage (PPRL) solutions face challenges in balancing confidentiality, integrity, and availability, particularly in decentralized systems, where central storage of patient data poses security risks and decentralized systems struggle with long-term data availability, while centralized systems lack robustness against attacks.
Implementing Secure Multiparty Computation (SMPC) methods to distribute the Trusted Third Party (TTP) across multiple entities, ensuring that no single entity has access to the entire dataset, and using symmetric or asymmetric encryption to create and manage pseudonyms across a network of services, preventing unauthorized access and ensuring unique key assignment.
Enhances security by reducing trust assumptions and ensuring that a malfunction or attack on a single service does not compromise all data, while maintaining data linkability and preventing unauthorized access to patient information.
Smart Images

Figure AT2025060214_11122025_PF_FP_ABST
Abstract
Description
[0001] Methods for pseudonymizing patient data
[0002] The invention relates to a method for anonymously creating personal identifiers.
[0003] Background of the invention
[0004] The linking of datasets while preserving privacy, also known as "Privacy-Preserving Record Linkage" (PPRL), presents a challenge at the interface between biomedical research and clinical routine: different datasets must be linked together for use in research without revealing the identities of the individuals concerned.
[0005] Research datasets typically originate from research projects, particularly clinical trials, biobanks that store samples such as blood, tissue, bone marrow, etc., or from registries, etc. These datasets are often pseudonymized to prevent the disclosure of personal information such as name, date of birth, etc., to third parties. In many cases, there are also legal regulations governing the collection of such identifying data (IDAT), for example, the General Data Protection Regulation (GDPR) within the EU.
[0006] In the field of rare diseases, routine patient care is also very closely linked to clinical trials, because study protocols largely take on the role of guidelines. As a result, most patients with rare diseases are enrolled in at least one clinical trial along their treatment pathway, their samples and genomic profiles are stored in at least one biobank, and their data are recorded in at least one registry. The importance of PPRL in pediatric oncology was recently highlighted by Vassal et al.
[0007] To protect patient privacy, each context (study, biobank, registry) generates a pseudonym for each patient, with different contexts generally requiring different pseudonyms to be used according to the GDPR. This situation leads to two serious problems:
[0008] Firstly, if a patient is treated by more than one healthcare provider, it can very easily happen that the same patient is registered more than once in the same context, because the healthcare providers cannot recognize from the pseudonyms that the patient already exists in the context.
[0009] Secondly, datasets from different contexts cannot be linked or related to each other because different pseudonyms are used in the contexts.
[0010] PPRL solves this problem by creating opportunities for linking without compromising patient privacy.
[0011] State of the art
[0012] There are various PPRL solution providers on the market, whose solutions are based on different technological approaches. In 2013, Vatsalan et al. proposed a taxonomy for comparing PPRL solutions, which provides a good overview of 15 dimensions of PPRL. The following discussion focuses specifically on the dimensions "Privacy Technique" and "Number of Parties".
[0013] On the one hand, there are solutions like Findata in Finland, the Health Data Hub in France, or the Mainzeiliste in Germany, which are based on a trusted independent institution (a "Trusted Third Party," or TTP). The TTPs manage the IDAT (International Data Address) of all patients, while the clinical data is stored in the contexts. The TTP provides the pseudonyms to these contexts. Therefore, in such solutions, the TTP can also enable the linking of data from different contexts without itself having access to the sensitive clinical data. Privacy is thus ensured by separating IDAT and clinical data. The advantage of these TTP-based solutions is that the information required for linking is stored centrally, which simplifies the linking process compared to decentralized solutions.However, the central storage of all IDAT data poses a certain security risk that cannot be accepted for all applications.
[0014] On the other hand, there are solutions that apply PPRL directly between partners who know the IDAT, e.g., directly between hospitals, without requiring a TTP. Laud and Pankova, for example, applied such an approach to the EUP ID Services (see below) to avoid TTP. Lazrig et al. and researchers from the Mainzelliste group published approaches based on Bloom filters. Both groups use Secure Multiparty Computation (SMPC) methods for PPRL. A disadvantage of such approaches is that, when applied to pseudonymized contexts (studies, registries, biobanks, etc.), all local data sources (hospitals) must be available for each link, and for the entire duration of the linking process.In contrast, centralized solutions do not require the availability of all local sources (hospitals), but only the availability of pseudonymized data sources (studies, registries, biobanks). While privacy can be well ensured in these approaches, guaranteeing the long-term availability of all data is difficult.
[0015] Finally, there are solutions that compromise security and availability based on the approaches mentioned above by storing only encrypted and / or hashed IDAT data, rather than unencrypted IDAT. The European Patient Identity (EUPID) services were specified in the FP7 EU project European Network for Cancer Research in Children and Adolescents (ENCCA) and are currently used for PPRL in various projects, particularly in the fields of rare diseases and pediatric oncology. Although the current implementation of the EUPID services employs some additional security measures, the PPRL algorithm is fundamentally based on the comparison of hashed IDAT. The Joint Research Centres (JRC) of the European Commission operate the pseudonymization service SPIDER, which was developed based on the EUPID services and hashed IDAT.
[0016] For all solutions based on hashes, Bloom filters, etc., it is true that the TTP (Technology Transfer Process) or an attacker who gains access to the TTP's services can theoretically launch attacks on the derived data (especially dictionary attacks, rainbow table attacks, or brute-force attacks) in which all possible or expected combinations of IDAT are tried and the resulting data is compared with the data in memory. Even though existing services like EUPID or SPIDER have implemented measures to make these attacks more difficult, they cannot be completely prevented.
[0017] Every PPRL solution must therefore find a suitable compromise between confidentiality, integrity, and availability (also known as the CIA Triad) for the respective application. The European Union Agency for Cybersecurity (ENISA) recently published white papers with recommendations regarding pseudonymization services. However, even these recommendations cannot define a universally valid operating point within the CIA Triad. Rather, the choice of the optimal operating point for a specific application is based on data sensitivity, number of users, required response times, consequences in case of errors, etc. Objective of the invention
[0018] The present invention aims to further enhance security, particularly for solutions such as EUPID Services or SPIDER, without sacrificing the advantages of good data linkability. This is achieved by applying SMPC methods to hashed data to reduce the necessary trust assumptions regarding the TTP, for example, by distributing the TTP across multiple entities. In particular, such a solution should ensure that a malfunction of a single central service or a successful attack on a single central service does not lead to the disclosure of all data concerning the subjects.
[0019] Additionally, in such an environment, it is also the task of the invention to ensure that keys are only assigned once for the same person or subject and / or to determine whether a key has already been created for a subject.
[0020] The invention solves this problem in a method of the type mentioned at the outset with the characterizing features of claim 1.
[0021] It is provided that a client, for the purpose of accessing data associated with the identified subject, will: a) create or provide identification data of the subject according to predefined criteria, b) create or provide at least one key, c) if necessary, prepare the identification data for encryption with the key, d) encrypt the identification data or data derived therefrom with the at least one key, and create an encryption result, whereby the encryption method ensures that the identification data or data derived therefrom, if necessary prepared for encryption, can be recovered from the at least one key or a key associated with this key and the encryption result by decryption.and e) transmits the encryption result and, where applicable, at least one key in the form of an account request to a service network comprising at least two separate services, wherein the individual services of the service network: f) do not grant each other access to the keys and / or encryption results transmitted to them and, where applicable, stored by them, g) manage a, in particular distributed, storage for subject accounts in which they store the encryption results and any transmitted keys together with an account identifier in the form of data records by which the jointly transmitted encryption results, and where applicable, also the keys, can be assigned to the respective subject account, and h) upon receiving an account request, jointly check whether a subject account with a data record has already been created for the same subject,whether a subject account already exists that was created on the basis of the same identification data, with the services interacting in such a way that,
[0022] - all services together have sufficient information to determine, based on two data sets alone, whether the identification data used in creating the data are identical,
[0023] - without any of the services disclosing information about the content of the identification data stored by it to the other services, and
[0024] - without either service receiving additional information about the contents of the identification data.
[0025] A preferred variant of the invention, which allows distributed storage across a multitude of services, provides that
[0026] - in step b) at least one key is created for use in a symmetric encryption method, in particular exclusively for one-time use,
[0027] - in step d) the identification data is encrypted with at least one key using a symmetric encryption method, and an encryption result is created, whereby the encryption method ensures that the identification data can be recovered from the at least one key and the encryption result by decryption,
[0028] - in step e) the account request includes at least one key and the encryption result, wherein, upon transmission to the service network, each service receives only one key or the encryption result, but the entirety of the services contained in the service network receives both the encryption result and the entirety of the keys used to create the encryption result.
[0029] - the individual services of the service network do not grant each other access to the keys and / or encryption results transmitted to them and possibly stored by them,
[0030] - the individual services in step g) manage a storage for subject accounts in which they store the keys and / or encryption results transmitted to them together with an account identifier, by means of which the jointly transmitted keys and / or encryption results can be assigned to each other and to the respective subject account, and
[0031] - the individual services in step h) upon receiving an account request for the purpose of determining whether a subject account already exists for the same subject, together with the other services, using SMPC to check whether a data record already exists that is assigned to the same subject account and stored distributed across the services, comprising at least one key and one encryption result,
[0032] - for which the result of the decryption operation based on the keys and / or encryption results assigned to the same subject account corresponds to the result of the decryption operation based on the transmitted keys and / or encryption results.
[0033] Another preferred variant of the invention, which allows the use of an asymmetric encryption method with a public and private key, provides that
[0034] - the client in step b) provides a public key part of an asymmetric key pair ,
[0035] - the client in step d) encrypts the identification data or the data derived therefrom using the public key part according to a homomorphic asymmetric encryption method and thus determines the encryption result ,
[0036] - the private key part of the asymmetric key pair is stored and made available in one of the services, - the account request transmitted in step e) includes the encryption result and preferably no keys,
[0037] - at least two services exist, of which a first service manages the encryption results and their respective assignment to the respective subject accounts, and a second service has access to the private key portion and / or is trained to decrypt the encryption result with this private key,
[0038] - in step h) to determine whether a subject account with a data record has already been created for the same subject ,
[0039] — the first service determines a combination value, in particular a difference or quotient or a value derived therefrom, from the encryption results contained in the respective data record assigned to a subject account and in the account request, and
[0040] — this combination value is transmitted to the second service and the second service decrypts the combination value using the private key part (RPr) with an asymmetric homomorphic decryption method, and
[0041] — it is checked whether a predefined standard value is obtained after decryption, which indicates that the identification data contained and encrypted in the respective data record and in the account request, or the data derived from it, are identical.
[0042] To ensure that keys are assigned only once for the same person or subject and that no duplicate entries are created, it may be provided that, in the event that no matching subject account was found in step h), j), a new subject account is created and the keys and / or the encryption result transmitted to the individual services are assigned to the subject account and stored by the services.
[0043] In order to obtain the necessary access to data linked to the key, it may be necessary that
[0044] - in at least one of the services, the following access information is assigned to the subject accounts:
[0045] 1) Pseudonyms used for the subject in question,
[0046] 2) other services (contexts) in which data relating to the subject is stored, in particular specifying the respective pseudonym,
[0047] - that after the identification of a subject account corresponding to the identification data in step i) or after the creation of a new subject account in step j), the access information assigned to the respective subject account is at least partially transmitted to the client.
[0048] To prevent the individual central services that jointly manage the pseudonyms from learning additional information in the context of requests for personal data, it may be provided that in step c) when preparing the identification data for encryption, a hashing procedure is applied to the identification data and a hash value is determined.
[0049] - that the hash value preferably has the same length as the at least one key used for encryption, and
[0050] - that the hash value is used for encryption.
[0051] This effectively prevents such information from being used by any of the services to compromise the protection of the data jointly managed by the services. Special handling of various input errors, particularly those resulting from incorrect name entries, provides that
[0052] - in step c) identification data of a subject are preprocessed separately using several preprocessing methods, in particular hashing methods and / or phonetic hashing methods, so that a number of derived values are obtained,
[0053] - that steps d) to e) are executed separately for the values derived in this way, and in step h) a separate determination is made for each transferred derived value as to whether these derived values correspond to the derived values already assigned to a subject account, and the determination of whether a subject account with a data record has already been created for the same subject is made based on these individual determinations, whereby, in particular, depending on how many and / or which of the various individual determinations identify existing subject accounts, different further processing steps are carried out, such as queries to the client, determination of whether it is the same subject, or creation of a new subject account.
[0054] Figure description
[0055] Figure 1 shows an overall architecture of a network for information exchange. Figure 2 schematically shows a specific procedure for comparing whether a particular entry already exists in the system, which, however, is carried out in a distributed manner within the scope of the invention. A preferred first embodiment (Figure 1) of the invention is based on the distribution of hashed information from the EUPID services to two separate servers, EUPID 1 and EUPID 2. Neither service has access to the data of the other. Figure 1 shows an overall architecture for carrying out a first embodiment of a method according to the invention for managing patient information across multiple clients (Hospital 1-n). These patients, also referred to as subjects S within the scope of the invention, are identified by identification data (IDAT) I.This identification data I, or data derived from it, is then used to identify data concerning the subject, which is stored in different contexts (respectively, contexts).
[0056] In a first step a), identification data I of the subject is created according to predefined criteria. For example, the subject's name, date of birth, insurance number, and other personal data can be used as identification data I. This data is then, for example, concatenated into a string, whereby normalization steps can be performed (e.g., taking special characters, different alphabets (Cyrillic, etc.), etc.), and the string is made available as an identification data structure.
[0057] In a second step b), a key or random value R is generated for use in a symmetric encryption method. Common methods for generating random numbers or keys are suitable for this purpose.
[0058] In the present particular embodiment of the invention, the identification data are prepared for encryption in a third step c) by applying a hashing method to the identification data structure and thereby determining a hash value of fixed length H. Commonly available hashing methods can be used to generate the hash value H. Preferably, collision-resistant hashing methods can be used, which are designed to avoid identical hashes for different identification data. The hash value H preferably has the same length as the key R used for encryption. Preferably, to increase the security of the data transmission, a separate key is generated for each query, i.e., each key R is used only once to encrypt identification data I.
[0059] In a fourth step d), the hash values H generated from the identification data I are encrypted using at least one key R in a symmetric encryption method. An encryption result S is created, whereby the encryption method ensures that the identification data I, or the hash values H generated from the identification data I, can be recovered from the at least one key R and the encryption result S by decryption.
[0060] As part of the encryption process, an encryption result - hereinafter also referred to as Secret S - is created or provided, and the random value R is generated.
[0061] In one exemplary embodiment of the invention, a user (e.g., a physician / study assistant) starts a web browser on their local computer and opens a specific web page to register a patient in a particular context and generate a pseudonym for that context. There, they enter the first name, last name, and date of birth of the patient to be registered. On the computer, this data is combined (e.g., using JavaScript) into a text (e.g., as...). <vorname> | <nachname> | <geburtsdatum>) and a hash Hl is calculated from the text. Furthermore, a random value RI of the same length is generated. Then a secret S is calculated as a logical bitwise XOR of the two values: S1 = Hl © RI .
[0062] In this setting, RI acts as a symmetric key used to encrypt Hl and to decrypt the encrypted secret S1. Alternatively, other encryption methods can be used.
[0063] In a fifth step e), the two generated values, namely the encryption result or . Secret S and the key R, are transmitted to a service network in the form of an account request A.
[0064] Upon receiving an account request A, the service network should now determine whether a subject account K with a data record already exists for the same subject. For this purpose, all services of the service network cooperatively check whether a subject account K already exists that was created based on the same identification data I.
[0065] In the event that no subject account K has yet been created for a subject specified in an account request A, a new subject account K can preferably be created for the subject in question.
[0066] If a matching subject account K is found according to the criteria described below, the client can be granted access to further information. Access to the data associated with subject account K, or the creation of a subject account K, will be described in more detail later.
[0067] In this case, the service network comprises two separate services, EUPID1 and EUPID2. In this embodiment, the two services have separate storage for this data, with the first service, EUPID1, storing the encryption result or secret S, and the second service storing the random number or key R. Each service, EUPID1 and EUPID2, has access only to the data or portions of the account request A transmitted to it. The two services, EUPID1 and EUPID2, are configured such that they do not grant each other access to the keys or random numbers R and / or encryption results or secrets S transmitted to and potentially stored within them, nor to any data derived from the key R and / or encryption results S.
[0068] The data R and S created during the creation or definition of the subject account K are transferred to the two services EUIPD1 and EUPID2 in such a way that each service receives only one of the two data sets and the other data sets are kept secret from that service.
[0069] S1 is sent to EUPID 1, RI is sent to EUPID 2. The original hash Hl remains hidden from both EUPID servers. Only together can the two servers recalculate the hash as Hl = S1 © RI. This calculation between the servers is performed exclusively using a secure protocol, such as an SMPC protocol, ensuring that neither server receives any information from the other during the calculation. Therefore, the hash is never revealed to either server, and the patient's IDAT always remains hidden.
[0070] In this way, a distributed storage area for subject accounts K is created among the services, in which the encryption results or secrets S and the keys or random numbers R are stored together with an account identifier in the form of data records. Only if both the encryption result or secret S and the random number R are known is a correct assignment to the respective subject account K possible; that is, none of the services can determine on its own which subject account K an account request should be assigned to.
[0071] The services work together in such a way that all services together have sufficient information to determine, based on two data sets, whether the identification data I used in creating the data are identical.
[0072] - without either service EUPID1 or EUPID2 disclosing information about the parts of the data records stored with it to the other services, and
[0073] - without either service receiving additional information about the contents of the identification data I .
[0074] In this preferred embodiment, upon receiving an account query, the two services check whether they already hold an entry belonging to the same subject or patient by comparing all previously stored values of S and R with the new value using Secure Multiparty Computation (SMPC), without either service having to disclose EUPID1 or EUPID2 information to the other server. The generation of pseudonyms, which are assigned to the contexts and subject accounts, occurs after the account query.
[0075] No match can be found for the very first patient. Therefore, if services EUPID1 and EUPID2 do not yet contain data from subject accounts K, a new data record is created and a subject account K is generated. EUPID1 stores the secret S1 transmitted to it and generates a unique identification number EUPID 1, which it uses only internally and does not transmit to anyone else, as well as a new pseudonym PSN1 for the context chosen by the user, which is returned to the user and displayed. EUPID 2 stores RI.
[0076] If subject accounts already exist in EUPID1 and EUPID2, SMPC is used to calculate for each of these values Si and Ri whether RI © S1 equals Ri © Si, i.e., whether the hashes Hl and Hi, and therefore also the IDAT of patients 1 and i, are identical. The two servers use, for example, the Yao protocol for this purpose.
[0077] In binary, for the comparison between two values RI © SI = R2 © S2, for each bit i, Rli © Sli == R2i © S2i must hold. This comparison for all bits i can be implemented, for example, in a logic circuit like the one in Figure 2, which is known to both EUPID 1 and EUPID 2.
[0078] Fig. 2 shows an example of a logic circuit for comparing all bits i of RI (R1 l ... Rin) and SI ( Si l ... Sin) as well as R2 (R21 ... R2n) and S2 ( S21...S2n) , so that in the end the equation RI © S1 = R2 © S2 is solved .
[0079] According to the Yao protocol, the EUPID1 service calculates a truth table for all gates of the circuit and then "garbled" these tables (a "garbled circuit"). RSA encryption can be used to encrypt the table entries. The garbled circuit and the garbled input values S1 and S2 are then transferred from EUPID1 to EUPID2. EUPID2 then requests the garbled values for RI and R2 from EUPID1 using oblivious transfer, without having to disclose the actual values of RI and R2 to EUPID1. Now EUPID2 has all the information to traverse the garbled circuit with all bits of RI, R2, S1, and S2 and calculate the garbled value for RI ≤ SI == R2 ≤ S2, i.e., to determine the garbled value for H1 ≤ H2. Finally, the Garbled result is passed to EUPID1, which then determines whether this value corresponds to the result Hl == H2 or Hl != H2.
[0080] If a match is found, the procedure can be terminated. Otherwise, it continues with the next stored value S3 and R3 until either a match is found or all values have been processed.
[0081] In the present embodiment, the first service EUPID1 also includes an additional memory in which further data associated with a subject account K can be stored. If an account query A has successfully identified a subject account K, the requesting client can access this additional memory. In principle, any data can be stored in this memory. For example, the following information can be associated with the subject accounts in the first service EUPID1.
[0082] 1) Pseudonyms used for the subject in question in different contexts,
[0083] 2) Further data concerning the subject. In particular, different pseudonyms can be stored for different contexts.
[0084] If the same subject or patient participated in multiple medical studies, a different pseudonym was assigned to each study. To access the data from each study, the pseudonym and the context—that is, information about the respective study—are stored within the data.
[0085] Pseudonyms are stored together with the subject account K, which are associated with further data concerning the subject in another service. For example, several pieces of medical information concerning the same subject can be stored together in such a way that different third-party keys or pseudonyms are assigned to the subject account K, which can then be used to retrieve data concerning the subject specified in the subject account K from other services. After determining or identifying a subject account K corresponding to the identification data I (step i) or after creating a new subject account K (step j), the data associated with the respective account are...
[0086] The access information associated with the subject account K is at least partially transmitted to the client that made the account request A.
[0087] For the aforementioned first embodiment of a method according to the invention, it is also readily possible to use a larger number of services in the service network. Instead of a single key, several keys are generated client-side, all of which are subjected to an encryption process together with each other and with the identification data. In this case, one obtains an encryption result S and several keys RI, ..., Rn. The account query A transmitted to the service network contains the encryption result S and the used keys RI, ..., Rn. The encryption result S and the used keys RI, ..., Rn are stored in separate services, whereby n+1 services are available when using n keys. Such methods are described, for example, in Mohassel / Rosulek / Zhang, Fast and Secure Three-party Computation: The Garbled Circuit Approach, https: / / dl.acm.org / doi / 10.1145 / 2810103 . 2813705 described .
[0088] A second embodiment of the invention is described in more detail below. The units involved and the general purpose of the process correspond to the first embodiment of the invention, with only the deviations being described in more detail here, whereby some of the steps shown are carried out differently.
[0089] Instead of symmetric encryption, this embodiment of the invention uses asymmetric encryption. A key pair is generated beforehand (step b), wherein one key of the key pair is known to all computers as the public key RPu, and the other key of the key pair is stored only in the second service EUPID2 of the service network as the private key RPr. This second service EUPID2 has the sole task of decrypting the data transmitted to it using the asymmetric encryption method and the private key RPr, and of evaluating the decryption results.
[0090] After the data has been entered, in the fourth step (step d) the client encrypts the identification data I using the public key part RPu according to a homomorphic asymmetric encryption method. In this way, an encryption result S is determined.
[0091] The encryption method is preferably designed such that when the same identification data I is encrypted multiple times, different encryption results S always result, but all of them yield the original identification data I when decrypted with the private key part RPr.
[0092] The account request A transmitted in the fifth step (step e) includes the encryption result S and preferably no keys.
[0093] As in the first embodiment of the invention, the service network also has two services EUPID1 and EUPID2, of which, as already mentioned, only one service EUPID2 has access to the private key part RPr.
[0094] The other – first – service EUPID1 does not have access to the private key part RPr. The first service EUPID1 stores the encryption results S and manages their respective assignment to the respective subject accounts. This service EUPID1 also allows the storage of data or the assignment of individual data points to the relevant encryption result S of the subject account K. In step h), it is determined according to the following specifications whether a subject account K with a data record already exists for the same subject.
[0095] The first service EUPID1, in which all previous encryption results of all previous account requests are stored and to which the new account request A or its encryption result S is also sent, now examines according to the following criterion whether the encryption result originates from the same subject as an encryption result that has already been transmitted and assigned to a subject account K.
[0096] For this purpose, the first service, EUPID1, determines a combination value from the respective encryption result Si assigned to a subject account Ki and the encryption result S contained in the account request A. Depending on the type of homomorphism (additive or multiplicative homomorphism) of the encryption method in question, this can be a difference or a quotient. Since multiple encryptions of the same identification data always yield different results, this combination value of Si and S does not indicate whether Si equals S. Subsequently, a rerandomization is performed (e.g., by multiplication with an encrypted random value in the case of an additive homomorphic method).
[0097] This encrypted combination value, i.e., this difference (Si-S) or this quotient (Si / S), is then transmitted from the first service EUPID1 to the second service EUPID2. The second service EUPID2 then decrypts the combination value using the private key part RPr and the asymmetric homomorphic decryption method. Due to the homomorphism of the decryption method, the contributions of the encryption result Si assigned to the respective subject account K and the last transmitted encryption result S cancel each other out, so that after decryption a predetermined default or neutral value is obtained (zero for difference calculation, 1 for quotient calculation). The occurrence of this value indicates that the encrypted identification data I contained in the respective data record and in the account query are identical.
[0098] However, if encryption results S are combined that do not originate from the same identification data I, the respective contributions do not cancel each other out, which can be determined by the second service EUPID2. Due to the mutual overlap of the encryption results S, however, it is not possible for the second service to deduce the two encryption results S. Due to the aforementioned rerandomization, no information about the difference / quotient is revealed to the second service EUPID2.
[0099] These measures ensure that neither service EUPID1 nor EUPID2 is capable of assigning the transmitted subject request A to a subject account K on its own.
[0100] For each subject account, several types of hashing or derivations from the identification data I can also be used, which will yield identical values in the case of a partial match of the specified identification data I.
[0101] For example, using a phonetic hash in addition to the phonetic hash makes it possible to identify spelling errors in names, resulting in a match, whereas using a collision-resistant hash will produce different hash values. If a match of phonetic hashes is detected, for example, because the surname "Maier" was incorrectly entered as "Mayer," while the collision-resistant hashes do not match, it is possible to inform the user that there may be an input error in the identification data. Another type of hash yields the same results if dates or numbers have been transposed; for example, the day and month numbers are regularly swapped in patients' birth dates.
[0102] The method described according to the invention can also be applied to these additional hashes.
[0103] If multiple such hashes are used in a single request, they are encrypted separately and transmitted to the service separately. The subject accounts K each contain entries for the differently generated hash values or derived values. When checking whether a subject account K already exists, the individual hash values or derived values generated in the same way are compared.
[0104] If all of the transmitted hash values or derived values correspond to the respective hash values or derived values assigned to a subject account K, a match can be established. However, if all of the transmitted hash values or derived values differ from the respective hash values or derived values assigned to a subject account K, then no match can be established. If, however, only some of the transmitted hash values or derived values are the same, a query can be initiated with the client.
[0105] For example, if a collision-resistant hash shows no match, but a phonetic hash does, it can be assumed that the subject's name was misspelled. The client will be prompted to correct it.< / geburtsdatum> < / nachname> < / vorname>
Claims
Patent claims:
1. A method for managing information relating to a subject identified by identification data (I), wherein a client (C) for the purpose of accessing data associated with the identified subject: a) creates or provides identification data (I) of the subject according to predefined criteria, b) creates or provides at least one key (R), c) optionally prepares the identification data (I) for encryption with the key (R), d) encrypts the identification data (I) or data derived therefrom with the at least one key (R), and produces an encryption result (S), wherein the encryption method ensures that the at least one key (R) or a key (RPr) associated with that key and the encryption result (S) can be decrypted to reveal the data that may have been prepared for encryption.Identification data (I) or data derived therefrom are recoverable, and e) transmits the encryption result (S) and, where applicable, at least one key (R) in the form of an account request (A) to a service network comprising at least two separate services, wherein the individual services of the service network: f) do not grant each other access to the keys (R) and / or encryption results (S) transmitted to them and, where applicable, stored by them, g) manage a, in particular distributed, storage for subject accounts (K) in which they store the encryption results (S) and, where applicable, transmitted keys (R) together with an account identifier in the form of data records, about which the jointly, transmitted encryption results (S), and possibly also the keys (R), are attributable to the respective subject account (K), and h) upon receiving an account request (A) for the purpose of determining whether a subject account (K) with a data record already exists for the same subject, jointly check whether a subject account (K) already exists that was created on the basis of the same identification data (I), the services cooperating in such a way that - all services together have sufficient information to determine, based on two data sets alone, whether the identification data (I) used in creating the data are identical, - without any of the services disclosing information about the contents of the identification data (I) stored by it to the other services, and - without either service receiving additional information about the contents of the identification data (I).
2. Method according to claim 1, characterized in that - in step b) at least one key (R) is created for use in a symmetric encryption method, in particular for one-time use only, - in step d) the identification data ( I ) are encrypted with at least one key ( R) using a symmetric encryption method, and an encryption result ( S ) is created, whereby the encryption method ensures that the identification data ( I ) can be recovered from the at least one key ( R) and the encryption result ( S ) by decryption, - in step e) the account request (A) includes at least one key (R) and the encryption result (S), wherein in its Transmission to the service network: each service receives only one key (R) or the encryption result (S), but the entirety of the services contained in the service network receives both the encryption result (S) and the entirety of the keys (R) used to create the encryption result (S). - the individual services of the service network do not grant each other access to the keys (R) and / or encryption results (S) transmitted to them and possibly stored by them, - the individual services in step g) manage a storage for subject accounts in which they store the keys (R) and / or encryption results (S) transmitted to them together with an account identifier by which the jointly transmitted keys (R) and / or encryption results (S) can be assigned to each other and to the respective subject account (K), and - the individual services in step h) upon receiving an account request (A) for the purpose of determining whether a subject account (K) has already been created for the same subject, together with the other services, check using SMPC whether a data record already exists that is assigned to the same subject account (K) and stored distributed across the services, comprising at least one key (R) and one encryption result (S). - for which the result of the decryption operation based on the keys (R) and / or encryption results (S) assigned to the same subject account (K) corresponds to the result of the decryption operation based on the transmitted keys (R) and / or encryption results (S).
3. Method according to claim 1, characterized in that - the client in step b) provides a public key part (RPu) of an asymmetric key pair , - the client in step d) encrypts the identification data ( I ) or the data derived therefrom using the public key part (RPu) according to a homomorphic asymmetric encryption method and thus determines the encryption result ( S ). - the private key part (RPr) of the asymmetric key pair is stored and made available in one of the services, - the account request (A) transmitted in step e) includes the encryption result (S) and preferably no keys , - at least two services exist, of which a first service manages the encryption results (S) and their respective assignment to the respective subject accounts, and a second service has access to the private key part (RPr) and / or is trained to decrypt the encryption result (S) with this private key, - in step h) to determine whether a subject account (K) with a data record has already been created for the same subject , - the first service determines a combination value, in particular a difference or a quotient or a value derived therefrom, from the encryption results (S) in the respective data record assigned to a subject account (K) and in the account request (A), and - this combination value is transmitted to the second service and the second service decrypts the combination value using the private key part (RPr) with an asymmetric homomorphic decryption method, and - it is checked whether a predefined standard value is obtained after decryption, which indicates that the identification data (I) contained and encrypted in the respective data record and in the account request, or the data derived from it, are identical.
4. Method according to one of the preceding claims, characterized in that, in the event that no matching subject account (K) was found in step h), j) a new subject account (K) is created and the keys (R) and / or the encryption result (S) transmitted to the individual services are assigned to the subject account (K) and stored by the services.
5. Method according to one of the preceding claims, characterized in that , - that in at least one of the services the following access information is assigned to the subject accounts: 1) Pseudonyms used for the subject in question, 2) other services (contexts) in which data relating to the subject is stored, in particular specifying the respective pseudonym, - that after the identification of a subject account (K) corresponding to the identification data ( I ) in step i ) or after the creation of a new subject account (K) in step j ) the access information assigned to the respective subject account (K) is at least partially transmitted to the client (C .
6. Method according to one of the preceding claims, characterized in that , - that in step c) during the preparation of the identification data ( I ) for encryption, a hash procedure is applied to the identification data ( I ) and a hash value (H) is determined, - that the hash value (H) preferably has the same length as the at least one key used for encryption, and that the hash value (H) is used for encryption.
7. Method according to one of the preceding claims, characterized in that , - that in step c) identification data (I ) of a subject are preprocessed separately using several preprocessing methods, in particular hashing methods and / or phonetic hashing methods, so that a number of derived values are obtained, - that steps d) to e) are executed separately for the derived values in this way, and in step h) a separate determination is made for each transferred derived value as to whether these derived values correspond to the derived values already assigned to a subject account (K), and the determination of whether a subject account (K) with a data record already exists for the same subject is made based on these individual determinations, whereby, in particular, depending on how many and / or which of the various individual determinations identify existing subject accounts (K), different further processing steps are carried out, such as queries to the client, determination of whether it is the same subject, or Creating a new subject account (K) .