Method and system for re-associating anonymized data with data owner

By generating and managing global confidential identifiers for data owners, technical means to securely reassociate anonymous data with data owners in fields such as healthcare and scientific research, solve the problem of insufficient data privacy and security in the prior art, and realize selective identification and re-identification of data.

CN120153372APending Publication Date: 2025-06-13MEDISSEDA DAUS DE SAUD AG
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202380077056.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2022-11-14
Filing Date
2023-11-10
Publication Date
2025-06-13

AI Technical Summary

Technical Problem

The prior art is difficult to securely reassociate anonymous data with data owners while ensuring data privacy, especially in areas such as healthcare and scientific research.

Method used

De-identification and re-identification of data is achieved by generating and managing the global confidential identifier of the data owner, including three password keys k1, k2 and k3. The data owner uses personal software applications to send requests to the data provider server, performs de-identification of the data, and copies the anonymous data to the service provider server. Meanwhile, the name and personal identifier of the data owner are transmitted to the user identification data server for later re-identification.

Benefits of technology

The ability to selectively identify and re-identify data owners and their personal data while meeting data privacy and security requirements is realized, solving the technical difficulties of re-associating data with data owners.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120153372A_ABST
    Figure CN120153372A_ABST
Patent Text Reader

Abstract

A computer-implemented method for re-associating anonymized data with a data owner is described, wherein the data owner has an associated personal code. The method comprises the steps of accessing, by a third-party computer, anonymized data stored in a service provider computer server and transmitting a first form of personal code from the third-party computer to the service provider computer server. The method further includes matching, at the service provider computer server, the personal code in the first form with a personal code in a second form, and transmitting the personal code in the second form from the service provider computer server to a user identification data computer server. The method further includes matching, by the user identification data computer server, the personal code in a second form with the data owner identifier.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application claims priority to Portuguese Patent Application No. PT118342, filed on November 14, 2022. The entire disclosure of Portuguese Patent Application No. PT118342 is hereby incorporated by reference herein.

[0002] The field of the present invention relates to a method and system for re-associating anonymized data with a data owner. The present invention belongs to the fields of computer systems and cryptography and coordinates the conflicting requirements for securely and selectively identifying, de-identifying, and re-identifying a previously de-identified data owner and their personal data, based on the identity and function of the data recipient. Background Art

[0003] The draft EU regulation on the European Health Data Space, published on May 3, 2022, describes strict conditions for protecting personal health data and the privacy rights of citizens, based on data de-identification and other measures. Data de-identification includes: anonymization, i.e., the data is irreversibly anonymized and cannot be traced back to the data owner; and pseudonymization, i.e., the data is anonymous to those who do not need to know the identity of the data owner, but if the identity of the data owner needs to be known or the data owner needs to be contacted, the data can subsequently be re-identified and re-associated with the data owner.

[0004] The proposed regulation indicates two different uses of data: primary use, i.e., the data is health data used for treating the diseases and health of the data owner; or secondary use, which includes all other uses, including creating new knowledge by processing health data from many different data owners on a large scale.

[0005] Methods for implementing the identification and de-identification of data owners are known. The prior art describes methods for identifying and de-identifying data, namely WO 2020 / 221778 and WO 2020 / 165174.

[0006] It is necessary to make the de-identification process so secure as to avoid identity leakage between the data subject or data owner or patient, the data provider that stores and shares the personal data of the data owner, the service provider that receives, stores, manages and shares the personal data of the data owner, the third parties (such as researchers and data experts) that receive the personal data of the data owner for further processing, and the user who can re-identify the data owner to identify the data processor. The method must be very secure so that even if there is collusion between the members of the data loop, it is impossible for those individuals or computers that do not need to know the identity of the data owner to re-identify the previously de-identified data. Re-identification of such previously de-identified data is possible if and only if a) there is a legitimate reason for such re-identification of the data, and the legitimate reason can be legal or moral, or in line with the consent or request expressed by the data owner; b) the re-identification of the data is only carried out for the authorized recipients of the re-identified data, and the re-identification method will continue to hide the identity of the data owner from all others who do not need to know it.

[0007] However, the prior art does not disclose a system or method for re-associating previously anonymized data with the data owner. Summary of the Invention

[0008] The present invention solves the technical problem that de-identification and re-identification are essentially conflicting requirements. One of the most effective protection measures when dealing with personal data is to de-identify the data so that unauthorized third parties (such as computer technicians or researchers who are interested in the data and do not need to know the identity of the data owner) are prevented from accessing their identity details and their identified data. However, the health data transmitted to healthcare professionals during consultations with patients must clearly indicate the patient's name and date of birth so that healthcare professionals can ensure that the patient is treated according to their own health data rather than that of other patients. Therefore, it is necessary to re-identify the patient data or re-associate the patient data with the owner's name, and take the necessary precautions to prevent the identified health data from being leaked to unauthorized third parties or individuals.

[0009] In another scenario, health data must be de-identified to protect the identity of the owner, i.e., the patient's health data is included in a large population dataset and is sent to a third party (such as a researcher) for scientific research purposes. The researcher must not have any information about the patient's name or other characteristics that could identify the patient. However, in some cases, such as when the researcher discovers new and important information about the patient's health, diagnosis, or prescription through accessing and processing the dataset and this information must be communicated to the patient's healthcare professional for action, there must be a way to re-identify the patient involved. This disclosure reconciles these conflicting requirements.

[0010] The uses of the present invention are not limited to health data. The uses can include all data systems that store and process personal and personally sensitive data and require confirmation that the data truly belongs to the data owner and can be identified, de-identified, and re-identified. Voting and elections are an example, where there are several requirements regarding election data: a) registering citizens or data owners to vote electronically; b) verifying the identity and voting eligibility of citizens; c) enabling electronic means for citizens to vote on election day; d) properly recording that the identified citizen has voted; and e) properly recording the citizen's voting choice, without any association with the citizen's name, so that the vote is as secret as a paper ballot.

[0011] Other uses include all cases where data belonging to sensitive categories need to be processed, such as racial or ethnic origin, political opinions, religious or philosophical beliefs, trade union membership, genetic data or biometric data used to uniquely identify a natural person or data owner, data about the sex life and sexual orientation of a natural person, and data for which people have a legal right or privacy expectation, such as money, banking, taxation, assets, property records, and even in cases where a person who has joined a social network wants to be known by friends and is willing to share data with them, but also wants to be completely anonymous and able to prevent sharing of their data outside the group.

[0012] The described technical problem is solved by a method and system for re-associating anonymized data with the data owner, where the data of interest initially resides in a data provider computer server operated by a government, enterprise, or healthcare organization. The method includes: the data owner uses a personal software application on a personal computer device, which sends a request to the data provider computer server, instructing the de-identification of the data of interest to the data owner and its regular copying to a service provider computer server selected by the data owner. In a healthcare scenario, the health data of citizens may be scattered and stored on a large number of hospital servers, making access and data sharing difficult or impossible. To solve this problem, the European Health Data Space Regulation gives citizens the right to designate a service provider who will manage their health data and enable citizens, healthcare professionals, and scientific researchers to easily access and share the data. In a non-healthcare scenario, the service provider can be an organization responsible for managing the personal data of citizens. At the same time, for ease of future re-identification, the name, other personal identifiers, and contact data of the data owner are transmitted to a separate user identification data computer server. Then, the de-identified data of interest can be transmitted by the service provider computer server to a third-party computer, where the data of interest can be specifically processed to generate additional new information. In some cases, it may be necessary to re-associate this additional information with the name and personal identifiers of the original data owner. This re-association is carried out through the coordinated operation of the computing devices, servers, and computers described in this specification.

[0013] According to a first aspect of the invention, a computer-implemented method for re-associating anonymized data with the data owner is described, where the data owner is associated with three cryptographic keys k1, k2, and k3, collectively referred to as the global secret identifier of the data owner.

[0014] These keys are generated by the personal software application upon first installation. Key k1 always remains in the personal software application, key k2 is transmitted to the service provider computer server, and key k3 is transmitted to the data provider computer server, the user identification data computer server, and the service provider computer server, and is subsequently transmitted by the service provider computer server to the third-party computer. Key k3 will be the personal code of the data owner, which is a pseudonym that can identify the data of interest to the user in the absence of the personal name and personal identification details and allows subsequent re-identification.

[0015] The method includes the following steps: receiving, by a data provider computer server, a request from a personal software application to search, obtain, de-identify, and periodically transmit to a service provider computer server data of interest to a data owner, where the data is uniquely identified by a key k3.

[0016] The method further includes: receiving, by a service provider computer server, the anonymized data of interest from the data provider computer server, where the data is uniquely identified by a key k3. The method further includes the following steps: receiving, by a third-party computer, the anonymized data of interest from the data provider computer server, where the data is uniquely identified by a key k3, and where data processing steps are performed on the data to produce a useful result.

[0017] The method further includes the following steps: receiving, by a user identification data computer server, a request from the third-party computer to re-identify the anonymized useful result, information of the anonymized data owner, and records of interest, where they are uniquely identified by a key k3. It further includes the following steps: receiving, by the personal software application of the data owner and by the computer system of an authorized party having a legitimate need and reason to know, the useful result regarding the data owner and their records of interest and information, this time for re-identification using the name and personal identification details of the data owner.

[0018] The data owner identifier can be at least one of the name of the data owner and another personal identifier.

[0019] The step of accessing the anonymized data can further include the step of obtaining a requirement to associate the anonymized data with the identity of the data owner.

[0020] The key k1 can be a private asymmetric cryptographic key of the data owner. The key k2 can be a public asymmetric cryptographic key of the data owner. The key k3 can be a hash function of the public asymmetric cryptographic key k2, or can be a randomly generated number recorded in a computer file associated with the key k2.

[0021] This document further describes a system for re-associating anonymized data with a data owner. The system includes a service provider computer server for storing the anonymized data, matching the key k3 with the key k2, and transmitting the key k3 to a user identification data computer server. The system further includes a third-party computer for accessing the anonymized data and transmitting the key k3 to the user identification data computer server to match the key k3 with a data owner identifier (such as a name and / or other personal identifier).

[0022] The system may further include a data provider computer server for transmitting anonymized data of interest to the service provider computer server.

[0023] The system may further include a personal computing device of the data owner for generating keys k1, k2, and k3 of a global confidential identifier of the data owner. Keys k2 and k3 are transmitted to the service provider computer server, and key k3 is subsequently transmitted to a computer of a third party, where, without any record of a name or other personal identifier, key k3 anonymously and confidentially identifies the personal data of the data owner. Key k3, as well as the name and personal identifier of the data owner, are also transmitted to the data provider computer server, where they are used to search for, obtain, and confidentially identify the data of interest of the data owner, and are transmitted to the user identification data computer server, where they will be stored for the purpose of re-identifying the data at a later time. Thus, the service provider computer server is the only server storing keys k2 and k3 and the data of interest, while the user identification data computer server is the only server storing key k3, the name, personal identifier, and contact information of the data owner, but not storing other data of interest. This separation of data is key to ensuring the confidentiality of the privacy-by-design, and allows for the re-identification of the data owner when needed.

[0024] The method may be used to store at least one of health data and election data.

[0025] A computer program is also described, which includes instructions that, when executed by a computer, cause the computer to perform the method.

[0026] A computer-readable medium is also described, which includes instructions that, when executed by a computer, cause the computer to perform the method.

[0027] The systems and methods described in this document provide selective reversible and irreversible anonymization. The system processes data of the data owner, also known as citizen, data subject, user, or patient in the health scenario. It is the data owner who controls access to their personal data, requests the transfer of data to parties storing the personal data of the data owner, with the aim of transmitting a copy to a service provider chosen by the data owner, or processing the personal data for purposes consented to by the data owner. Whenever the law defines conditions for the processing of personal data, it is desired to confirm the identity of the data owner via a data owner identifier, such as their name and / or other available personal identifiers, such as date of birth, address, postal code, email address, phone number, citizen number, social security number, national insurance number, tax identification number, etc., which are contained in official documents or from trusted sources. Other personal identifiers can include username, nickname, financial account number, names of employers and insurance companies.

[0028] The personal software application running on the data owner's personal computing device is managed by the data owner and is used to confirm the identification data of the data owner, as well as record the data owner's authorizations, preferences, and data transfer requests. The personal software application is where the data owner specifies which types of personal data can be shared, with whom, and for how long; or who can process the personal data and for what purposes. This is the data owner's expression of consent, and according to the law, withdrawing consent is as easy as granting it by the data owner, and the personal software application must support the granting and withdrawal of consent.

[0029] When the personal software application is first installed on the data owner's personal computing device, the personal software application containing the cryptographic software module generates cryptographic keys k1, k2, and k3, which not only identify and encrypt the identity of the data owner, but also are used to uniquely and confidentially identify the data owner in the absence of a name and other personal identifiers. These cryptographic keys can be used for subsequent anonymous identification and for re-identifying personal data or re-associating personal data with the name of its owner when needed.

[0030] A user identification data computer server that is operated (or under the control of) a public entity in one aspect stores the name and / or other personal identifiers of the data owner, the personal code k3 of the data owner, preferences, authorizations, and the requests of the data owner, which are to transfer their personal data to a selected service provider or for purposes consented to by the data owner or for legal purposes. According to the EU General Data Protection Regulation (GDPR), the entity managing the user identification data computer server is the data controller of the received personal identification data. This server disseminates the requests and complete identification data of the data owner to the computer servers of all data providers. At the end of the data transfer cycle, it is this user identification data computer server that receives, verifies, and executes requests for re-identification of the data owner.

[0031] A data provider computer server that receives data from a personal software application or from the user identification data computer server stores at least the name and / or other personal identifiers of the data owner, the secret key k3 of the data owner, and the data owner's data transfer request. The data provider computer server searches for the data of interest of the data owner contained in the data provider database and the application computer server, obtains the data of interest, removes the name and / or all other personal identifiers from the data record, replaces the personal identifiers with the secret key k3 of the data owner, and periodically sends the de-identified data records of interest to the computer server of the service provider selected by the data owner.

[0032] The service provider computer server can be controlled by a private or public entity, depending on the choice of the data owner. The service provider computer server receives the data of interest of the data owner from the data provider computer server and identifies only using the secret key k3 of the data owner. Thus, the service provider computer server has no information about the name and / or other personal identifiers of the data owner and cannot know the name and personal identifiers of the data owner. The personal data of the user is irreversibly anonymized for the service provider computer server. The service provider computer server stores the data of interest of the data owner and can upload the data of interest to the personal computing device or personal software application of the data owner for local use, or transfer the de-identified data to the computer of a third party for further processing. According to the GDPR, the service provider entity is the data controller of the received data records of interest of the data owner.

[0033] A third-party computer receives the data of the data owner separately or receives the data of the data owner as part of a large data set, where the data records to be used for further processing to obtain useful results are identified only by the corresponding data owner's key k3 and are thus anonymized. Also in this case, the data of the data owner is irreversibly anonymized because the third-party computer has no means to re-identify it. If re-identification of the data owner is required, the data owner's key k3 and other information of interest (such as the useful result) must be sent to the user identification data computer server.

[0034] All computer devices, servers, and systems of the present disclosure include a data exchange software application and a cryptographic software module in their software applications, and the data exchange software application and the cryptographic software module are configured to securely and confidentially send data to and receive data from connected computer devices, servers, and systems, as well as encrypt, decrypt, and digitally sign the data. They can be physical servers or run in the cloud.

[0035] The use of the data owner's personal code, cryptographic key, and the associated selective masking of the data owner, the use of encryption and decryption, the separation of data processing functions between different types of computer servers controlled and managed by different legal entities, and the use of data minimization techniques (such that only the information relevant to each type of computer processing action is sent to the corresponding computer performing the corresponding computer processing action) together provide a very high level of security. Programs running in the computer systems of the various agents described above control whether the identity of the data owner is selectively displayed or hidden.

[0036] Selective reversible and irreversible anonymization, or "strong pseudonymization", creates conditions such that only authorized parties can re-identify. This function is targeted at data operators and third parties involved in further processing to discover new information using the data owner's data, because data operators and third parties can generate information that is more sensitive than the original data they started with. For example, a researcher may discover that the data owner belongs to a certain risk group, which makes the data owner more likely to develop a serious disease in the future and thus reduces the data owner's employment potential; or an insurance company may discover that the risk classification of an insured data owner needs to be changed, which may lead to an increase in premiums.

[0037] Accordingly, it is desired to prevent any computer used by a researcher or technician from further processing personal data, accessing the name of the data owner and / or any other personal identifier of the data owner, or cooperating or colluding illegally with another party operating the computer server of the currently described system in an attempt to re-identify the data owner. To prevent this from happening, the amount of information available to each operator and their computer system, or even two colluding operators and their computer systems that pool their data, must be insufficient to successfully re-identify the data owner. At least three different entities or operators running different computers or computer servers must cooperate to effectuate data owner and data re-identification in such a way that only the computer systems of authorized entities can access the name and / or other personal identifiers of the data owner. These three entities are the user identification data entity, the service provider entity, and the third party engaged in further processing of the data. Each of these three entities operates a computer system that contains a piece of information required by the other computer systems for useful re-identification.

[0038] It should be noted that the data provider computer server contains the identification data - name, personal identifier, and key k3 - of each data owner that is the same as that of the user identification data computer server, but only the user identification data computer server contains the complete and verified name and contact data of the data owner, making it an ideal system for transmitting the re-identification information of interest to an individual or entity that needs to obtain the re-identification information. In any case, the fact that the data provider computer server also stores the name of the data owner and thus has part of the means to re-identify new information generated by further processing of personal data does not change the fact that, according to this disclosure, at least three computer systems must cooperate to re-identify the de-identified data owner and their data.

[0039] In this process, the security level is determined by the sophistication and complexity of the encryption system, and the encryption keys k1, k2, and k3 are generated when the data owner's personal software application is first installed on the data owner's personal computing device. This provides an operation centered around the data owner.

[0040] Specifically, the data owner contributes by providing their name and / or other personal identifiers and a request consenting that their personal data will be used for purposes useful to the data owner and other computer systems participating in the data exchange. The data provider computer server provides personal data of interest to other parties according to the data owner's request and in accordance with the data portability right of citizens, which is effective under the GDPR in many jurisdictions, especially in the European Union. Examples of such data providers include banks, hospitals, insurance companies, and other public and private institutions. The service provider is an entity trusted by the data owner that processes the anonymized personal data of the data owner in a secure and efficient manner.

[0041] The third party is an expert in data processing and has a scientific, regulatory, or commercial interest in the personal data managed by the service provider, but without access to the name identifier of the data owner. The user identification data entity stores the user preferences and requests of the data owner and the means required to identify and re-identify the data owner and the means required to contact the data owner, but no other personal data of the data owner. All these parties perform these operations through computers that communicate and collaborate with each other for useful or valuable purposes to process the personal data of the data owner, but in a manner that complies with the law and maintains the protection rights of the personal data and privacy rights of the data owner.

[0042] This disclosure describes the means by which the user identification data computer server combines all data items (information of interest, additional information, and the name and other personal identifiers of the data owner) to successfully re-identify the data owner. In the health data scenario described above, important new information will be received by the user identification data computer server, relayed by the service provider computer server, or directly received from a third-party computer, where the information is created by researchers and other experts by processing the original health data. Once the user identification data computer server re-identifies the anonymized data and its data owner, the user identification data computer server will transmit the new information to the computer of the healthcare professional regarding the data owner to take appropriate medical measures, or to other computer systems that need to associate the anonymized data of the data owner with the identity of the data owner.

[0043] In the voting example, voter registration can be periodically verified and compared to the place of residence to clean up the voting rolls of voters who have moved, but the re-identification function will not be able to view how a voter voted. In the insurance scenario, the re-identification function will be disabled to prevent increasing the insurance premiums of individuals with a given risk profile, but the function will be enabled if the data is shared with healthcare professionals or other entities with legitimate and ethical purposes. Thus, depending on the identity and function of the recipient of the personal data, the present disclosure is able to manage the name and / or other personal identifiers of the data owner, thereby allowing their data to be anonymized or re-identified as needed.

[0044] Cryptography is used in the present disclosure. The public key / private key cryptosystem used herein can employ RSA asymmetric key cryptography, as well as variants of the elliptic curve digital signature algorithm (ECDSA), RSA. Future asymmetric key systems can also be used in the present disclosure, provided that these asymmetric key systems use mathematically related private and public keys. In use, the data is encrypted by the data owner's personal software application using the data owner's private key k1 before transmission, and the private key k1 is known only to the data owner's personal software application. The computer receiving the encrypted data uses the data owner's public key k2 (which is generally publicly known) to decrypt the message. This method can use a public key k2 that has itself been encrypted, such that only a computer system with the decryption key for the encrypted public key k2 can decrypt the initiator message encrypted using the private key k1. This can employ different public key / private key pairs, or use a symmetric key (which is used for both encryption and decryption and is confidentially shared between the sender and the receiver), or alternatively employ a hash function. Thus, encrypting or otherwise manipulating the public key introduces an additional layer of security and helps to control which computer systems are authorized to decrypt the data, the name and / or other personal identifiers of the data owner.

[0045] Other important cryptographic methods used herein include cryptographic hashing and digital signatures.

[0046] Modern cryptographic hash algorithms, such as SHA-2, SHA-256, and BLAKE2, are considered secure enough for most applications. SHA-2 is a set of Strong cryptographic hash functions for the password concept are considered highly secure. SHA-3 is considered highly secure and has been published as the official U.S. recommended cryptography standard. The Digital Signature Algorithm (DSA) is a Federal Information Processing Standard for digital signatures. Any of these modern secure hash algorithms or their successors are useful in this disclosure. These functions are typically part of the standard libraries of modern programming languages and platforms. The input value for this function can be a document, a string, or a cryptographic key, and the output is an alphanumeric expression that cannot be reversed back to the original input value since the hash function is a one-way function. This is useful for securely storing passwords and cryptographic keys. Whenever a computer receives a login request or a key that the computer needs to verify, the computer calculates its hash value according to the selected method and compares the calculated hash value with the hash value stored in its storage medium. If the calculated value and the stored value are the same, the input value is confirmed and access is authorized.

[0047] A digital signature is a cryptographic tool used to sign digital messages and verify message signatures in order to provide proof of authenticity for digital messages or electronic documents. Digital signatures provide message authentication, integrity, and non-repudiation, which are features in this disclosure where the receiving computer must verify the identity of the sending computer as well as the identity (name or anonymity) of the data owner of the transmitted information. The digital signature binds the digital message to the data owner's public key or to the transmitting computer's public key.

[0048] In one example, the digital message to be signed includes a request for the sender to log in to the computer of interest, which will prove its content and origin after being successfully read. The sending computer will digitally sign the message using the sender's private key k1 and send the signed message to the receiving computer. The receiving computer will read the signature in the message and verify whether the user is a known user by using the public key k2 of the sending computer known to the receiving computer as well as the methods for asymmetric keys, hash functions, and digital signatures described above. If the content of the signed message is equal to the public key of the sending computer, the digital signature has been successfully verified and login access is authorized.

[0049] To digitally sign a longer digital message, it is useful to first hash the entire message, encrypt the digital message using the sender's private key, and then transmit the resulting encrypted digital message to the receiving computer for verification. Experts in the field will be able to replicate the above methods, which are described in detail at the websites cryptobook.nakov.com / digital-signatures and https: / / www.cisa.gov / uscert / ncas / tips / ST04-018.

[0050] In addition, the blockchain can be used to record information transactions between a user's computer, user data identifying entities, data providers, other data operators, and third-party entities, and can give the data owner the possibility to customize the consent and viewing rights of their personal data via the use of smart contracts (a feature of the blockchain). Smart contracts allow the data owner to specify via their computer which computers of whom can view personal data and data of interest, which categories of data can be viewed and by whom, which data can be written and by whom, which actions can be authorized based on the underlying data of interest, and globally or selectively grant or revoke consent and computer access rights, as well as associate monetization rules with each element of personal data.

[0051] These computer cryptographic methods (symmetric keys, asymmetric keys, hashes, digital signatures, and blockchains) can be effectively combined in the practical aspects of this disclosure. Other cryptographic functions (such as session keys, key rings, ephemeral keys, and tokens) and any other cryptographic systems can also be used.

[0052] To make the system more secure, the system can be supplemented with a secure processing environment where third-party computers operated by researchers or experts involved in further processing the personal data of the data owner cannot display the data itself and cannot physically access the data itself, nor can they access the personal code k3 of the data owner, but can only remotely access the former, i.e., personal data or other data of interest. By preventing direct access to the public key k2 of the data owner, the secure processing environment of this disclosure prevents the factorization of the public key k2 to extract the private key k1, which can technically be achieved using currently developing quantum computers.

[0053] Therefore, the personal data is invisible to the researchers, but only its metadata - data that describes the personal data, such as the number of data owners, classified by attributes of interest to the researchers. Examples of such metadata include statistical data and the number of data owners by gender, by year of birth, by postal code, by occupation, or by type of goods purchased.

[0054] In a healthcare scenario, examples of metadata include the number of patients in each specific diagnosis, treatment, prescription, clinical test result, and medical outcome category. In such a secure processing environment, researchers do not have direct access to the data, which, if subsequently interconnected with other data sources, could potentially reveal the identity of the data owner. For example, a blood test contains 10 to 20 alphanumeric results and a date. The name of the data owner can be found by subsequently cross-referencing this dataset with data from a clinical test service operator's database to which the researcher may have access. In the secure processing environment of the present application, without using the methods and systems described herein, it would be impossible to find the name of the data owner by subsequently cross-referencing the dataset.

[0055] To illustrate how a researcher can further process personal data, the computers used in this secure processing environment allow searches based on selection criteria - such as searching for all diabetic patients aged 50 to 60 with hypertension and a body mass index greater than 30 - and then running a program that calculates the correlation between these health and disease metrics and the medications prescribed to these same patients. This allows the computer to process and reveal which medications are most effective and which ones result in a higher rate of side effects. This is crucial information that must be communicated to each patient and their healthcare professional to confirm or modify the treatment plan. The identification, de-identification, and re-identification methods described in this disclosure, when combined with the secure processing environment, substantially reduce the risk of re-identification by unauthorized persons or computers to zero.

[0056] In another aspect of the present invention for enhancing data security, the number of keys that make up the global secret identifier can be increased from three to four or even more, such that each receiving computer has a different form of the data owner's public key k2 and personal code k3. In this case, there must be a computer table where each registered user entry contains all the different forms of the public key k2 and personal code k3 for that data owner, and the table can be stored in the service provider's computer server. Alternatively, the table can be divided into two parts, one stored in the service provider's computer server and the other stored in the user identification data computer server. However, one private key k1, one public key k2, and one personal code k3 are sufficient for the present method and system to work as described in this disclosure.

[0057] This disclosure describes a fully integrated personal data protection system and data protection method operated by a computer that starts from the data owner's personal computing device and links the participating computers in an unbroken thread, where data processing is enabled and controlled by the personal computing device.

[0058] The system and method include a data owner or data subject, a personal computing device of the data owner running a personal software application, a user identification data computer server, a data provider computer server, a service provider computer server, and a third-party computer. All these computing devices and servers run applications, data exchange programs, encryption and decryption programs, and store data in files and databases. Using these computing devices and servers configured to run the software allows the system and method of the present disclosure to operate and provides a solution to the technical problem determined for protecting personal data and privacy rights based on the leakage, hiding, and re-leakage of personal names and identities within the computers of this system.

[0059] The personal computing device of the data owner can be, for example, a personal computer, a smartphone, or a tablet, but a smartphone is preferred because it is a very personal device. The smartphone or other personal computing device can be configured to perform the steps described herein through a personal software application downloaded to the personal computing device from a website or from an application distribution service (such as the App Store or Google Play). The purpose of this personal software application is to enable the user and the data owner to manage personal data and implement the process by which this user data is identified, de-identified, and re-identified.

[0060] The computer server system described herein includes computer servers operated by data providers, user identification data entities, third parties participating in further processing, and other data operators, and these servers are generally systems with relatively strong computing power, storage, and communication capabilities. Data exchange occurs between the personal computing device, the user identification data computer server, the data provider computer server, the service provider computer server, and the third-party computer through appropriate data communication software stored respectively.

[0061] All computer systems in the present invention include a processor, a memory capable of storing program instructions, a communication subsystem, a storage medium, input devices (such as a keyboard, a mouse, a pointer, a touch screen, a microphone, or a camera), and output devices (such as a display screen and speakers). The computer systems are capable of communicating with each other using private or public telecommunication networks, but a public network is preferred, and the preferred medium is the Internet. They can also be in a private network or a virtual private network. All communications between the parties are preferably encrypted using standard Internet protocols, such as HTTP + TLS / SSL and IPsec or any subsequent protocol with equivalent or higher security.

[0062] The data owner is represented by the personal code k3, which serves as the data owner's identification in cases where names or regular personal identifiers have been removed from the data records of interest. The personal code is large enough to uniquely identify each member of the target population. In addition to the key k3, the data owner can also be represented and uniquely identified by the public key k2, but only the data owner's personal software application and the service provider's computer server know k2. The keys k2 and k3 are unique codes, each uniquely representing the data owner. Using these two components of the global secret identifier (the use of each component depends on the identity and function of the data recipient) to specify the same data owner increases the complexity and security level of the protection measures.

[0063] In one aspect, the data owner or user downloads a personal software application to a personal computing device, which can be a smartphone, tablet, laptop, or desktop computer. Installing the personal software application by the data owner includes defining a locally stored PIN or password and entering a name and / or other personal identifiers, user preferences, authorizations, and requests specific to the actual use of the personal software application. During the installation of the personal software application, the personal software application can include steps to confirm the identity of the data owner, and this confirmation can be achieved using government-provided authentication methods, in-person confirmation at a registry, biometric methods, connection to a trusted database with a verified identity, or any other acceptable confirmation method. This confirmation step is desirable because it can create certainty about the identity of the user, which is crucial in activities such as healthcare and voting.

[0064] The personal software application also generates a cryptographic key, namely a pair of asymmetric public key k2 and private key k1, as well as the data owner's personal code k3.

[0065] During the installation of the personal software application, the personal computing device makes two contacts, the first contact is with the user data identification computer server, and the second contact is with the service provider computer server.

[0066] During the first contact with the user data identification computer server, the personal software application sends the data owner's name and / or other personal identifiers, user preferences, authorizations, and requests, as well as the personal code k3. The user identification data computer server receives this data, stores it, and then transfers this data to the data provider computer server(s), where there can be one or more data provider computer servers.

[0067] The data provider computer server receives data including the name of the data owner, personal identifiers, a request to copy the personal data to a designated service provider computer server, and a personal code k3. The data provider computer server searches its application database for data related to the data owner, and upon finding the data, obtains the data and replaces the name of the data owner and all other personal identifiers with the personal code k3, and transmits the thus anonymized and de-identified data to the service provider computer server. This action is repeated periodically whenever new data of interest is stored in the application database of the data provider computer server.

[0068] In a second contact, almost simultaneous with the first contact, the personal software application sends a message to the service provider computer server to indicate that a new anonymized data owner has been created, and transmits the keys k2 and k3, as well as preferences, authorizations, and requests, without transmitting any name or any other type of personal identifier. The installation of the personal software application is complete.

[0069] When the data owner first uses the newly installed application, the personal computing device of the data owner with the personal software application running thereon needs to be linked to the service provider's computer server. The data owner enters a personal PIN to open the personal software application, and then the application sends a digitally signed message to the service provider computer server. This process will prove to the service provider computer server that the message was sent by the data owner previously registered during the above-mentioned second contact. The digitally signed message may contain a timestamp so that the recipient can only process the most recent message. The digitally signed message consists of at least the personal code k3 in plain text and the public key k2 encrypted with the private key k1. The service provider computer server reads the personal code k3 and uses it to retrieve the registration data of the data owner from its user registration database, and reads the stored key k2 of the data owner therefrom. The service provider computer uses this key k2 and attempts to decrypt the encrypted part of the digitally signed message. If the decryption returns the public key k2 of the data owner identical to the key k2 stored in the user registration database, the message is considered valid and the data owner is authenticated.

[0070] Successfully reading the digital signature proves that the digital signature was signed by the private key k1 associated with the user's public key k2. This allows the service provider computer server to confirm that the personal computing device used to install the personal software application of the data owner, who wishes to confirm their identity, is the same device that the same data owner is currently using, and that the data owner has entered the same PIN defined previously. The service provider computer server now stores the keys k2 and k3 as anonymous and confidential identifiers of the same data owner. Apart from the data owner's personal computing device, the service provider computer server is the only computer server in this disclosure that can access the public key k2 and the personal code k3.

[0071] All subsequent communication sessions between the personal software application and the service provider computer server are always initiated by the data owner using the personal computing device and logging into the personal software application, such that the personal software application sends a digitally signed message that is the same as or substantially similar to the digitally signed message generated during the first communication session with the service provider computer server. The service provider computer server successfully reads the message and determines that the data owner is a known and valid data owner, enabling data communication between the service provider computer server and the personal software application, thereby uploading data of interest to the personal software application or downloading from the personal software application. Thus, the service provider computer server is able to communicate securely with a data owner whose identity is unknown but who is indeed the data owner of personal data.

[0072] Subsequently, the service provider computer server can enable data communication between its application database and a third-party computer to further process the data owner's data that is regularly obtained from the data provider's computer server. Access to the personal data of interest will be provided, where each data record is now uniquely identified by the personal code k3. The third-party computer can only access the de-identified data.

[0073] If it is necessary to re-identify the data owner to convey new or important information, such as but not limited to important information in the context of health, or for any other reason, the third-party computer will send the new information, along with the personal code k3 of the data owner associated with the new information, to the service provider computer server. The service provider computer server receives the information, and the personal code k3 sends the information and the personal code k3 to the user identification data computer server.

[0074] The user identification data computer server receives the message and uses the personal code k3 stored in its user registration database to read the name of the data owner and / or other personal identifiers associated with the data owner, as well as any contact details for the data owner. The user identification data computer server uses these contact details to forward new significant or useful information generated by a third-party computer using anonymized data to the data owner or other interested parties, but this time re-identifying using the name of the data owner and / or other personal identifiers.

[0075] In addition to the data owner's personal software application, the only entity with access to the data owner's public key k2 and personal code k3 is the service provider computer server, which plays an important role in restoring the anonymization of the de-identified data without knowing the name and / or other personal identifiers of the data owner. Furthermore, only the service provider computer server can confirm that the data owner's personal code was created in the same personal computing device that is subsequently used for the actual purpose of the personal software application. If the desired steps for user identification are included during or after the installation process of the personal software application, then the service provider computer server can also provide the ability to trace the de-identified personal data back to its legal owner without any risk of misidentification and without knowing the name of the data owner. It is through the collaboration of the service provider computer server and the user identification data computer server that the data owner can be securely identified, de-identified, and re-identified through the cryptographic keys and different keys of the global secret identifier generated by the data owner's personal computing device and selectively transmitted to the computer system of the present disclosure. BRIEF DESCRIPTION OF THE DRAWINGS

[0076] Figure 1A A block diagram showing the system architecture for downloading and installing a personal software application onto a personal computing device according to an example aspect of the present disclosure.

[0077] Figure 1B A block diagram showing the computer system architecture for connecting a personal computing device, a user identification data computer server, a data provider computer server, a service provider computer server, and a third-party computer to establish communication between them according to an example aspect of the present disclosure.

[0078] Figure 1C A flowchart showing a computer-implemented method for re-associating anonymized data with a data owner according to an example aspect of the present disclosure.

[0079] Figure 2A A data flow diagram showing the installation of a personal software application in a personal computing device according to an example aspect of the present disclosure.

[0080] Figure 2B Shows a data flow diagram for receiving updates in a personal software application between a personal computing device and a service provider computer server according to an example aspect of the present disclosure.

[0081] Figure 2C Shows a data flow diagram for describing the transmission of de-identified personal data to a third-party computer and the re-identification by a service provider computer server and a user identification data computer server according to an example aspect of the present disclosure.

[0082] Figure 3 Shows a block diagram of a personal computing device according to an example aspect of the present disclosure.

[0083] Figure 4 Shows a block diagram of a user identification data computer server according to an example aspect of the present disclosure.

[0084] Figure 5 Shows a block diagram of a data provider computer server according to an example aspect of the present disclosure.

[0085] Figure 6 Shows a block diagram of a service provider computer server according to an example aspect of the present disclosure.

[0086] Figure 7 Shows a block diagram of a third-party computer server according to an example aspect of the present disclosure. Detailed Description

[0087] The present invention will now be described based on the accompanying drawings. It should be understood that the aspects and aspects of the present invention described herein are merely examples and do not limit the scope of protection of the claims in any way. The present invention is defined by the claims and their equivalents. It should be understood that the features of one or more aspects of the present invention can be combined with the features of one or more different aspects and / or aspects of the present invention.

[0088] In Figure 1A , the data owner 10 or the data subject or user, patient or individual connects their personal computing device 110 to a communication network. The communication network can be a public mobile digital communication network, such as the Global System for Mobile Communications (GSM) or the Internet. Then, the data owner 10 downloads a software application 1000, which contains the programs required for the operation of the methods and systems of the present disclosure. The software application 1000 is downloaded from an appropriate software distribution system 100 (such as the App Store or Google Play) or an Internet website to which the data owner 10 connects using the personal computing device 110.

[0089] Now by reference to Figure 2A ,Figure 2B and Figure 2C supplement with the data flow diagrams in Figure 1B for a detailed explanation.

[0090] In Figure 1B the software application 1000 is installed as the personal software application 1100 of the data owner 10, including the identifier 300, such as the name of the data owner 10 and / or other personal identifiers, preferences, authorizations, and requests input or provided by the data owner 10. The personal software application 1100 records the PIN or password defined by the data owner 10 and generates a pair of asymmetric cryptographic keys, namely the private key 190, the public key 200, and the personal code 210. This corresponds to Figure 2A step 1 in

[0091] The personal software application 1100 of the personal computing device 110 sends (arrow a) a data message to the user identification data computer server 120. The data message includes the identifier 300 of the data owner 10 (such as the name and / or other personal identifier 300, preferences, authorizations, requests) and the personal code 210, but does not include the PIN, the private key 190, or the public key 200. This corresponds to Figure 2A step 1 in

[0092] The software application 1200 contained in the user identification data computer server 120 receives, stores, and processes the data message from the personal software application 1100 and sends the data message (arrow b) to the data provider computer server 130. This corresponds to Figure 2A step 3 in

[0093] The personal software application 1100 of the personal computing device 110 also sends a data message (arrow c) to the service provider computer server 140 - including preferences, authorizations, requests, the public key 200 digitally signed with the private key 190, the plaintext personal code 210, but not including the name of the data owner 10, the personal identifier 300, the PIN, or the private key 190. This corresponds to Figure 2A step 2 in

[0094] The service provider computer server 140 receives and stores the data received from the data provider computer server 130. This corresponds to Figure 2A step 4 in

[0095] The software application 1300 contained in the data provider computer server 130 receives, stores, and processes data messages received from the software application 1200 in the user identification data computer server 120. Based on the preferences, authorizations, and requests of the received data owner 10, and using the name and / or other personal identifiers 300, the data provider computer server 130 searches for personal data of interest to the data owner 10. After finding the personal data, the name and other personal identifiers 300 are removed from the personal data, and the name and other personal identifiers 300 are replaced with the personal code 210 of the data owner. To improve the interoperability, usability, and effectiveness of the transmitted data, the parties operating the data provider computer server 130 and the service provider computer server 140 can agree on data standards, encoding / decoding standards, communication protocols, and database formats so that when the service provider computer server 140 receives the personal data, the personal data is already of high quality, confidential, identified, and very suitable for further processing. Then, the data provider computer server 130 sends (arrow d) the de-identified, structured, and organized personal data of interest and the personal code 210 of the data owner 10 to the service provider computer server 140. This corresponds to Figure 2A Step 5 in

[0096] As long as the data transfer request of the data owner 10 remains valid and in effect, and whenever new personal data of interest is stored at the data provider computer server 130, the de-identified, structured, and organized data of interest and the personal code 210 of the data owner 10 are periodically transferred to the service provider computer server 140. For example, this is the case when a patient visits and a new diagnosis is established, or when a voter changes their residence.

[0097] The software application 1400 contained in the service provider computer server 140 receives and processes data messages from the software application 1300 in the user identification data computer server 120 and stores the data messages as de-identified data, uniquely identified by the public key 200 and the personal code 210. Since the service provider computer server 140 has access to both the public key 200 and the personal code 210, it does not matter which of the two forms is used internally by the service provider computer server. This corresponds to Figure 2A Step 6 in

[0098] For regularly updating the data of interest in the personal software application 1100 during daily operations, when the data owner 10 opens the personal software application 1100, the personal software application 1100 contacts the service provider computer server 140 by sending (arrow c) the personal code 200 of the data owner 10 (digitally signed using the private key 190 of the data owner 10) and the personal code 210. This corresponds to Figure 2B Step 1.

[0099] After receiving the personal code 210 of the data owner 10, the service provider computer server 140 reads the digital signature, and if the reading is successful, it confirms that the public key 200 and the personal code 210 of the data owner 10 are valid requests for the previously registered data owner 10. The service provider computer server 140 uploads (arrow c) the new data of interest stored in its database since the last update from the data provider computer server 130 to the personal software application 1100. This corresponds to Figure 2B Step 2.

[0100] The personal software application 1100 in the personal computing device 110 receives and stores the updated data of interest. This corresponds to Figure 2B Step 3.

[0101] Thus, the personal software application 1100 of the data owner 10 regularly receives the latest information or data of interest originating from the data provider computer server 130. The advantage is that the information belonging to the same data owner 10 stored in multiple data provider computer servers 130 can be uploaded to the personal computing device 110, where the information is stored in a structured and well-organized manner in the personal software application 1100. Applications that particularly benefit from the combination of multi-source data are electronic health record applications, where the user or data owner 10 can have their complete health history from different hospitals and can easily access it on their personal computing device 110, and the data owner 10 can easily share clinical data with designated healthcare professionals from it.

[0102] Another advantage is that even if the service provider computer server 140 does not know the identifier 300 of the data owner 10 (such as name and / or other personal identifiers), the service provider computer server 140 is confident that the data owner 10 and the owner of the personal computing device 110 are the same person. The data of interest will always be transmitted to its owner, and it is technically impossible to have an incorrect identity when the service provider computer server 140 uploads personal data to the personal computing device 110 of the data owner 10.

[0103] Still in Figure 1BIn this case, third-party researchers, data experts, or data scientists are interested in the personal data contained in the service provider's computer server 140, and use a third-party computer 150 that includes a software application 1500 to connect (arrow e) to the service provider's computer server 140 and request access to the identified data for further processing. This corresponds to Figure 2C step 1.

[0104] Due to the large scale of the database containing personal data accumulated over time at the service provider's computer server 140, this database is valuable to data scientists and can be processed or used for algorithm development using state-of-the-art technologies such as big data, machine learning, and artificial intelligence, as well as traditional statistics. Accordingly, the third-party computer 150 connects to the service provider's computer server 140 and preferably to a secure processing environment included in the software application 1400, which is shown as 1480 in Figure 6 In some embodiments, the secure processing environment 1480 can be a separate computer server, different from the service provider's computer server 140.

[0105] The secure processing environment 1480 of the service provider's computer server 140 only allows authorized data scientists to access the service provider's computer server 140. The de-identified data of interest will not be stored in the third-party computer 150, nor will the de-identified data be downloaded to the third-party computer 150. The de-identified data always remains in the service provider's computer server 140 and is processed by sending commands (arrow e) from the third-party computer 150, thereby triggering data processing operations in the service provider's computer server 140. Accordingly, the data scientist operating the third-party computer 150 is granted access to the personal data, rather than access to its physical ownership, and the secure processing environment 1480 enables the third-party computer 150 to read data records identified only by the personal code 210 of the data owner. This corresponds to Figure 2C step 2 in

[0106] Data scientists do not have direct access to the data of interest, so it is impossible to match any elements of the data of interest with other databases that may have data elements related to the same person, creating conditions for uncontrolled user re-identification. Instead, data scientists specify a data strategy by defining search criteria to obtain the target population and studying the criteria to obtain the desired results. The former will include data science programs and algorithms, preferably contained in a secure processing environment 1480 so that the source data and the operating computer programs that process them are part of the same computer space. This further data processing may produce information or results 220, such as useful information that needs to be re-associated with its data owner 10, or sometimes even new discoveries that are crucial to the data owner 10. This corresponds to Figure 2C step 3.

[0107] When it is found that the result 220 or data element or new information or new important or useful information needs to be associated with its data owner 10, the third-party computer 150 of the data scientist only has the personal code 210 of the data owner 10 to identify the relevant data owner 10. The third-party computer 150 sends (arrow f) the result 220 together with the new information and the personal code 210 of the relevant data owner 10 to the service provider computer server 140. This corresponds to step 4 of Figure 2c.

[0108] The service provider computer server 140 receives the new information or result 220 and the personal code 210 of the data owner 10, searches its user registration database to verify whether the personal code 210 corresponds to a known data owner 10, and also collects more personal information that was not included in the original data scientist's search parameters but may be useful and provide a better context. In a healthcare scenario, this could be adding a complete clinical history to the new information or result 220 of interest to be sent to the relevant but still de-identified patient or data owner 10 and the treating healthcare professional. The service provider computer server 140 sends (arrow g) the complete data message to the user identification data computer server 120. This corresponds to Figure 2C step 5.

[0109] The user identification data computer server 120 receives the data message from the service provider computer server 140, including the information or result 220 of interest to a specific user or group of data owners 10 and their personal code 210. Since the user identification data computer server 120 has previously recorded the identifier 300 of the data owner 10, such as name, personal identifier, and personal code 210 ( Figure 2AStep 3), so that the user identification data computer server 120 can look up the personal code 210 of the data owner 10 in its user database and identify the identifier 300 of the data owner 10, such as name, other personal identifiers, and contact details. This corresponds to Figure 2C Step 6.

[0110] The user identification data computer server 120 can now send the new information or result 220 associated with the relevant data owner 10 to the computers of all authorized parties (such as the data owner 10 or the healthcare professionals of the patient) that need to know the new information or result 220, or even to data scientists if there is a valid and legal motivation, and is now identified by its name and identifier 300. Ideally, the re-identified data should never be sent to the service provider computer server 140 so that its anonymized data always remains anonymous, and the service provider computer server 140 is technically unable to recover the de-identification of the data owner, and the data controller operating it can claim to be a trusted home for the secure custody of people's personal data.

[0111] The following table illustrates which computer systems can access which user identifiers 300, password keys 190, 200, and personal code 210, where 1 indicates access and 0 indicates no access.

[0112]

[0113]

[0114] The table shows that only those computer systems that are personal devices or are designed to manage identifiable data can access the identifier 300 (such as name and / or other personal identifiers) of the data owner 10. The most well-known confidential identifier is the personal code 210. Since the third-party computer 150 is the first to discover the result 220, and the result 220 may contain sensitive or confidential information of the data owner 10, specific security measures should be provided. The third-party computer 150 cannot directly and physically access the public key 200 or personal code 210 of the data owner 10 when processing data in the secure processing environment 1480 via the access interface 1585 of the secure processing environment 1480. The service provider computer server 140 operating the secure processing environment 1480 cooperates with the user data identification computer server 120 to match and re-associate the personal code 210 with the name and identifier 300 of the data owner 10.

[0115] To re-identify data owner 10, three computer systems in this disclosure must cooperate: a third-party computer 150 derives information or result 220 that needs to be associated with the name and identifier 300 of the relevant data owner 10, and a service provider computer server 140 verifies that data owner 10, whose personal data has been processed to obtain result 220, is a validly registered user, and links result 220 and any other useful information for transmission to a user identification data computer server 120. The user identification data computer server 120 uses personal code 210 to retrieve the identifier 300 of data owner 10, such as name and / or other personal identifiers and / or contact details.

[0116] Figure 1C A flowchart depicting a computer-implemented method for re-associating anonymized data with data owner 10 is shown. In step S100, anonymized data stored on a service provider computer server 140 is accessed by a third-party computer 150. The step S100 of accessing anonymized data includes processing the anonymized data S100 and obtaining result 220 resulting from accessing and / or processing the anonymized data in step S100. Step S102 includes the need to associate the anonymized data used to create result 220 with data owner 10 in step S104.

[0117] Personal code 210 is transmitted from the third-party computer 150 to the service provider computer server 140 in step S110. Result 220 is transmitted from the third-party computer 150 to the service provider computer server 140 together with personal code 210 in step S112. Personal code 210 is matched with public key 200 at the service provider computer server 140 in step S120, and personal code 210 is transmitted from the service provider computer server 140 to the user identification data computer server 120 in step S130.

[0118] Result 220 is transmitted from the service provider computer server 140 to the user identification data computer server 120 together with personal code 210 in step S132. Personal code 210 is matched by the user identification data computer server 120 with data owner identifier 300 in step S140. Result 220 is transmitted to the personal computing device 110 of data owner 10 and the computer of the attending healthcare professional of data owner 10 in step S150.

[0119] Figure 3 The hardware and software components of personal computing device 110 are shown. Personal computing device 110 includes a main processor 1101, a communication subsystem 1110 (designed to communicate with participating computer systems via a communication network, such as Figure 1B andFigure 2A , Figure 2B and Figure 2C as shown), an input device 1120, a display 1130, and a storage media subsystem 1135 that stores computer programs and data. The processor 1101 interacts with a memory 1102 that contains personal software applications 1100 and data retrieved from the media storage subsystem 1135. The processor 1101 loads program instructions 1140, application programs 1150, and data from files storing interesting information 1170 received from a service provider computer server 140 into the memory 1102 as needed.

[0120] Figure 4 shows the hardware and software components of one or more user identification data computer servers 120. One or more user identification data computer servers 120 include a main processor 1201, a communication subsystem 1210 (designed to communicate via a communication network with Figure 1B and Figure 2A , Figure 2B and Figure 2C the participating computer systems shown in), an input device 1220, a display 1230, and a storage media subsystem 1235 that stores computer programs and data. The processor 1201 interacts with a memory 1202 that contains all software programs 1200 and data retrieved from the media storage subsystem 1235. The processor 1201 loads program instructions 1240, application programs 1250, and a user registration database 1260 into the memory 1202 as needed.

[0121] Figure 5 shows the hardware and software components of a data provider computer server 130. The data provider computer server 130 includes a main processor 1301, a communication subsystem 1310 (designed to communicate via a communication network with, such as Figure 1B and also 2A, Figure 2B and Figure 2Cshown participating in communicating with a computer system), an input device 1320, a display 1330, and a storage media subsystem 1335 that stores computer programs and data. The processor 1301 interacts with a memory 1302 that contains a software program 1300 and data retrieved from the media storage subsystem 1335. The processor 1301 loads program instructions 1340, application programs 1350, a user registration database 1360, and an application database 1370 containing information of interest to the data owner 10 into the memory 1302 as needed. In use, the program instructions 1340 will search the user registration database 1360 using an identifier 300 (such as the name and / or other personal identifier of the data owner 10), and after finding the data owner identifier 300, search the application database 1370 for the data of interest to that data owner 10. Then, before sending the information of interest to the service provider computer server 140, the program instructions 1340 will replace the identifier 300 (such as the name and / or other personal identifier) of the data owner 10 with the personal code 2 of the data owner 10.

[0122] Figure 6 The hardware and software components of the service provider computer server 140 are shown. The service provider computer server 140 includes a main processor 1401, a communication subsystem 1410 (designed to communicate via a communication network with, for example, Figure 1B and Figure 2A , Figure 2B and Figure 2C shown participating in communicating with a computer system), an input device 1420, a display 1430, and a storage media subsystem 1435 that stores computer programs and data. The processor 1401 interacts with a memory 1402 that contains a computer program 1400 and data retrieved from the media storage subsystem 1435. The processor 1401 loads program instructions 1440, application programs 1450, a user registration database 1460, and an application database 1470 containing the data of interest to the data owner 10 into the memory 1402 as needed. In use, the program instructions 1440 will receive from the data provider computer server 130 the data of interest related to the data owner 10, which is identified only by the personal code 210 of the data owner 10, and store it in the application database 1470.

[0123] Program instruction 1440 will also regularly upload information of interest to the personal software application 1100 of the data owner 10 and allow access by the third-party computer 150 to further process the data of interest contained in the application database 1490 of the service provider computer server 140. Program instruction 1440 will also authorize the third-party computer 150 to access the secure processing environment 1480 running in the memory 1402, receive the result 220 containing new information of interest of the relevant data owner 10 identified by its personal code 210 from the third-party computer 150, and transmit the data of interest 200 and the personal code 210 of the data owner 10 to the user identification data computer server 120 for re-identification.

[0124] Figure 7 The hardware and software components of the third-party computer 150 for further processing personal data stored in the service provider computer server 140 are shown. The third-party computer server 150 includes a main processor 1501, a communication subsystem 1510 designed to communicate with the Figure 1B and Figure 2A , Figure 2B and Figure 2C participating computer systems shown in, an input device 1520, a display 1530, and a storage medium subsystem 1535 for storing computer programs and data. The processor 1501 interacts with the memory 1502 containing the computer program 1500 and the data retrieved from the medium storage subsystem 1535. The processor 1501 loads the program instruction 1540 and the software module providing the access interface 1585 to the secure processing environment 1480 of the service provider computer server 140 into the memory 1502 as needed. When the third-party computer 150 creates the result 220 with information of interest for one or more data owners 10 associated with the original data of interest, the third-party computer 150 transmits the result 220 with information of interest together with the personal code 210 of the data owner 10 to the service provider computer server 140 via the secure processing environment 1480.

[0125] Although specific aspects of the systems and methods of the present invention have been shown in the drawings and described in the foregoing detailed description, it should be understood that the present invention is not limited to the disclosed aspects, namely the use of more or less encryption means and their various combinations, but is capable of various rearrangements, modifications, and substitutions without departing from the spirit of the present invention.

[0126] Reference numerals

[0127] 1 System

[0128] 10 Data owner

[0129] 100 Software Distribution System

[0130] 110 Personal Computing Device

[0131] 120 User Identification Data Computer Server

[0132] 130 Data Provider Computer Server

[0133] 140 Service Provider Computer Server

[0134] 150 Third-Party Computer

[0135] 190 Data Owner's Private Key

[0136] 200 Data Owner's Public Key

[0137] 210 Data Owner's Personal Code

[0138] 220 Result

[0139] 300 Data Owner's Identifier and Contact Information

[0140] 1000 Software Application

[0141] 1100 Personal Software Application

[0142] 1101 Main Processor / Processor

[0143] 1102 Memory

[0144] 1110 Communication Subsystem

[0145] 1120 Input Device

[0146] 1130 Display Screen

[0147] 1135 Media Storage Subsystem

[0148] 1140 Program Instructions

[0149] 1150 Application

[0150] 1170 Information of Interest

[0151] 1200 Software Application

[0152] 1260 User Registration Database

[0153] 1300 Software Application / Program

[0154] 1400 Software Application

[0155] 1480 Secure Processing Environment

[0156] 1500 Software application / computer program

[0157] 1585 Access interface to the security processing environment

[0158] S100 Access anonymized data

[0159] S102 Obtain the re-association requirement

[0160] S104 Obtain the result

[0161] S110 Transmit the personal code in the first form

[0162] S112 Transmit the result

[0163] S120 Match the personal code in the first form

[0164] S130 Transmit the personal code in the second form

[0165] S132 Transmit the result

[0166] S140 Match the personal code in the second form

[0167] S150 Transmit the result.

Claims

1. A computer-implemented method for re-associating anonymized data with a data owner (10), wherein the data owner (10) has an associated public key (200) and a personal code (210), the method comprises the following steps: a. accessing (S100) the anonymized data stored on a service provider computer server (140) by a third-party computer (150); b. transmitting (S110) the personal code (210) from the third-party computer (150) to the service provider computer server (140); c. matching (S120) the personal code (210) with the public key (200) at the service provider computer server (140); d. transmitting (S130) the personal code (210) from the service provider computer server (140) to a user identification data computer server (120); and e. matching (S140) the personal code (210) with a data owner identifier (300) by the user identification data computer server (120).

2. The computer-implemented method according to claim 1, wherein, the data owner identifier (300) is at least one of the name of the data owner (10) or other personal identifiers.

3. The computer-implemented method according to claim 1 or 2, wherein, the step of accessing (S100) the anonymized data includes the step of obtaining (S102) a requirement for associating the anonymized data with the data owner (10).

4. The computer-implemented method according to any one of claims 1 to 3, wherein the step of obtaining (S102) a requirement for associating the anonymized data comprises: obtaining (S104) a result (220) obtained by the third-party computer (150) accessing and processing (S100) the anonymized data.

5. The computer-implemented method according to claim 4, further comprises the following steps: a. transmitting (S112) the result (220) together with the personal code (210) from the third-party computer (150) to the service provider computer server (140); and b. transmitting (S132) the result (220) together with the personal code (210) from the service provider computer server (140) to the user identification data computer server (120).

6. The computer-implemented method according to claim 5, further comprises the step of transmitting (S150) the result (220) to the personal computing device (110) of the data owner (10) and the computer of any authorized entity that needs to know the result (220).

7. The computer-implemented method according to any one of the preceding claims, wherein the public key (200) is the public asymmetric cryptographic key of the data owner (10), and the personal code (200) is either a hash function of the public key (200) of the data owner (10) or a random number recorded in a computer file associated with the public key (200).

8. The computer-implemented method according to any one of the preceding claims, wherein the public key (200) is known only to the service provider computer server (140) and the personal computing device (110) of the data owner (10).

9. A system (1) for re-associating anonymized data with a data owner (10), wherein, the data owner (10) has an associated private key (190), public key (200) and personal code (210), and the system (1) comprises: a. A service provider computer server (140) for: i. Storing anonymized data, ii. Matching the public key (200) with the personal code (210), and iii. Transmitting the personal code (210) to a user identification data computer server (120); b. A third-party computer (150) for accessing the anonymized data and transmitting the personal code (210) to the service provider computer server (140); and c. The user identification data computer server (120) for matching the personal code (210) with a data owner identifier (300).

10. The system (1) according to claim 9, further comprising a data provider computer server (130) for transmitting the anonymized data to the service provider computer server (140).

11. The system (1) according to claim 9 or 10, further comprising a personal computing device (110) of the data owner (10) for a. Generating the private key (190), the public key (200) and the personal code (210); b. Transmitting the public key (200) only to the service provider computer server (140); and c. Transmitting the personal code (210) to the user identification data computer server (120), the data provider computer server (130), the service provider computer server (140) and the third-party computer (150).

12. Using the method according to any one of claims 1 to 8 to store at least one of health data and election data.

13. A computer program comprising instructions which, when executed by a computer, cause the computer to perform the method according to any one of claims 1 to 8.

14. A computer-readable medium comprising instructions which, when executed by a computer, cause the computer to perform the method according to any one of claims 1 to 8.

Citation Information

Patent Citations

  • A one-click login procedure

    WO2020165174A1

  • A computer system and method of operating same for handling anonymous data

    WO2020221778A1