Method and apparatus for de-identifying shared data

By using hashing algorithms and encryption processing of trusted execution environments during data sharing, the privacy information leakage problem caused by a single encryption algorithm is solved, secure data sharing and privacy information protection are achieved, and the maximum utilization of data value is ensured.

CN114462088BActive Publication Date: 2025-07-22ALIPAY (HANGZHOU) INFORMATION TECH CO LTD
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202210117022.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-02-07
Publication Date
2025-07-22
Estimated Expiration
2042-02-07

AI Technical Summary

Technical Problem

In the process of data sharing in the prior art, the de-identification processing of a single encryption algorithm poses a risk of private information leakage, which cannot effectively ensure data security and cannot realize secure data sharing among all data parties.

Method used

The hashing algorithm is used to hash the privacy information in the shared data, generate the first byte array, and perform truncation processing to obtain the second byte array. These arrays are encrypted and fused in the data fusion device of the trusted execution environment to generate shared data.

Benefits of technology

It improves the security of private information in shared data, ensures that private information during data sharing is not leaked, meets the requirements of laws and regulations, and maximizes the use of data value.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114462088B_ABST
    Figure CN114462088B_ABST
Patent Text Reader

Abstract

The embodiments of this specification provide a method and apparatus for de-identifying shared data. In this method, at each data party for sharing data, a hash algorithm is used to perform a hash calculation on the privacy information in the data to be shared, so as to obtain a corresponding first byte array; at each data party, the obtained first byte array is truncated to obtain a second byte array; at a data fusion device with a trusted execution environment, each second byte array obtained from each data party is respectively encrypted in the trusted execution environment; and at the data fusion device, in the trusted execution environment, the data to be shared including the encrypted second byte arrays is fused to obtain shared data.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments of this specification relate to the technical field of data processing, and specifically, to a method and apparatus for de-identifying shared data. Background Art

[0002] A data party as a legal entity may legally collect data, including user data, operation data, etc. Each data party may use the collected data. For example, the collected data may be used as a sample for model training. However, due to privacy data security issues, generally, each data party can only use its own data and cannot cross-share with other data parties.

[0003] However, only through data sharing can the value of data be maximally released and the utilization of data be maximized. To maximize the value of data, data sharing can be carried out among data parties. However, data sharing needs to ensure the security of privacy data and meet the requirements of laws and regulations. To meet these requirements, privacy data needs to be de-identified before data sharing to hide privacy information and avoid leakage of privacy information. Currently, the main de-identification means is to perform encryption processing using data encryption algorithms. Summary of the Invention

[0004] In view of the above, the embodiments of this specification provide a method and apparatus for de-identifying shared data. Through the embodiments of this specification, a more secure de-identification means is provided, improving the security of privacy information in shared data.

[0005] According to one aspect of the embodiments of this specification, a method for de-identifying shared data is provided, including: at each data party for shared data, performing a hash calculation on privacy information in the data to be shared using a hash algorithm to obtain a corresponding first byte array; at each of the data parties, truncating the obtained first byte array to obtain a second byte array; at a data fusion device having a trusted execution environment, performing encryption processing on each of the second byte arrays obtained from the data parties in the trusted execution environment; and at the data fusion device, performing data fusion on the data to be shared including the encrypted second byte arrays in the trusted execution environment to obtain shared data.

[0006] According to another aspect of the embodiments of the present specification, there is also provided a method for de-identifying shared data, which is executed by each data party for sharing data. The method includes: performing a hash calculation on the privacy information in the data to be shared by using a hash algorithm to obtain a corresponding first byte array; performing a truncation process on the obtained first byte array to obtain a second byte array; and sending the data to be shared including the second byte array to a data fusion device having a trusted execution environment, so that the data fusion device encrypts the second byte arrays obtained from each data party in the trusted execution environment, and fuses the data to be shared including the encrypted second byte arrays to obtain shared data.

[0007] According to another aspect of the embodiments of the present specification, there is also provided a method for de-identifying shared data, which is executed by a data fusion device having a trusted execution environment. The method includes: obtaining, from each data party of the shared data, the data to be shared including a second byte array, where the second byte array is obtained by performing a truncation process on a first byte array obtained by each data party performing a hash calculation on the privacy information in the data to be shared by using a hash algorithm; performing an encryption process on each of the obtained second byte arrays in the trusted execution environment; and fusing, in the trusted execution environment, the data to be shared including the encrypted second byte arrays to obtain shared data.

[0008] According to another aspect of the embodiments of the present specification, there is also provided a device for de-identifying shared data, which is applied to each data party for sharing data. The device includes: a hash calculation unit that performs a hash calculation on the privacy information in the data to be shared by using a hash algorithm to obtain a corresponding first byte array; an array truncation unit that performs a truncation process on the obtained first byte array to obtain a second byte array; and a data sending unit that sends the data to be shared including the second byte array to a data fusion device having a trusted execution environment, so that the data fusion device encrypts the second byte arrays obtained from each data party in the trusted execution environment, and fuses the data to be shared including the encrypted second byte arrays to obtain shared data.

[0009] According to another aspect of the embodiments of the present specification, there is also provided a device for de-identifying shared data, which is applied to a data fusion device with a trusted execution environment. The device includes: a data acquisition unit that acquires the data to be shared including a second byte array from each data party of the shared data, where the second byte array is obtained by truncating a first byte array obtained by each data party through hashing the privacy information in the data to be shared using a hashing algorithm; an encryption unit that encrypts each of the acquired second byte arrays in the trusted execution environment; and a data fusion unit that fuses the data to be shared including the encrypted second byte arrays in the trusted execution environment to obtain the shared data.

[0010] According to another aspect of the embodiments of the present specification, there is also provided an electronic device, including: at least one processor, a memory coupled to the at least one processor, and a computer program stored on the memory. The at least one processor executes the computer program to implement the method for de-identifying shared data as described in any one of the above.

[0011] According to another aspect of the embodiments of the present specification, there is also provided a computer-readable storage medium storing a computer program, where the computer program, when executed by a processor, implements the method for de-identifying shared data as described above.

[0012] According to another aspect of the embodiments of the present specification, there is also provided a computer program product including a computer program, where the computer program, when executed by a processor, implements the method for de-identifying shared data as described in any one of the above. Description of the Drawings

[0013] By referring to the following drawings, a further understanding of the essence and advantages of the content of the embodiments of the present specification can be achieved. In the drawings, similar components or features may have the same reference numerals.

[0014] Figure 1 FIG. shows a schematic diagram of an example of the application scenario of each data party and the data fusion device according to the embodiments of the present specification.

[0015] Figure 2 FIG. shows a flowchart of an example of the method for de-identifying shared data according to an embodiment of the present specification.

[0016] Figure 3 FIG. shows a flowchart of an example of the method for de-identifying shared data according to another embodiment of the present specification.

[0017] Figure 4A flowchart showing an example of a method for de-identifying shared data according to another embodiment of the present specification.

[0018] Figure 5 A block diagram showing an example of a shared data de-identification device according to an embodiment of the present specification.

[0019] Figure 6 A block diagram showing an example of a shared data de-identification device according to another embodiment of the present specification.

[0020] Figure 7 A block diagram showing an electronic device for implementing the method for de-identifying shared data according to an embodiment of the present specification.

[0021] Figure 8 A block diagram showing an electronic device for implementing the method for de-identifying shared data according to an embodiment of the present specification. Detailed implementation manners

[0022] The subject matter described herein will be discussed with reference to example embodiments below. It should be understood that discussing these embodiments is only to enable those skilled in the art to better understand and thus implement the subject matter described herein, and is not a limitation on the scope of protection, applicability, or examples set forth in the claims. The functions and arrangements of the elements discussed can be changed without departing from the scope of protection of the content of the embodiments of the present specification. Each example can omit, substitute, or add various processes or components as needed. Additionally, the features described relative to some examples can also be combined in other examples.

[0023] As used herein, the term "comprising" and its variants represent open terms, meaning "including but not limited to". The term "based on" means "at least partially based on". The term "one embodiment" and "an embodiment" mean "at least one embodiment". The term "another embodiment" means "at least one other embodiment". The terms "first", "second", etc. can refer to different or the same objects. Other definitions, whether explicit or implicit, may be included below. Unless explicitly specified in the context, the definition of a term is consistent throughout the specification.

[0024] A data party as a legal entity can legally collect data, including user data, operation data, etc. Each data party can use the collected data. For example, the collected data can be used as samples for model training. However, due to privacy data security issues, generally, each data party can only use its own data and cannot cross-share with other data parties.

[0025] However, only data sharing can maximize the release of data value and achieve the maximization of data utilization. To maximize the data value, data sharing can be carried out among various data parties. However, data sharing needs to ensure the security of privacy data and meet the requirements of laws and regulations. To meet the requirements, it is necessary to perform de-identification processing on privacy data before data sharing to hide privacy information and avoid the leakage of privacy information. Currently, the main de-identification means is to perform encryption processing using data encryption algorithms.

[0026] However, the current encryption processing method only uses a single encryption algorithm. General single encryption algorithms also have various degrees of disadvantages in the de-identification process, which may lead to incomplete de-identification and even the ability to restore the plaintext, resulting in the risk of privacy information leakage. For example, the MD5 hash algorithm is a common data hash algorithm and is also often used for de-identifying data. However, the ciphertext obtained by using a single MD5 hash algorithm for de-identification can be restored to the original plaintext only through a rainbow table.

[0027] In view of the above, the embodiments of the present specification provide a method and device for de-identifying shared data. In this method, at each data party for sharing data, a hash algorithm is used to perform a hash calculation on the privacy information in the data to be shared to obtain a corresponding first byte array; at each data party, the obtained first byte array is truncated to obtain a second byte array; at a data fusion device with a trusted execution environment, each second byte array obtained from each data party is encrypted in the trusted execution environment; and at the data fusion device, the data to be shared including the encrypted second byte arrays is fused in the trusted execution environment to obtain shared data. Through the embodiments of the present specification, a more secure de-identification means is provided, improving the security of privacy information in shared data.

[0028] The following will describe in detail the method and device for de-identifying shared data provided by the embodiments of the present specification with reference to the accompanying drawings.

[0029] Figure 1 FIG. shows a schematic diagram of an example of the application scenario of each data party and the data fusion device according to the embodiment of the present specification.

[0030] As Figure 1 shown, the application scenario of the de-identification solution provided by the embodiment of the present specification may include multiple data parties and at least one data fusion device. Each data party among the multiple data parties has its own data, and each data party can share all or part of the data it owns with other data parties.

[0031] The data fusion device can be a machine independent of each data party. The data fusion device is communicatively connected to each data party and can receive data from each data party. The data fusion device can include a server, a terminal device, etc. The terminal device can include a notebook, a computer, a tablet, etc.

[0032] The data fusion device can be configured with a Trusted Execution Environment (TEE), so that the processor of the data fusion device has TEE capabilities. In addition, the data fusion device can also be configured with a Rich Execution Environment (REE). The processor of the data fusion device can be shared by the TEE and the REE, or can be configured only as the TEE. In one example, the TEE in the data fusion device can be implemented based on TmstZone. The TEE is isolated from the REE, and communication can occur between the TEE and the REE.

[0033] Figure 2 The flowchart of an example 200 of a method for de-identifying shared data according to an embodiment of this specification is shown.

[0034] As Figure 2 shown, at 210, at each data party, a hash algorithm can be used to perform a hash calculation on the privacy information in the data to be shared, so as to obtain a corresponding first byte array.

[0035] In the embodiments of this specification, the hash algorithm can include at least one of algorithms such as SHA-1, SHA-2, MD5, CRC-32, etc. Among them, SHA-2 can include at least one of SHA-224, SHA-256, SHA-384, and SHA-512, etc.

[0036] In the embodiments of this specification, the first byte array can be composed of multiple values belonging to the same base. The same base can include binary, hexadecimal, etc. For example, the first byte array can be composed of multiple binary bytes, or can be composed of multiple hexadecimal values.

[0037] The first byte arrays obtained using different hash algorithms can be different. In addition, the number of values included in the first byte arrays obtained using different hash algorithms can be different. Taking binary as an example, using the CRC-32 hash algorithm outputs 32bit, so as to obtain a first byte array including 32-bit binary bytes. Using the MD5 hash algorithm outputs 128bit, so as to obtain a first byte array including 128-bit binary bytes. Using the SHA-256 hash algorithm outputs 256bit, so as to obtain a first byte array including 256-bit binary bytes.

[0038] In the embodiments of this specification, the privacy information to be de-identified may include information that can point to and identify a specific user, such as an ID number, a name, a home address, etc. After each piece of privacy information undergoes a hash calculation, a corresponding first byte array can be obtained. The first byte array and the privacy information can have a one-to-one correspondence, and each first byte array can represent the corresponding privacy information.

[0039] In one example, at each data party, the privacy information to be de-identified can be determined from the data to be shared according to at least one of the characteristics of the privacy information type, the data lineage, and the fields corresponding to the privacy information type.

[0040] Regarding the characteristics of the privacy information type, different types of privacy information have different characteristics. Each type of privacy information can have at least one characteristic, and the characteristics of the privacy information can be unique to that type of privacy information, so as to have a strong directivity for that type of privacy information. This directivity can mean that the type of privacy information can be deduced through the characteristics of the privacy information.

[0041] For example, the characteristics of a name can include the number of Chinese characters in the name. The number of Chinese characters in the name is more than two and less than six. Generally speaking, a name includes two or three Chinese characters. In addition, the characteristics of a name can also include common surnames, such as Zhao, Wang, Sun, Li, etc. For another example, the characteristics of an ID number include being composed of 18 digits, or being composed of 17 digits plus an "X" at the last digit. For another example, the characteristics of an address can include words such as "city", "district", "street", and house number.

[0042] When determining privacy information according to the characteristics of the privacy information type, the data to be shared can be scanned according to the characteristics of the privacy information type to determine whether there are relevant characteristics of each privacy information type in the data to be shared. When relevant characteristics are scanned, according to the corresponding relationship between each characteristic and the privacy information type, the privacy information type to which the scanned characteristics belong can be determined, so that the information to which the scanned characteristics belong can be determined as privacy information. For example, when it is scanned that there are 18 consecutive digits, it can be considered that these 18 digits are an ID number.

[0043] Regarding data lineage, data lineage refers to the origin and development of data, including the source of data, the data processing method, the mapping method, and the data export, etc. There is a certain correlation between two pieces of information with data lineage. For example, among two pieces of information with data lineage, the downstream information can inherit all or part of the characteristics of the upstream information.

[0044] Based on this, when determining privacy information according to data lineage, if one of multiple pieces of information with data lineage is privacy information, then it can be determined that the other information among the multiple pieces of information is also privacy information. In one example, when it is determined that a piece of information in the data to be shared is privacy information, it can be checked whether there is information in the data to be shared that has a data lineage relationship with the privacy information. If so, it can be determined that the information thus checked and found to exist is also privacy information.

[0045] Regarding the fields corresponding to privacy information types, in one example, each type of privacy information can be stored in a fixed field. For example, privacy information related to a name is stored in the "Name" field, and privacy information related to an ID number is stored in the "ID Number" field. In this example, each field can correspondingly store one type of privacy information. Based on this, when determining privacy information according to the fields corresponding to privacy information types, the fields corresponding to each privacy information type can be searched in the data to be shared, and each piece of information in the searched fields can be determined as privacy information.

[0046] In this example, after determining the privacy information to be de-identified, the hash algorithm can be used to perform hash calculation on each determined privacy information to obtain the corresponding first byte array.

[0047] At 220, at each data party, the obtained first byte array is truncated to obtain a second byte array.

[0048] In the embodiments of this specification, the number of values in the second byte array is less than the number of values in the first byte array, and the set of values included in the second byte array is a subset of the set of values included in the first byte array.

[0049] In one example of the truncation process, the first byte array can be truncated into multiple sub-byte arrays, each sub-byte array including at least one value. Then, some of the truncated sub-byte arrays can be selected to form the second byte array. The number of the selected sub-byte arrays can be random or a specified number.

[0050] The number of sub-byte arrays after truncation can be random or specified. When the number of sub-byte arrays after truncation is specified, the number of values in each sub-byte array can be randomly allocated or evenly allocated. For example, in the first byte array obtained by using the SHA256 hash algorithm, there are 256-bit binary bytes. The specified number of sub-byte arrays after truncation is 2, and each sub-byte array is evenly allocated. Then, the 256-bit binary bytes can be truncated into two sub-byte arrays each including 128-bit binary bytes, and one of the sub-byte arrays can be used as the second byte array.

[0051] In one example, the number of values in the second byte array obtained by truncation can be random or specified. The specified number is less than the number of values in the first byte array. When the number of values in the second byte array is specified, truncation can be performed according to the specified rule. In one example of truncation according to the specified rule, the specified number of values can be sequentially selected from the first byte array in the order from front to back or from back to front for truncation. In another example, a specified number of values can be selected starting from a certain position in the middle of the first byte array for truncation.

[0052] Through the truncation process, the second byte array corresponding to the privacy information is further changed based on the obtained hash value (i.e., the first byte array), preventing the privacy information from being restored according to the rainbow table, thereby avoiding rainbow table attacks and further enhancing the security of the privacy information.

[0053] At 230, each data party sends the de-identified data to be shared to the data fusion device.

[0054] The de-identification process here includes hash calculation and truncation processing. Thus, the de-identified data to be shared can be the data to be shared that has undergone hash calculation and truncation processing. Compared with the data to be shared without de-identification processing, the de-identified data to be shared does not include privacy information, but includes the second byte array corresponding to the privacy information and the non-privacy information that has not been de-identified. The data to be shared sent by each data party is the data to be shared that it owns and has undergone de-identification processing.

[0055] For example, in a set of data to be shared owned by a data party, it includes the user's name, mobile phone number, Internet access time period information, frequently visited websites, etc. Among them, the name and mobile phone number are privacy information. After de-identification processing, the privacy information generates the corresponding second byte array. The Internet access time period information and the frequently visited website information belong to non-privacy information. Then, the data to be shared sent by this data party to the data fusion device includes the second byte arrays corresponding to the name and mobile phone number, as well as the Internet access time period information and the frequently visited websites.

[0056] In one example, before sending the de-identified data to be shared to the data fusion device, each data party can also perform a first detection on the de-identified data to be shared for de-identification completeness.

[0057] In this example, the first detection for de-identification completeness is to detect whether all the privacy information in the data to be shared has been de-identified. In the data to be shared that has been de-identified, after the privacy information has been hashed and truncated, it is replaced with the corresponding second byte array.

[0058] Through the first detection of de-identification completeness, when it is detected that there is no privacy information in the data to be shared that has not been de-identified, the data to be shared can be sent to the data fusion device. When it is detected that there is privacy information in the data to be shared that has not been de-identified, the privacy information needs to be de-identified, that is, the privacy information is hashed and truncated to obtain the corresponding second byte array to avoid the leakage of privacy information. After the de-identification process is completed, the first detection for de-identification completeness can be performed on the data to be shared again until the de-identification in the data to be shared reaches completeness, that is, all the privacy information in the data to be shared has been de-identified.

[0059] In one example, the first detection for de-identification completeness may include: performing a scan detection based on the characteristics of the privacy information, and / or performing a reverse derivation detection on the de-identified data to be shared, etc.

[0060] For the detection method of scanning based on the characteristics of the privacy information type, at each data party, the data to be shared to be sent can be scanned according to the characteristics of various privacy information types to detect whether there are characteristics corresponding to various privacy information types in the data to be shared. If it is detected that there are characteristics corresponding to the privacy information type, it can be determined that there is privacy information in the data to be shared that has not been de-identified.

[0061] In one example, customized scanning can be implemented, that is, scanning is performed on the characteristics of the specified privacy information type, and the specified privacy information type can include one or more. Through customized scanning, targeted scanning can be carried out, which can improve the efficiency of completeness detection.

[0062] For the detection method of reverse derivation, the object of the reverse derivation detection is the de-identified data to be shared to be sent. The reverse derivation detection is to detect whether a specific user can be reverse-located based on the de-identified data to be shared. When a specific user can be located through the reverse derivation detection, it means that there is privacy information in the data to be shared that can point to a specific user, such as name, mobile phone number, ID number, etc.

[0063] In a reverse derivation calculation, the de-identified data to be shared can be arranged according to the data distribution mode of users and their relevant characteristic information. The relevant characteristic information can include privacy information and non-privacy information, and each type of privacy information can be used as a piece of relevant characteristic information. For example, in the data distribution based on a data table, each data row corresponds to a user, and each data column corresponds to a relevant characteristic information of the user. Then, based on the data distribution of the data to be shared, it is detected whether the de-identification is complete. Specifically, when the number of columns for each user in the data to be shared is the same as the number of columns before the de-identification process, it can be considered that there is still privacy information in the data to be shared that has not been de-identified. For example, the relevant characteristic information for a user includes name, mobile phone number, ID number, education level, occupation, etc., that is, the characteristic information includes 5 columns, and after the de-identification process, the user still corresponds to 5 columns of characteristic information, then it can be determined that there is privacy information for this user.

[0064] At 240, at the data fusion device, each of the second byte arrays obtained from each data party can be encrypted in the trusted execution environment.

[0065] In one example, after the data fusion device receives the de-identified data to be shared from each data party and before encrypting each of the second byte arrays, it can perform a second detection on the received data to be shared for de-identification completeness. The methods of the second detection can include scanning detection according to the characteristics of the privacy information type, and / or performing reverse derivation detection on the received data to be shared, etc.

[0066] In this example, the method used for the second detection can be different from the method used for the first detection. For example, if the method used for the first detection is scanning detection according to the characteristics of the privacy information type, then the method used for the second detection is to perform reverse derivation calculation on the received data to be shared.

[0067] Before encrypting the received data to be shared, the data fusion device performs a second detection on the received data to be shared again for de-identification completeness, further ensuring that the privacy information in the data to be shared has been de-identified, thereby further improving the security of the privacy information. In addition, the second detection and the first detection use different detection methods. By cross-checking the data to be shared with two different detection methods for multiple de-identification completeness detections, the accuracy of the detection for de-identification completeness can be further improved.

[0068] In the embodiments of this specification, the encryption process for each second byte array is a process of strengthening the security of the de-identification of private information. The encryption algorithms used in the encryption process may include at least one of AES, RSA, DES, national encryption algorithms, RC2, and RC4, etc. The same encryption algorithm can be used for encrypting each second byte array, or different encryption algorithms can be used. In an example of using different encryption algorithms, different encryption algorithms can be used for different data parties. By using different encryption algorithms, the second byte array corresponding to the private information can be further secured, the security of the second byte array can be improved, and thus the security of the private information can be improved.

[0069] The encryption algorithm code and the key used in the encryption process are both stored in the trusted execution environment of the data fusion device. The encryption algorithm code runs in the trusted execution environment and, in combination with the key in the trusted execution environment, performs the encryption operation in the trusted execution environment. As an independent operating environment, the trusted execution environment can ensure the confidentiality and integrity of the encryption algorithm code and the key in the trusted execution environment, as well as the security of the encryption operation during the execution process.

[0070] In an example, the key and the encryption algorithm code used in the encryption process are stored separately in the trusted execution environment. In the trusted execution environment, the key and the encryption algorithm code can be stored in different storage spaces, and different applications can also be used to manage the stored key and encryption algorithm code. In an example, the algorithm hosting application can manage the encryption algorithm code, and the key hosting application can manage the key stored in the trusted execution environment. The algorithm hosting application and the key hosting application belong to different applications in the trusted execution environment.

[0071] In an example, in the trusted execution environment of the data fusion device, when an encryption operation needs to be performed, the algorithm hosting application can obtain the key from the key hosting application. In an example of key acquisition, the algorithm hosting application can send a key request to the key hosting application. The key hosting application obtains the corresponding key from the storage space according to the key request and sends the key to the algorithm hosting application.

[0072] In an example, in response to the key request, the key hosting application can verify the algorithm hosting application that sends the key request. If the verification passes, the key can be sent to the algorithm hosting application. When the verification passes, it can indicate that the algorithm hosting application is trusted. The verification methods can include at least one of authentication, password verification, signature verification, etc.

[0073] When authentication is adopted, the identity information of the algorithm hosting application may be included in the key request sent by the algorithm hosting application. When the key hosting application receives the key request, it verifies the identity information in the key request to determine that the algorithm hosting application sending the key request is trustworthy. In the case of authentication, the key hosting application may provide key services only for specified applications, so that only the specified applications can obtain keys from the key hosting application to perform encryption operations, improving the credibility and security of the encryption process.

[0074] When password authentication is adopted, the algorithm hosting application may encrypt the key request using a private key. When the key hosting application receives the key request, the key hosting application may decrypt the key request using the public key corresponding to the private key. When the decryption is successful, it can be determined that the algorithm hosting application sending the key request is trustworthy. When signature authentication is adopted, the algorithm hosting application may sign the key request using a private key. When the key hosting application receives the key request, the key hosting application may perform signature verification on the key request using the public key corresponding to the private key. When the signature verification passes, it can be determined that the algorithm hosting application sending the key request is trustworthy.

[0075] After obtaining the key, the algorithm hosting application may perform encryption operations based on the obtained key and the encryption algorithm code to encrypt each of the second byte arrays obtained from each data party respectively.

[0076] By storing the key and the encryption algorithm code separately, the security of the encryption process is improved. By verifying the algorithm hosting application through the key hosting application, it is avoided that the key is obtained and used by other illegal applications, further improving the credibility and security of the encryption process.

[0077] In one example, data sharing can be applied to different application scenarios, such as machine model training, federated learning, user behavior analysis, etc. The keys used in different application scenarios for data sharing may be different, thus avoiding using the same key for encryption processing and improving the robustness of the encryption algorithm.

[0078] In this example, different importance levels can be set for each application scenario. For application scenarios with a low importance level, keys with a relatively short character length can be used. For application scenarios with a high importance level, keys with a longer character length can be used.

[0079] Back to Figure 2 , at 250, at the data fusion device, in the trusted execution environment, the data to be shared including each encrypted second byte array can be data-fused to obtain the shared data.

[0080] The obtained shared data does not include privacy information, but only includes the encrypted second byte array generated corresponding to the privacy information and non-privacy information. After obtaining the shared data, the data fusion device can distribute the shared data to each data party to achieve data sharing among the data parties.

[0081] Figure 3 Fig. 4 shows a flowchart of an example 300 of a method for de-identifying shared data according to another embodiment of this specification. Figure 3 The method shown can be executed by each data party for shared data.

[0082] As Figure 3 shown, at 310, a hash algorithm is used to perform a hash calculation on the privacy information in the data to be shared to obtain the corresponding first byte array.

[0083] In one example, according to at least one of the characteristics of the privacy information type, data lineage, and the fields corresponding to the privacy information type, the privacy information to be de-identified is determined from the data to be shared; and the determined privacy information is hashed using a hash algorithm to obtain the corresponding first byte array.

[0084] At 320, the obtained first byte array is truncated to obtain a second byte array.

[0085] At 330, the data to be shared including the second byte array is sent to a data fusion device with a trusted execution environment, so that the data fusion device encrypts the second byte arrays obtained from each data party in the trusted execution environment, and fuses the data to be shared including the encrypted second byte arrays to obtain shared data.

[0086] In one example, before sending the data to be shared including the second byte array to the data fusion device, a first detection for de-identification completeness can also be performed on the de-identified data to be shared.

[0087] In one example, the first detection for de-identification completeness can include: scanning according to the characteristics of the privacy information type, and / or performing reverse derivation calculation on the de-identified data to be shared.

[0088] Figure 4 Fig. 5 shows a flowchart of an example 400 of a method for de-identifying shared data according to another embodiment of this specification. Figure 4 The method shown can be executed by a data fusion device with a trusted execution environment.

[0089] At 410, obtain the data to be shared including the second byte array from each data party of the shared data. The second byte array is obtained by truncating the first byte array obtained by each data party through hashing the privacy information in the data to be shared using a hashing algorithm.

[0090] In one example, a second detection for de-identification completeness can be performed on the received data to be shared. In one example, the second detection for de-identification completeness can include: performing a scan detection according to the characteristics of the privacy information type, and / or performing a reverse derivation detection on the de-identified data to be shared.

[0091] At 420, perform encryption processing on each of the obtained second byte arrays in a trusted execution environment.

[0092] At 430, perform data fusion on the data to be shared including the encrypted second byte arrays to obtain the shared data in a trusted execution environment.

[0093] In one example, the key and the encryption algorithm code used for the encryption processing are separately stored in the trusted execution environment.

[0094] In one example, in the trusted execution environment, the algorithm hosting application obtains the key from the key escrow application for managing the key; and the algorithm hosting application performs an encryption operation based on the obtained key and the encryption algorithm code to perform encryption processing on each of the second byte arrays obtained from each data party.

[0095] In one example, in the trusted execution environment, the algorithm hosting application sends a key request to the key escrow application; and when the key escrow application verifies that the algorithm hosting application is trusted, it sends the key to the algorithm hosting application.

[0096] Figure 5 A block diagram of an example of a shared data de-identification device 500 according to an embodiment of the present specification is shown. The shared data de-identification device 500 can be applied to each data party for sharing data.

[0097] As Figure 5 shown, the shared data de-identification device 500 includes a hashing calculation unit 510, an array truncation unit 520, and a data sending unit 530.

[0098] The hashing calculation unit 510 can be configured to perform hashing calculation on the privacy information in the data to be shared using a hashing algorithm to obtain the corresponding first byte array.

[0099] In one example, the hash calculation unit 510 may also be configured to: determine, from the data to be shared, the privacy information to be de-identified based on at least one of the characteristics of the privacy information type, the data lineage, and the fields corresponding to the privacy information type; and perform a hash calculation on the determined privacy information using a hash algorithm to obtain a corresponding first byte array.

[0100] The array truncation unit 520 may be configured to perform a truncation process on the obtained first byte array to obtain a second byte array.

[0101] The data sending unit 530 may be configured to send the data to be shared including the second byte array to a data fusion device having a trusted execution environment, so that the data fusion device encrypts the second byte arrays obtained from each data party in the trusted execution environment, and performs data fusion on the data to be shared including the encrypted second byte arrays to obtain shared data.

[0102] In one example, the shared data de-identification device 500 may further include a completeness detection unit, which is configured to perform a first detection on the data to be shared that has undergone de-identification processing for de-identification completeness.

[0103] Figure 6 FIG. shows a block diagram of an example of a shared data de-identification device 600 according to another embodiment of the present specification. The shared data de-identification device 600 may be applied to a data fusion device having a trusted execution environment.

[0104] As Figure 6 shown, the shared data de-identification device 600 includes: a data acquisition unit 610, an encryption unit 620, and a data fusion unit 630.

[0105] The data acquisition unit 610 may be configured to acquire, from each data party of the shared data, the data to be shared including the second byte array, where the second byte array is obtained by performing a truncation process on the first byte array obtained by each data party performing a hash calculation on the privacy information in the data to be shared using a hash algorithm.

[0106] The encryption unit 620 may be configured to perform an encryption process on each of the acquired second byte arrays in the trusted execution environment.

[0107] In one example, the encryption unit 620 may also be configured to: in the trusted execution environment, the algorithm hosting application obtains a key from the key hosting application for managing keys; and the algorithm hosting application performs an encryption operation based on the obtained key and the encryption algorithm code to perform an encryption process on each of the second byte arrays obtained from each data party.

[0108] In one example, the encryption unit 620 may also be configured to: in a trusted execution environment at the data fusion device, the algorithm hosting application sends a key request to the key hosting application; and when the key hosting application verifies that the algorithm hosting application is trusted, the key hosting application sends the key to the algorithm hosting application.

[0109] The data fusion unit 630 may be configured to perform data fusion on the data to be shared including each encrypted second byte array in the trusted execution environment to obtain the shared data.

[0110] In one example, the shared data de-identification device 600 may further include a completeness detection unit, which is configured to perform a second detection on the received data to be shared for de-identification completeness. The second detection for de-identification completeness may include: performing a scan detection according to the characteristics of the privacy information type, and / or performing a reverse derivation detection on the data to be shared after de-identification processing.

[0111] The above references Figures 1 to 6 describe the embodiments of the method and device for de-identifying shared data according to the embodiments of the present specification.

[0112] The device for de-identifying shared data according to the embodiments of the present specification may be implemented in hardware, or may be implemented by software or a combination of hardware and software. Taking software implementation as an example, as a logically meaningful device, it is formed by the processor of the device where it is located reading the corresponding computer program instructions in the memory into the memory for operation. In the embodiments of the present specification, the device for de-identifying shared data may be implemented by using, for example, an electronic device.

[0113] Figure 7 The block diagram of an electronic device 700 for implementing the method for de-identifying shared data according to the embodiments of the present specification is shown.

[0114] As Figure 7 shown, the electronic device 700 may include at least one processor 710, a memory (e.g., a non-volatile memory) 720, a memory 730, and a communication interface 740, and at least one processor 710, the memory 720, the memory 730, and the communication interface 740 are connected together via a bus 750. At least one processor 710 executes at least one computer-readable instruction stored or encoded in the memory (i.e., the above elements implemented in software form).

[0115] In one embodiment, computer-executable instructions are stored in a memory, which when executed cause at least one processor 710 to: perform a hashing calculation on privacy information in the data to be shared using a hashing algorithm to obtain a corresponding first byte array; perform a truncation process on the obtained first byte array to obtain a second byte array; and send the data to be shared including the second byte array to a data fusion device having a trusted execution environment, so that the data fusion device encrypts the second byte arrays obtained from each data party in the trusted execution environment, and performs data fusion on the data to be shared including the encrypted second byte arrays from each data party to obtain shared data.

[0116] It should be understood that the computer-executable instructions stored in the memory, when executed, cause at least one processor 710 to perform the various operations and functions described above in the respective embodiments of this specification in combination with Figures 1 - 6 the description.

[0117] According to one embodiment, a program product such as a machine-readable medium is provided. The machine-readable medium may have instructions (i.e., the elements implemented in software as described above), which when executed by the machine cause the machine to perform the various operations and functions described above in the respective embodiments of this specification in combination with Figures 1 - 6 the description.

[0118] Figure 8 FIG. shows a block diagram of an electronic device 800 for implementing a method for de-identifying shared data according to an embodiment of this specification.

[0119] As Figure 8 shown, the electronic device 800 may include at least one processor 810, a memory (e.g., a non-volatile memory) 820, a memory 830, and a communication interface 840, and at least one processor 810, the memory 820, the memory 830, and the communication interface 840 are connected together via a bus 850. At least one processor 810 executes at least one computer-readable instruction stored or encoded in the memory (i.e., the elements implemented in software as described above).

[0120] In one embodiment, computer-executable instructions are stored in a memory, which when executed cause at least one processor 810 to: obtain data to be shared including a second byte array from each data party of the shared data, where the second byte array is obtained by each data party performing a truncation process on a first byte array obtained by performing a hashing calculation on privacy information in the data to be shared using a hashing algorithm; perform an encryption process on each of the obtained second byte arrays in a trusted execution environment; and perform data fusion on the data to be shared including the encrypted second byte arrays from each data party in the trusted execution environment to obtain shared data.

[0121] It should be understood that the computer-executable instructions stored in the memory, when executed, cause at least one processor 810 to perform the various operations and functions described above in the respective embodiments of this specification in conjunction with Figures 1 - 6 the descriptions.

[0122] According to one embodiment, a program product such as a machine-readable medium is provided. The machine-readable medium may have instructions (i.e., the elements implemented in software as described above), which, when executed by the machine, cause the machine to perform the various operations and functions described above in the respective embodiments of this specification in conjunction with Figures 1 - 6 the descriptions.

[0123] Specifically, a system or device equipped with a readable storage medium may be provided, on which software program code for implementing the functions of any one of the above embodiments is stored, and the computer or processor of the system or device is caused to read and execute the instructions stored in the readable storage medium.

[0124] In this case, the program code read from the readable medium itself can implement the functions of any one of the above embodiments, so the machine-readable code and the readable storage medium storing the machine-readable code constitute a part of the present invention.

[0125] The computer program code required for the operations of each part of this specification can be written in any one or more programming languages, including object-oriented programming languages such as Java, Scala, Smalltalk, Eiffel, JADE, Emerald, C++, C#, VB, NET, and Python, conventional procedural programming languages such as C, Visual Basic 2003, Perl, COBOL2002, PHP, and ABAP, dynamic programming languages such as Python, Ruby, and Groovy, or other programming languages. The program code can run on the user's computer, or run on the user's computer as an independent software package, or part of it runs on the user's computer and another part runs on a remote computer, or all of it runs on a remote computer or server. In the latter case, the remote computer can be connected to the user's computer in any network form, such as a local area network (LAN) or a wide area network (WAN), or connected to an external computer (e.g., via the Internet), or in a cloud computing environment, or used as a service, such as software as a service (SaaS).

[0126] Examples of readable storage media include floppy disks, hard disks, magneto-optical disks, optical disks (such as CD-ROM, CD-R, CD-RW, DVD-ROM, DVD-RAM, DVD-RW, DVD-RW), magnetic tapes, non-volatile memory cards, and ROM. Optionally, program code can be downloaded from a server computer or the cloud via a communication network.

[0127] The above describes specific embodiments of this specification. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims can be performed in a different order than in the embodiments and still achieve the desired result. Additionally, the processes depicted in the figures do not necessarily require the specific order or sequential order shown to achieve the desired result. In certain embodiments, multitasking and parallel processing are also possible or may be advantageous.

[0128] Not all steps and units in the above process flows and system structure diagrams are necessary, and some steps or units can be ignored according to actual needs. The execution order of the steps is not fixed and can be determined as needed. The device structures described in the above embodiments can be physical structures or logical structures, that is, some units may be implemented by the same physical entity, or some units may be implemented by multiple physical entities separately, or some components in multiple independent devices may be jointly implemented.

[0129] The term "exemplary" used throughout this specification means "serving as an example, instance, or illustration" and does not mean "preferred" or "advantageous" over other embodiments. For the purpose of providing an understanding of the described technology, the specific embodiments include specific details. However, these technologies can be implemented without these specific details. In some instances, well-known structures and devices are shown in block diagram form to avoid obscuring the concepts of the described embodiments.

[0130] The optional embodiments of the embodiments of this specification have been described in detail above in conjunction with the accompanying drawings. However, the embodiments of this specification are not limited to the specific details in the above embodiments. Within the technical concept scope of the embodiments of this specification, various simple modifications can be made to the technical solutions of the embodiments of this specification, and these simple modifications all fall within the protection scope of the embodiments of this specification.

[0131] The foregoing description of the content of this specification is provided to enable any person of ordinary skill in the art to make or use the content of this specification. Various modifications to the content of this specification will be apparent to those of ordinary skill in the art, and the general principles defined herein can also be applied to other variations without departing from the scope of protection of the content of this specification. Therefore, the content of this specification is not limited to the examples and designs described herein, but is consistent with the broadest scope that conforms to the principles and novel features disclosed herein.

Claims

1. A method for de-identifying shared data, comprising: At each data party for sharing data, using a hashing algorithm to perform hashing calculation on the privacy information in the data to be shared, so as to obtain a corresponding first byte array; At each of the data parties, truncating the obtained first byte array to obtain a second byte array and sending the de-identified data to be shared to a data fusion device with a trusted execution environment, where the de-identified data to be shared includes the second byte array and the non-privacy information in the data to be shared; At the data fusion device with a trusted execution environment, performing encryption processing on each of the second byte arrays obtained from the respective data parties in the trusted execution environment; And At the data fusion device, performing data fusion on the data to be shared including the encrypted second byte arrays in the trusted execution environment to obtain shared data, Wherein, before each of the data parties sends the de-identified data to be shared to the data fusion device, the method further includes: Performing a first detection on the de-identified data to be shared for de-identification completeness, where the first detection includes: performing a reverse derivation detection on the de-identified data to be shared, Performing a reverse derivation detection on the de-identified data to be shared includes: Arranging the de-identified data to be shared according to the data distribution mode of users and the relevant characteristic information of users; and Detecting whether the de-identification completeness is achieved based on the data distribution of the de-identified data to be shared.

2. The method according to claim 1, wherein At each data party for sharing data, using a hashing algorithm to perform hashing calculation on the privacy information in the data to be shared to obtain a corresponding first byte array includes: At each data party for sharing data, determining the privacy information to be de-identified from the data to be shared according to at least one of the characteristics of the privacy information type, data lineage, and the fields corresponding to the privacy information type; and Using a hashing algorithm to perform hashing calculation on the determined privacy information to obtain a corresponding first byte array.

3. The method according to claim 1, wherein, The first detection further includes: performing a scan detection according to the characteristics of the privacy information type.

4. The method according to claim 1, wherein After the data fusion device receives the de-identified data to be shared from the respective data parties, the method further includes: Performing a second detection on the received data to be shared for de-identification completeness, where the manner used in the second detection is different from the manner used in the first detection.

5. The method according to claim 1, wherein The key and encryption algorithm code used in the encryption processing are separately stored in the trusted execution environment.

6. The method according to claim 5, wherein, At the data fusion device with a trusted execution environment, performing encryption processing on each of the second byte arrays obtained from the respective data parties in the trusted execution environment includes: In the trusted execution environment of the data fusion device, the algorithm hosting application obtains a key from the key hosting application for managing keys; and The algorithm hosting application performs an encryption operation based on the obtained key and the encryption algorithm code to encrypt each of the second byte arrays obtained from the respective data parties.

7. The method according to claim 6, wherein In the trusted execution environment of the data fusion device, the algorithm hosting application obtaining a key from the key hosting application for managing keys includes: In the trusted execution environment of the data fusion device, the algorithm hosting application sends a key request to the key hosting application; and When verifying the trustworthiness of the algorithm hosting application, the key hosting application sends the key to the algorithm hosting application.

8. The method according to claim 5, wherein, The keys used for data sharing are different in different application scenarios.

9. A method for de-identifying shared data, which is executed by each data party for sharing data, the method including: Performing a hash calculation on the privacy information in the data to be shared by using a hash algorithm to obtain a corresponding first byte array; Performing a truncation process on the obtained first byte array to obtain a second byte array; Performing a first detection on the de-identified data to be shared for de-identification completeness, the first detection including: performing a reverse derivation detection on the de-identified data to be shared, the de-identified data to be shared including the second byte array and the non-privacy information in the data to be shared; and Sending the de-identified data to be shared to a data fusion device having a trusted execution environment, so that the data fusion device encrypts the second byte arrays obtained from each data party in the trusted execution environment, and fuses the data to be shared including the encrypted second byte arrays to obtain shared data, Performing a reverse derivation detection on the de-identified data to be shared includes: Arranging the de-identified data to be shared according to the data distribution manner of the user and the relevant characteristic information of the user; and Detecting whether the de-identification completeness is achieved based on the data distribution of the de-identified data to be shared.

10. A method for de-identifying shared data, which is executed by a data fusion device having a trusted execution environment, the method including: Obtaining de-identified data to be shared from each data party of the shared data, the de-identified data to be shared including a second byte array and the non-privacy information in the data to be shared, the second byte array being obtained by each data party through performing a truncation process on a first byte array, the first byte array being obtained by performing a hash calculation on the privacy information in the data to be shared by using a hash algorithm, and the obtained de-identified data to be shared passes a first detection for de-identification completeness at each data party; Performing an encryption process on each of the obtained second byte arrays in the trusted execution environment; And Fusing the data to be shared including the encrypted second byte arrays in the trusted execution environment to obtain shared data, Among them, the first detection includes: performing reverse derivation detection on the de-identified data to be shared, Performing reverse derivation detection on the de-identified data to be shared includes: Arranging the de-identified data to be shared according to the data distribution mode of users and their relevant characteristic information; and Detecting whether the completeness of the de-identification process is achieved based on the data distribution of the de-identified data to be shared.

11. A device for de-identifying shared data, applied to each data party for sharing data, the device includes: A hash calculation unit that uses a hash algorithm to calculate the hash of the privacy information in the data to be shared to obtain a corresponding first byte array; An array truncation unit that truncates the obtained first byte array to obtain a second byte array; A completeness detection unit that performs a first detection on the de-identified data to be shared for de-identification completeness, where the de-identified data to be shared includes the second byte array and the non-privacy information in the data to be shared; And A data sending unit that sends the de-identified data to be shared that has passed the first detection to a data fusion device with a trusted execution environment, so that the data fusion device encrypts the second byte arrays obtained from each data party in the trusted execution environment, and fuses the data to be shared including the encrypted second byte arrays to obtain shared data, Among them, the first detection includes: performing reverse derivation detection on the de-identified data to be shared, Performing reverse derivation detection on the de-identified data to be shared includes: Arranging the de-identified data to be shared according to the data distribution mode of users and their relevant characteristic information; and Detecting whether the completeness of the de-identification process is achieved based on the data distribution of the de-identified data to be shared.

12. A device for de-identifying shared data, applied to a data fusion device with a trusted execution environment, the device includes: A data acquisition unit that acquires de-identified data to be shared from each data party of the shared data, where the de-identified data to be shared includes a second byte array and the non-privacy information in the data to be shared, the second byte array is obtained by truncating a first byte array, and the first byte data is obtained by each data party using a hash algorithm to calculate the hash of the privacy information in the data to be shared, and the acquired de-identified data to be shared passes the first detection for de-identification completeness at each data party; An encryption unit that encrypts each of the acquired second byte arrays in the trusted execution environment; And A data fusion unit that fuses the data to be shared including the encrypted second byte arrays in the trusted execution environment to obtain shared data, Among them, the first detection includes: performing reverse derivation detection on the to-be-shared data that has undergone de-identification processing, Performing reverse derivation detection on the to-be-shared data that has undergone de-identification processing includes: Arranging the to-be-shared data that has undergone de-identification processing according to the data distribution mode of users and their relevant characteristic information; and Detecting whether the completeness of the de-identification processing is achieved based on the data distribution of the to-be-shared data that has undergone de-identification processing.

13. An electronic device, comprising: At least one processor, a memory coupled to the at least one processor, and a computer program stored on the memory, the at least one processor executing the computer program to implement the method according to claim 9 or 10.

14. A computer-readable storage medium storing a computer program, the computer program, when executed by a processor, implementing the method according to claim 9 or 10.

15. A computer program product including a computer program, the computer program, when executed by a processor, implementing the method according to claim 9 or 10.

Citation Information

Patent Citations

  • Key authorization method and system

    CN111090865A

  • Implementation method of medical data sharing model based on block chain and IPFS

    CN111832038A

  • Service data processing method and device and electronic equipment

    CN113282959A

  • Data sharing method and device, equipment and system

    CN113434888A

  • Privacy data processing method and device

    CN113672977A