An Auditable Data Linking Method
Improve existing technology by splitting the Bloom filter (SBF), solving the problem of privacy leakage and malicious collusion in privacy protection data links, achieving lower privacy leakage risks and stronger privacy protection.
Patent Information
- Application Number
- CN202111226178.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-10-21
- Publication Date
- 2025-07-18
- Estimated Expiration
- 2041-10-21
AI Technical Summary
In the existing privacy-protected data linking scheme, blockchains have the risk of privacy leakage when copying data between untrusted parties, and the possibility of collusion between malicious parties is high.
Using split Bloom filter (SBF), the original Bloom filter is divided into s parts, and only a small part of iterative similarity calculation is used for reducing the amount of information sharing, and an auditable data linking process is performed through semi-trusted third-party STTP.
It reduces the risk of privacy leakage, reduces the possibility of collusion between malicious parties, and improves the privacy protection capabilities of the PPRL process.
Smart Images

Figure CN114117465B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the fields of information security technology and data link privacy protection. Background Art
[0002] Currently, in some PPRL (Privacy-Preserving Record Linkage) schemes, blockchain (smart contract) is used as a semi-trusted third party (STTP). However, blockchain does not provide a mechanism to protect entity privacy during the PPRL process. In fact, blockchain can cause privacy leakage by replicating the entire data among untrusted parties.
[0003] Chinese Patent (Publication No. CN110609831A) discloses a data link method based on privacy protection and secure multi-party computation. This invention uses an improved k-means classification method to partition local data, reducing the number of comparisons between data records. It has good scalability for large databases and also improves the execution efficiency of privacy-protected record linkage. By using the properties of reversible matrices and Shamir's threshold secret sharing scheme, this invention ensures good security when comparing similarities between two or more record-level Bloom filters, preventing adversaries from obtaining users' sensitive information. This invention has good scalability and relatively low computational overhead, and is applicable to real-world environments with a large amount of real data.
[0004] Chinese Patent (Publication No. 105138927A) discloses a privacy data protection method, which includes: a data access platform receives a key access request sent by a client, and the key access request includes a user identification number and a privacy data name; the data access platform obtains a first key corresponding to the user identification number and corresponding to the privacy data name according to the key access request. The data access platform pre-stores a first correspondence table, and the first correspondence table includes multiple privacy data names corresponding to each user identification number and a first key uniquely corresponding to each privacy data name; the data access platform obtains the privacy data corresponding to the privacy data name according to the first key. The data access platform pre-stores a second correspondence table, and the second correspondence table includes the privacy data of the user identified by the user identification number and corresponding to each privacy data name, and the first key uniquely corresponding to the privacy data; the data access platform generates a second key based on the user identification number, the privacy data name, and the current timestamp, and replaces the first key in the first correspondence table and the second correspondence table with the second key. Summary of the Invention
[0005] In order to overcome the deficiencies of the prior art, the present invention proposes an auditable data link method.
[0006] The method of the present invention is specifically as follows:
[0007] Step 1. The system generates the parameters of the Bloom filter BF, the parameters of the split Bloom filter SBF, the number s of the split functions split(s), the value of the similarity error error, and the thresholds α, β, where β = α - error.
[0008] Step 2. Each party anonymizes the entities and randomly generates a unique ID for each entity.
[0009] Step 3. Each party sends the split information to the semi-trusted third party STTP.
[0010] Step 4. The semi-trusted third party STTP calculates the similarity between each party.
[0011] Step 5. The semi-trusted third party STTP publishes the list ζ, which consists of the IDs and the similarities between the entities.
[0012] Step 6. For the entities stored in the list ζ, the two parties take turns to exchange the split information, one split information at a time, and each participant finally receives copies of the split information.
[0013] Step 7. The semi-trusted third party STTP uses the split information sent by each party as input and calculates the similarity between the entities again. The number s of the split functions split(s) is considered in this calculation;
[0014] If the difference between the similarity obtained in this calculation and the similarity obtained in Step 4 is greater than error, an abnormal behavior is detected and the execution is terminated.
[0015] Step 8. The two parties exchange the similarities calculated in the previous step to update the overall similarities of the entities in the list ζ.
[0016] Step 9. Each party checks the difference between the value of the exchanged similarity and the similarity value stored in the list ζ. If the difference is greater than error, an error is detected and the execution is terminated.
[0017] Step 10. The semi-trusted third party STTP selects the entities with similarities higher than the α threshold and publishes the results.
[0018] The present invention improves the most common anonymization technology used in PPRL: the Bloom filter (BF) is improved to the "split Bloom filter" (SBF), enabling the present invention to provide a lower risk of privacy leakage. In addition, the possibility of collusion between malicious parties is reduced. The present invention provides stronger privacy guarantees by reducing the amount of information shared between PPRL parties. Description of the Drawings
[0019] Figure 1 These are the three stages of the present invention. Detailed implementation manners
[0020] As Figure 1 shown, the present invention performs comparative implementation in three stages to improve the privacy protection ability of PPRL. To execute the comparison step of PPRL, all parties need to share their entire anonymized entities, which is conducive to complex cryptographic analysis attacks (e.g., pattern mining attacks). The present invention designs a new Bloom filter, called Split Bloom Filter (SBF), to achieve auditable data linking.
[0021] The basic idea of SBF is to use only a small part (instead of the whole BF) of the original BF for iterative similarity calculation to reduce the amount of information shared in the privacy protection data linking comparison step. SBF divides the original BF into s parts, where each divided part is a small part of the length of the original BF.
[0022] Based on the above concept, the present embodiment includes the following steps:
[0023] System initialization stage
[0024] Step 1. The system generates BF parameters, SBF parameters, the number s of the split function split(s), the value of the similarity error error, and thresholds α, β, where β = α - error.
[0025] Step 2. All parties anonymize the entities and randomly generate unique IDs for each entity;
[0026] Step 3. All parties send the split information to STTP, where:
[0027]
[0028] l is the byte of the original BF, is the set of p participating parties, e t is the set of anonymized entities.
[0029] Similarity calculation stage
[0030] Step 4. The semi-trusted third party STTP uses to calculate the similarity between all parties.
[0031] Where are respectively the i-th split of, and ε is the set error rate. respectively represent two different anonymized entities.
[0032] Used to compare the similarity between finite sample sets. The larger the value, the higher the similarity.
[0033] Step 5. The semi-trusted third party STTP publishes the list ζ, which consists of IDs and the similarities between each entity.
[0034] In Step 4, the semi-trusted third party calculates the similarity according to the original BF formula. In the fifth step, the list ζ of relevant entities is disclosed to other parties for entities with a very high similarity.
[0035] In the similarity calculation stage, the STTP calculates the similarity of all received entity pairs. Then, the STTP publishes the list of entity pairs with a similarity value greater than β. The threshold β must be carefully selected. By choosing a lower threshold β, the number of entities forwarded to the next stage increases, reducing the probability of a successful cryptanalysis attack.
[0036] Step 6. For the entities stored in the list ζ, the two parties take turns to exchange the split information, one piece of split information at a time. Finally, each participant receives copies of the split information.
[0037] Step 7. The STTP uses the split information sent by each party as input and uses the improved formula to calculate the similarity between entities;
[0038] If the difference between the similarity calculated this time and the value calculated by the STTP in Step 4 is greater than error, an abnormal behavior is detected and the execution is terminated.
[0039] This embodiment provides audibility for the similarity calculation performed by each party in PPRL using this improved formula. In addition, only a small part (instead of the whole BF) of the original BF is used for iterative similarity calculation, making it difficult for the STTP to perform a cryptanalysis attack.
[0040] Result announcement stage
[0041] Step 8. The two parties exchange the similarities calculated in the previous step to update the overall similarity of the entities in ζ.
[0042] Step 9. Each party checks the difference between the exchanged similarity value and the value stored in ζ. If this difference is greater than error, an error is detected and the execution is aborted.
[0043] Step 10. Finally, the STTP selects the entities with a similarity higher than the α threshold and announces the result.
Claims
1. An auditable data link method, characterized in that The method includes the following steps: Step 1. The system generates the parameters of the Bloom filter BF, the parameters of the split Bloom filter SBF, the number s of the split functions split(s), the value of the similarity error error, and the thresholds α, β, where β = α - error; Step 2. Each party anonymizes the entities and randomly generates a unique ID for each entity; Step 3. Each party sends the split information to the semi-trusted third party STTP; Step 4. The semi-trusted third party STTP calculates the similarity between the parties; Step 5. The semi-trusted third party STTP publishes the list ζ, which consists of the IDs and the similarities between the entities; Step 6. For the entities stored in list ζ, the two parties take turns to exchange the segmentation information, one piece of segmentation information at a time, and each participant finally receives pieces of segmentation information; Step 7. The semi-trusted third party STTP uses the split information sent by each party as input and calculates the similarity between the entities again, and the number s of the split functions split(s) is considered in this calculation; If the difference between the similarity obtained from this calculation and the similarity obtained from the calculation in Step 4 is greater than error, an abnormal behavior is detected and the execution is terminated; Step 8. The two parties exchange the similarities calculated in the previous step to update the overall similarity of the entities in the list ζ; Step 9. Each party checks the difference between the value of the exchanged similarity and the similarity value stored in the list ζ. If the difference is greater than error, an error is detected and the execution is terminated; Step 10. The semi-trusted third party STTP selects the entities with similarities higher than the α threshold and publishes the results.
2. The audit-able data link method according to claim 1, wherein: The segmentation information described in Step 3 is Wherein: SBF(e t ,s) = [φ 0 ,..., φ s-1 , l is the byte of the original BF, is the set of p parties, e t is the set of anonymous entities.
3. The auditable data link method according to claim 1, wherein: The formula for the semi-trusted third party STTP to calculate the similarity between the parties in Step 4 is as follows: respectively the i-th segmentation of, where ε is the error rate, respectively represent two different anonymous entities, used to compare the similarity between finite sample sets.
4. A auditable data link method according to claim 1, wherein: The similarity value of the entity pairs in the list ζ in Step 5 is greater than β.
5. The auditable data link method according to claim 3, wherein: The improved formula is used to calculate the similarity between the entities again in Step 7:
Citation Information
Patent Citations
Privacy data protection method and apparatus
CN105138927A
Data link method based on privacy protection and secure multi-party computing
CN110609831A
Method for increasing and canceling elements of Bloom filter and Bloom filter
CN101923568A
Multi-keyword ciphertext retrieval method based on iterative encryption in cloud environment
CN109213731A