Privacy processing method for non-interactive Hamming distance
By implementing outsourcing computing on cloud servers and using technologies such as blinding, hashing, Bloom filters and pseudo-random functions, the trade-off between security and efficiency of existing privacy protection Hamming distance calculation solutions is solved, and efficient and secure Hamming distance calculation is achieved, which is suitable for application scenarios of big data dimensions and real-time requirements.
Patent Information
- Application Number
- CN202510059422.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-15
- Publication Date
- 2025-06-13
AI Technical Summary
The existing privacy protection Hanming distance computing solution has a large trade-off between security and efficiency, resulting in high computing overhead and unsuitable for application scenarios with high requirements for big data dimensions and real-time.
A non-interactive Hamming distance privacy processing method is proposed, using cloud servers to realize outsourcing computing, and by blinding and hashing preprocessing data, combining Bloom filters and pseudo-random functions, the computing and communication overhead on the user side is reduced, and the privacy protection of input and output data is realized.
This solution is better than existing non-interactive systems in terms of efficiency, can efficiently calculate Hamming distance, reduce the computing burden on the user side, and supports application scenarios with big data dimensions and real-time requirements.
Smart Images

Figure CN120145409A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of privacy protection, and particularly relates to a privacy processing method for non-interactive Hamming distance. Background Art
[0002] As a commonly used similarity metric, the Hamming distance has been widely applied in fields such as information theory, bioinformatics, machine learning, and graph inspection. In the calculation of the Hamming distance for privacy protection, users should obtain the Hamming distance between their data and other data without revealing their private data. Existing solutions are based on the properties of binary sequences, transform the Hamming distance into the relationship of calculating data sum or product, and utilize homomorphic encryption and secure multi-party computation to achieve it. However, this will lead to significant computational overhead, thus posing challenges in practical applications.
[0003] With the widespread application of cloud computing, many industries are undergoing profound changes, driving innovation in data storage, processing, and analysis. Through the cloud platform, users can flexibly access powerful computing resources, which makes real-time data analysis and business decision-making possible. For this reason, cloud computing has been widely applied in multiple fields (such as e-commerce, financial services, and social media, etc.), promoting digital transformation and innovation, and significantly improving work efficiency.
[0004] As an important similarity metric, the Hamming distance plays a key role in the process of data retrieval and analysis, and can evaluate the differences between different samples. For example, in the field of bioinformatics, it can be used to evaluate the similarity of different DNA sequences, or for biometric features such as faces or irises; in information analysis, it can be used to detect and correct errors in data transmission, for image retrieval, or to judge the relevance of search results in search engines, etc. With the growth of data volume, the powerful computing ability of cloud computing makes it possible to calculate a large number of Hamming distances, promoting the development of intelligent applications.
[0005] Although cloud computing provides powerful data processing capabilities, the security and privacy issues of untrusted cloud service providers in processing data have also attracted wide attention. Since data needs to be transmitted to the cloud during the calculation process, this makes sensitive information possibly exposed in an unencrypted state, thus triggering the risk of privacy leakage. Malicious cloud service providers may use this leaked information to obtain users' personal data and behavior patterns. Therefore, when calculating the Hamming distance, it is necessary to take effective privacy protection measures to ensure the security of user data and prevent any leakage of sensitive information during the calculation process.
[0006] In the existing literature, only a few studies have begun to explore privacy-preserving computing schemes for Hamming distance. To securely compute the Hamming distance, most schemes use homomorphic encryption methods to compute sensitive information and achieve secure computing by means of secure multi-party computing (SMC), such as Garbled Circuits (GC) and Oblivious Transfer (OT). However, these schemes are usually accompanied by high communication and storage overheads and are difficult to meet the requirements of efficient processing. In addition, some schemes compare the Hamming distance with the trapdoor in a non-interactive manner. Although the interactive cost can be reduced, there are still significant bottlenecks in computing efficiency and it is difficult to meet application scenarios with high real-time requirements. For example, the PPH-HAM scheme does not support computations with large data dimensions, and the PHDC scheme only supports the Hamming distance computation of binary strings and also requires a large amount of computational overhead. Therefore, there is a large trade-off between security and efficiency in existing privacy-preserving computing schemes, and further optimization and improvement are urgently needed. This paper proposes a non-interactive privacy computing scheme for Hamming distance. This scheme uses a cloud server to implement outsourced computing, minimizing the computing and communication overheads on the user side and solving the problem of multi-round interactions required by some methods. Through experiments, the algorithm efficiency of the present invention is much superior to that of current non-interactive systems. At the same time, the present invention carefully designs blinding and hashing processes to achieve privacy protection for input and output data. Summary of the Invention
[0007] To solve the above technical problems, the present invention proposes a privacy processing method for non-interactive Hamming distance, including:
[0008] S1. The user preprocesses the original data sequence through blinding and hashing to obtain a set S' b ={a' 1 ,…,a' n ,b' 1 ,…,b' n} and sends it to the administrator;
[0009] S2. After receiving the set S' b ={a' 1 ,…,a' n ,b' 1 ,…,b' n}, the administrator processes the data through a Bloom filter and a pseudorandom function, and uploads the index set of the processed data set to the cloud server;
[0010] S3. The cloud server queries users in the index set of the processed data set whose Hamming distance is less than a threshold t to implement false positive user judgment.
[0011] In a preferred embodiment, in step S1, for the original data S=(a1 , a 2 , …, a n ), first generate a random number key of dimension n (b 1 , b 2 , …, b n ), after splicing, a new sequence S' = (a 1 , a 2 , …, a n , b 1 , b 2 , …, b n ) is obtained; perform hash processing on each data in the new sequence H(index, value) → hashcode, where index is the position information of each data in the new sequence, value is the value corresponding to the position of each data, H is the hash function, and hashcode is the hash value, to obtain the set S' b = {a′ 1 , …, a′ n , b′ 1 , …, b′ n}.
[0012] In the preferred embodiment, in the step S2, for process through the pseudo - random function F(k 1 ||x) to obtain:
[0013] C S = {F(k 1 ||x 1 ), …, F(k 1 ||x n )};
[0014] where k 1 represents the key, C S represents the data set, use the Bloom filter BF to generate the BF S corresponding to this data set, and upload BF S as the index set to the cloud server.
[0015] In the preferred embodiment, in the step S3, the user generates a trapdoor Token according to the sequence Q = (q 1 , …, q n ), first perform hash processing on the sequence Q H(index, value) → hashcode to obtain the set Q' b = {q' 1 , …, q' n}, the user randomly selects a perturbation value k 2 ∈ [n], and add k b to the set Q' 2a piece of perturbation data, the format of the perturbation data is the same as that in Q' b and merge it into Q' b For the system administrator calculates the pseudo-random function F(k 1 ||x) and obtains
[0016] The cloud server will run Check whether the elements in the Token exist in each Bloom filter. The number of elements that do not belong to the Bloom filter is the Hamming distance, and all Hamming distances are returned.
[0017] In the preferred embodiment, based on the correlation between the Hamming distance and the data position index, the data value of each dimension is mapped to a hash value hashcode;
[0018] H(index,value)→hashcode.
[0019] In the preferred embodiment, for two data A and B of the same length, the data dimension is n, representing that the data values of A and B on the same dimension are the same, and the number of different dimensions is the Hamming distance, then there is:
[0020] n = |A∩B| + d(A,B);
[0021] Also, since:
[0022] |A| = |A\B| + |A∩B| (1);
[0023] |B| = |B\A| + |A∩B| (2);
[0024] Because |A| = n and |B| = n, adding the two equations (1) and (2) gives:
[0025] |AΔB| = 2n - 2|A∩B|
[0026] The relationship between the Hamming distance and the length of the symmetric set difference is:
[0027] |AΔB| = 2d(A,B).
[0028] Compared with the prior art, the present invention has the following advantages:
[0029] The present invention proposes a batch privacy processing method for non-interactively implementing the Hamming distance. This scheme uses a cloud server to implement outsourcing computing, minimizing the computing and communication overhead on the user side and successfully solving the problem of multiple rounds of interaction required by the existing methods.
[0030] By constructing a secure hash value to make the algorithm universal, the present invention extends the algorithm to the alphabet set (such as alphanumeric) and the string set, and can be used as a plug-and-play Hamming distance calculation module.
[0031] Meanwhile, considering the false positives of the Bloom filter, the Hamming distance calculation proposed by the solution designed in the present invention will only be smaller than the correct value. In most systems, when it is necessary to find data with a Hamming distance less than the threshold, false positives can be achieved.
[0032] The present invention designs a blinding scheme for the pseudorandom function and the Bloom filter to achieve input-output security and the ability to resist frequency attacks. Finally, the present invention proves through simulation experiments that the proposed solution is superior to the existing methods in terms of efficiency. Brief Description of the Drawings
[0033] Figure 1 It is a schematic diagram of a Bloom filter;
[0034] Figure 2 It is a schematic diagram of the system model of the present invention;
[0035] Figure 3 It is a schematic diagram of the encryption process of the present invention;
[0036] Figure 4 It is a schematic diagram of the process of calculating the Hamming distance of the present invention;
[0037] Figure 5 It is a schematic diagram of the running time of big data dimension for testing the algorithm in big data dimension;
[0038] Figure 6 It is a graph showing the relationship between the time consumed by the user and the data dimension;
[0039] Figure 7 It is a graph showing the relationship between the encryption time and the data dimension of different solutions;
[0040] Figure 8 It is a graph showing the relationship between the comparison time and the data dimension of different solutions;
[0041] Figure 9 It is a schematic diagram of the comparison of the user side - total duration between the present invention and other solutions. Detailed Embodiments
[0042] To make the objectives, technical solutions, and advantages of the embodiments of the present application clearer, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are some, but not all, of the embodiments of the present application. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in the present application without creative efforts shall fall within the protection scope of the present application.
[0043] First, introduce the technical terms related to the present invention:
[0044] 1 Symmetric set difference
[0045] The symmetric set difference refers to the elements that belong to neither set A nor set B among two sets A and B, which can be expressed by the formula: AΔB = (A\B) ∪ (B\A). (A\B) represents the elements that are in A but not in B, and (B\A) represents the elements that are in B but not in A. For example, if A = {1, 2, 3, 4, 5} and B = {4, 5, 6, 7, 8}, then AΔB = {1, 2, 3, 6, 7, 8}.
[0046] 2 Hamming distance
[0047] The Hamming distance is defined as the number of positions where two equal-length strings are different. We define two equal-length strings A and B, and their Hamming distance is d(A, B), which can be expressed as: where n is the length of the string. For example, for the strings A = "01001" and B = "01011", their Hamming distance is d(A, B) = 1.
[0048] In the present invention, not only focus on the differences of binary strings, but also extend the Hamming distance to the calculation of a given range of characters and numbers.
[0049] 3 Bloom filter
[0050] The Bloom filter is mainly used to determine whether an element belongs to a set. It consists of a binary vector with a relatively large length and a set of hash functions, and has the advantages of fast query speed and small memory occupancy. However, the Bloom filter may falsely report that an element is in the set (i.e., false positive), but there will be no problem of missed reports.
[0051] The Bloom filter mainly includes three calculation processes, as Figure 1 shown:
[0052] 1. BF.Init, create a binary vector of size m, set all bits to 0, and select k independent hash functions.
[0053] 2. BF.add(x), use k hash functions to calculate k positions, and set the corresponding bit of the binary vector to 1.
[0054] 3. BF.check(q), when querying, also use k hash functions to calculate k positions. If all the bits corresponding to the calculated positions in the binary vector are 1, then it is considered that the element may be in the set; if any one of the bits is 0, then the element must not be in the set.
[0055] The algorithm of the present invention can be applied to a variety of application scenarios. For the convenience of understanding, a specific application scenario will be taken as an example for illustration below.
[0056] In a gene sequence analysis system, the user does not want to directly send their gene sequence to the administrator. Therefore, the user preprocesses their sample and uploads it to the administrator. The system administrator will process the gene sequences of multiple users and upload them to the cloud server as database index information. At the same time, the system administrator can retrieve other users in the database whose gene sequences are similar to that of the user based on the data uploaded by the user, so as to predict potential disease risks.
[0057] The present invention uses the Hamming Distance to measure the similarity between gene sequences. For example, when the Hamming Distance is less than a specific threshold, the system will consider these gene sequences to be similar and predict the disease probability of the user based on the known disease data of similar users.
[0058] This scenario involves three entities. The user (User) can upload the data they own. The administrator (Admin) is responsible for processing the data submitted by the user and retrieving similar data for them. The cloud server (CS) is responsible for storing the data uploaded by the system administrator and providing a batch retrieval function.
[0059] This model can be regarded as a Client-Server framework. The user and the administrator are regarded as the client side, and the cloud server is regarded as the server side. The system model diagram is as Figure 2 :
[0060] Existing Hamming distance privacy computations generally face three challenges:
[0061] 1) Only support Hamming distance privacy computation for 0-1 strings: Many solutions utilize the rules of 0-1 sequences for computation. For example, some solutions transform the exclusive OR operation into addition and multiplication operations Or the PHDC solution utilizes the relationship of 0-1 strings and calculates based on the condition that the sum of different bits is odd All require the string to be in the 0-1 format, which limits the scalability of the Hamming distance.
[0062] 2) Large computational and communication overheads: Many solutions utilize homomorphic encryption addition and multiplication operations, as well as secure multi-party computation methods, resulting in a large amount of computational and communication overheads. The computational load on the user side is relatively large, which is not conducive to expansion to scenarios such as the Internet of Things.
[0063] 3) Only support small data dimensions: The PHDC scheme processes data bit by bit. However, due to the large computational overhead generated by using homomorphic encryption, it is not suitable for large data dimension calculations. The PPH-HAM scheme will generate a large number of interpolation operations when the data dimension is large, and it is also not suitable for large data dimension calculations.
[0064] The present invention assumes that the user and the system administrator are completely trustworthy (that is, the client is completely trustworthy), while the cloud server is regarded as honest but curious. At the same time, some malicious programs may spy on user data through the cloud server. The goal of the present invention is to protect the user's privacy information from being leaked, and the calculation result of the final Hamming distance is also not leaked.
[0065] The following introduces the Hash processing method
[0066] Based on the characteristic that the Hamming distance is related to the data position index, the present invention hopes to map the data value of each dimension to a hash value. H is a secure hash function, index is the subscript of the input sequence, value is the value corresponding to index, and H(index, value) → hashcode.
[0067] The following introduces the design principle
[0068] Specifically, the core of the present invention is realized by relying on the relationship between the Hamming distance and the symmetric set difference. The present invention transforms the calculation of the Hamming distance into calculating half of the size of the symmetric set difference. To ensure input security and frequency attacks, the present invention designs a data blinding scheme. First, the user adds a random number sequence to the data and converts the string or data set into a location-related hash value through a secure Hash function. In addition, the system administrator uses a pseudo-random function to process each element.
[0069] To ensure output security, when generating the trapdoor, k 2 perturbation data will be added. The final Hamming distance is the difference between the result returned by the cloud server and k 2 , and k 2 is only known to the user, and the cloud server cannot obtain the correct Hamming distance.
[0070] Next, the present invention will verify the relationship between the Hamming distance and the size of the symmetric set difference:
[0071] Lemma 1: |AΔB| = 2d(A,B) holds when the following conditions are met: Two data with the same length have a data dimension of n, and the value of each dimension of the data can be mapped into a hash value, obtaining two sets A and B with the same length.
[0072] Proof:
[0073] According to the known conditions, it can be obtained that It means that the data values of A and B in a certain dimension are the same, and the number of different dimensions is the Hamming distance. That is,
[0074] n = |A ∩ B| + d(A, B)
[0075] Also, since |A| = |A\B| + |A ∩ B| and |B| = |B\A| + |A ∩ B|, and because |A| = n and |B| = n, adding the above two equations gives
[0076] |A Δ B| = 2n - 2|A ∩ B|
[0077] Combining the above two equations, the present invention can obtain the relationship between the Hamming distance and the length of the symmetric set difference.
[0078] |A Δ B| = 2d(A, B)
[0079] The specific method is as follows:
[0080] 1. User preprocessing:
[0081] The user inputs S = (a 1 , a 2 , …, a n ). To resist frequency attacks, the present invention designs a blinding scheme and combines it with the hash function H(index, value) to obtain the new set S' b = {a 1 , a 2 , …, a n , b 1 , b 2 , …, b n};
[0082] 2. The system administrator performs encryption.
[0083] The administrator receives the new set of user information S' b = {a 1 , a 2 , …, a n , b 1 , b 2 , …, b n}; and generates a random number key k 1 . For process it through the pseudo-random function F(k 1 ||x) to obtain C S = {F(k 1 ||x 1 ), …, F(k 1 ||x n )}. And use the Bloom filter BF to run Generate the BF corresponding to this set S . Upload the BF S as an index set to the cloud server. The detailed data can be encrypted by traditional encryption methods, and the encryption process is as Figure 3 shown.
[0084] 3. Query users whose Hamming distance is less than the threshold t.
[0085] The user generates a trapdoor Token according to Q=(q 1 ,…,q n ). First, perform a Hash process on the data H(index, value)→hashcode to obtain Q' b ={q' 1 ,…,q' n}. To ensure the output security, in the present invention, the user randomly selects a perturbation value k 2 ∈[n], and adds k 2 perturbed data to the set. The perturbed data added here is some impossible data calculations, such as H(-i - 1, "inpossible").
[0086] For the system administrator calculates the cloud server will run to check whether the elements in the Token exist in each Bloom filter. The number of elements not belonging to the Bloom filter is the returned Hamming distance. Since for each user k 2 is different, and only the user has k 2 , so the cloud server cannot know the specific size of the Hamming distance, as Figure 4 shown.
[0087] Next, analyze the correctness and security of the algorithm of the present invention
[0088] 1. Correctness analysis:
[0089] First, the principle of the solution of the present invention is |AΔB| = 2d(A, B). For example, Lemma 1. Next, the present invention will give the correctness of the solution.
[0090] Theorem 1: Without considering the false positives of the Bloom filter, the solution of the present invention is correct under a secure hash function H(index, value)→hashcode.
[0091] Proof: The present invention assumes that the user has an original data S=(a 1 ,a 2 ,…,a n), To avoid frequency attacks and achieve input security, the present invention blinds the set. First, a random character string b of length n is selected i , By combining the two parts, the present invention obtains a sequence of length 2n. After performing H(index, value) → hashcode on the data, S' is obtained b , The administrator will process S' using a pseudorandom function b to obtain the encrypted dataset C S , and generate a Bloom filter and upload it to the cloud server.
[0092] The query trapdoor sequence Q = (q 1 , q 2 , …, q n ). To ensure output security, the present invention adds k 2 interference values, and the interference values are different from any original data. The administrator uses a pseudorandom function to obtain the encrypted trapdoor set Token. At this time, due to the same key k 1 of the pseudorandom function of the present invention, the same hash value will be generated for the same element, and it has no effect on BF.check.
[0093] Take the initial value of the count value con as 0. Next, the present invention will calculate |(S\Q)|. For each value in Token, query whether it exists in the Bloom filter. If it does not exist, then con+1. On the basis of not considering the false positives of the Bloom filter, the Hamming distance returned by the cloud server should be res = con + k 2 , After the user obtains the value, calculate the Hamming distance d(S, Q) = res - k2.
[0094] And since |S| = |Q|, and S∩Q = Q∩S, so |(S\Q)| = |(Q\S)|, and the above query |(S\Q)| = 1 / 2|(SΔQ)| of the present invention. Since Lemma 1 holds, the theorem is proven.
[0095] Generally, the present invention hopes to obtain more similar strings, that is, strings with a smaller Hamming distance. Due to the false positives of the Bloom filter, when the present invention queries data, a certain bit may be misjudged as existing in the Bloom filter, resulting in a decrease in the Hamming distance by one bit. For the case of judging d(A, B) < t, false positives are also achieved, and there will be no missed reports.
[0096] 2. Security analysis
[0097] The present invention defines a semi-honest cloud server and an external adversary. The external adversary can obtain some or all of the transmission information through eavesdropping. Since the present invention is a client-server framework, the maximum information that the external adversary can steal is less than that of the cloud server.
[0098] In the solution of the present invention, during the process of uploading to the cloud server, an adversary can see a series of Bloom filters, namely: VIEW input ={BF set} During the process of uploading the trapdoor and performing calculations and queries, an adversary can see the trapdoor and the result set, namely: VIEW output ={Token,Res set}
[0099] Theorem 2: Input Privacy: The proposed solution can guarantee the privacy of the input data, that is, an adversary cannot learn any information about the user's input data.
[0100] Proof: For BF set , first of all, for each Bloom filter, the present invention will eliminate the correlation between them. It should be noted that if the same k hash functions are used, the generated Bloom filters will directly expose their correlation, that is, the Hamming distance between the binary vectors of the Bloom filters. However, the k hash functions used in the present invention are different and randomly generated, which makes it possible to obtain completely different outputs even for the same input.
[0101] At the same time, in order to prevent an adversary from guessing the possible dimension of the data set based on the Bloom filters and the number of Tokens in the case of a small input data set, the present invention performs a blinding operation on the data, adding equal-length data to make the Bloom filters more complex, so that the attacker cannot guess the dimension of the data set. Most importantly, the present invention uses a collision-resistant hash function to implement a pseudo-random function, generating a hash value for the data and then generating a Bloom filter. An adversary cannot convert the Bloom filter back to a hash value, let alone convert it back to the original data through the hash value (the irreversibility of the hash function), further protecting the data of the present invention.
[0102] Theorem 3: Output Privacy: The proposed solution can guarantee the privacy of the output data, that is, an adversary cannot infer the true Hamming distance d(A,B).
[0103] Proof: Since the cloud server stores the uploaded Bloom filter values, the present invention will analyze the security of an adversary in the case of knowing BF set , Token, Res set . The size of Res set output by the present invention is the sum of the true Hamming distance d(A,B) and the key k 2 , but the value of k 2 is randomly generated and only the client has it, so the CS cannot know d(A,B). As long as k 2 is not leaked, the adversary cannot know the specific d(A,B).
[0104] In the present invention, it is considered that since the adversary (CS) can know whether each value in the trapdoor exists in the BF, and the Hamming distance is strongly correlated with the position information, the present invention uses a random mapping function to modify the specific order in the trapdoor set, so that the adversary cannot know which specific position data exists in the BF. Therefore, the output privacy of the proposed scheme is guaranteed.
[0105] The following introduces the experimental evaluation process
[0106] 1. Settings
[0107] Through simulation experiments, the present invention evaluated the performance of the proposed algorithm and reproduced the non-interactive PHDC scheme and the PPH-HAM scheme. When making comparisons, the present invention ignored the communication overhead existing in the PHDC scheme. All experiments were carried out on a computer equipped with an Intel(R) Core(TM) i5-8265U CPU@1.60GHz, 24GB RAM, and Ubuntu 18.04. The present invention implemented the algorithm of the present invention using the SHA-256 algorithm in the hashlib library, implemented the PHDC scheme based on the PHE and Crypto libraries, and implemented the PPH-HAM scheme based on the Charm and Crypto libraries.
[0108] The above algorithms were all implemented based on Python 3.6, and the present invention repeated the experiments five times and took the average value of the results.
[0109] 2. Algorithm evaluation
[0110] Due to the superiority of the Bloom filter in terms of space efficiency, the present invention tested the algorithm in large data dimensions to adapt to the current social situation of the rapid increase in data volume. At the same time, the present invention fixed the value of k 2 to be half of n, and it can be seen from the running time in large data dimensions that even when processing data with a very long dimension, the algorithm of the present invention still maintains efficient processing. Figure 5
[0111] The random number k 2 determines the output security of the algorithm, but at the same time has a slight impact on the algorithm efficiency. Next, the present invention sets the data dimension n = 50 and expands the dimension of the extended random number k 2 from 5 to 50. It can be seen from the results of Figure 6 that the time consumed by the user is basically not affected by the random number dimension k 2 2 This is because the change of k 2 only affects the trapdoor generation time at the user side, and the trapdoor generation basically only requires several hash calculations. Therefore, the value of k 2 basically does not affect the calculation cost on the user side.
[0112] 3. Comparative Experiments
[0113] The present invention tested the performance of the Paillier cryptosystem and the ElGamal cryptosystem on this platform. The test results show that the calculation time for the Paillier encryption operation is TE = 282.095 ms, and the calculation time for the decryption operation is TD = 78.654 ms; the calculation time for the ElGamal encryption operation is TEnc = 9.171 ms, and the calculation time for the decryption operation is TDec = 6.066 ms. In addition, the calculation time required for the homomorphic addition operation between two Paillier ciphertexts is TA = 291.945 ms.
[0114] Since the method adopted by the present invention is different from the comparative experiment, the present invention compares the three schemes at two levels: data encryption and querying the Hamming distance, covering the overall process of each algorithm. The scheme cannot obtain the specific Hamming distance, and the scheme theoretically belongs to a two-party algorithm. Therefore, the present invention only calculates the efficiency difference between two strings and can only be applied to Boolean data. Therefore, binary sequences of different dimensions n are generated in the experiment.
[0115] 3.1 Data Encryption
[0116] The data encryption part in the scheme of the present invention includes two parts: generating a Bloom filter and generating a trapdoor, which can be regarded as encrypting two strings. In the PPH-HAM scheme, the present invention calculates PPH.Hash for two strings as the data encryption part; in the PHDC scheme, the present invention takes the first part described in the literature as the data encryption, which does not include the process of calculating the Hamming distance using homomorphic encryption. Since the PHDC scheme and the PPH-HAM scheme are not suitable for large dimensions, the present invention only tests the data sets with data dimensions from 5 to 50. Since the value of k 2 has an impact on the PHDC scheme, the present invention takes the value of k 2 as n / 2. As Figure 7 shown, it can be seen that the encryption efficiency of the present invention is 6 orders of magnitude higher than that of the PHDC scheme and 4 orders of magnitude higher than that of the PPH-HAM scheme, and it can be used in Internet of Things devices with limited computing power, etc.
[0117] 3.2 Data Calculation
[0118] The PHDC scheme uses homomorphic encryption to calculate the Hamming distance of ciphertexts, resulting in a large amount of computational overhead. The PPH-HAM scheme uses rational function interpolation to judge the Hamming distance and threshold. The scheme of the present invention uses a Bloom filter to calculate the Hamming distance. As Figure 8As shown, it can be seen from the experiment that the efficiency of the present invention in calculation is much higher than that of the PHDC scheme. When the data dimension n increases, the number of trapdoors will also increase, and the amount of calculation will also increase. When the data dimension exceeds 20, the PPH-HAM scheme is more efficient. However, it should be noted that this scheme relies on the calculation of the public parameter Γ (a matrix composed of g rn ). When n > 50, the calculation of Γ will consume a large amount of computing costs and is not applicable to large data dimensions.
[0119] 3.3 Proportion of Client Overhead
[0120] As Figure 9 shown, the present invention compares the proportion of the calculation amount on the user side in the total calculation overhead between the present invention and the scheme in the case of ignoring the communication overhead. It can be seen that the scheme of the present invention greatly reduces the overhead on the user side, and the operations on the client side are only encrypting data and uploading trapdoors, which is linearly increasing and will not occupy too much computing cost on the user side. This scheme can also be used for Internet of Things devices with relatively poor computing and storage capabilities.
[0121] This paper proposes a practical privacy-preserving Hamming distance calculation scheme. This scheme combines data blinding, Bloom filters, and pseudorandom functions, can efficiently calculate the Hamming distance, and ensure privacy protection at the same time. The present invention adopts a client-server architecture, realizes non-interactive calculation, and transfers most of the calculation overhead to the cloud server, thus reducing the calculation burden on the client side. Compared with the existing privacy-preserving Hamming distance calculation schemes, the scheme of the present invention has obvious advantages in terms of efficiency. In addition, the present invention supports the calculation of Hamming distance for integers, characters, and strings through secure hash functions, solving the problem that most current schemes only support binary data. In practical applications, data often comes from different owners and needs to support multi-user queries and calculations.
[0122] The technical features of the above embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope recorded in this specification.
[0123] The above-described embodiments only represent several implementation manners of the present application. The description is relatively specific and detailed, but it should not be construed as a limitation on the patent scope of the present application. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present application, several deformations and improvements can still be made, and these all belong to the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the appended claims.
[0124] The technical features of the above embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope described in this specification.
[0125] The above-described embodiments merely represent several implementation manners of the present application. The description thereof is relatively specific and detailed, but it should not be construed as a limitation to the patent scope of the present application. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present application, several modifications and improvements can still be made, and these all belong to the protection scope of the present application. Therefore, the protection scope of the present application shall be subject to the appended claims.
Claims
1. A privacy processing method based on non-interactive Hamming distance, characterized in that: include: S1, the user preprocesses the original data sequence by blinding and hashing to obtain the set S' b ={a'1,…,a' n ,b'1,…,b' n } and send to the administrator; S2. The administrator receives the set S' b ={a'1,…,a' n ,b'1,…,b' n }, processing the data through a Bloom filter and a pseudo-random function, and uploading an index set of the processed data set to a cloud server; S3. The cloud server queries the index set of the processed data set for users whose Hamming distance is less than the threshold t to achieve false positive user judgment.
2. The privacy processing method of non-interactive Hamming distance according to claim 1, characterized in that: In step S1, for the original data S=(a1, a2, ..., a n ), first generate a random number key (b1, b2, ..., b n ), after splicing, we get a new sequence S'=(a1,a2,…,a n ,b1,b2,…,b n ); Hash each data in the new sequence H(index, value) → hashcode, where index is the position information of each data in the new sequence, value is the value corresponding to the position of each data, H is the hash function, hashcode is the hash value, and the set S' is obtained b ={a'1,…,a' n ,b'1,…,b' n }.
3. The privacy processing method of non-interactive Hamming distance according to claim 1, characterized in that: In step S2, for Through the pseudo-random function F(k1||x), we get: C S ={F(k1||x1),…,F(k1||x n )}; Among them, k1 represents the key, C S Represents a data set, using Bloom filter BF to generate the BF corresponding to the data set S· BF S Uploaded to the cloud server as an index collection.
4. The privacy processing method of non-interactive Hamming distance according to claim 2, characterized in that: In step S3, the user and the administrator follow the sequence Q=(q1,…,q n ) Generate a trapdoor Token, first perform hash processing on the sequence Q H(index, value) → hashcode to obtain the set Q' b ={q'1,…,q' n }, the user randomly selects a perturbation value k2∈[n] and in the set Q' b Add k2 perturbation data in the same format as Q' b The data in the same format are merged into Q' b ,for The system administrator calculates the pseudo-random function F(k1||x) and obtains The cloud server will run Check whether the element in Token exists in each Bloom filter. The number of elements that do not belong to the Bloom filter is the Hamming distance. Return all Hamming distances.
5. The privacy processing method of non-interactive Hamming distance according to claim 4, characterized in that: Based on the correlation between Hamming distance and data position index, the data value of each dimension is mapped to a hash value hashcode; H(index,value)→hashcode.
6. The privacy processing method of non-interactive Hamming distance according to claim 1, characterized in that: Two equal-length data A and B, with a data dimension of n. It means that the data values of A and B on the same dimension are the same, and the number of different dimensions is the Hamming distance, then: n=|A∩B|+d(A,B); And because: |A|=|A\B|+|A∩B| (1); |B|=|B\A|+|A∩B| (2); Since |A|=n and |B|=n, adding equations (1) and (2) together, we get: |AΔB|=2n-2|A∩B| The relationship between the Hamming distance and the symmetric set difference length is: |AΔB|=2d(A,B).