Two-party privacy data security query method and system based on Hash proof system
Through the hash proof system's privacy data security query method, the communication complexity problem in the imbalanced computing capabilities and data scale of the client and server side is solved, communication efficiency and computing adaptability are improved, results are ensured, and the results are accurate, and suitable for resource-constrained devices.
Patent Information
- Application Number
- CN202510640097.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-19
- Publication Date
- 2025-07-18
AI Technical Summary
The existing privacy data query methods have high communication complexity in scenarios where the computing power and data scale of the client and server side are unbalanced. The communication overhead is related to the data scale of the server side, and the computing complexity is high, making it difficult to deploy in resource-constrained devices, with large errors and affecting the accuracy of the result.
The hash proof system is used to calculate the hash value through the client and send it to the server. The server performs data matching and difference calculations, generates the final query results, and is verified by the client to reduce the communication complexity and calculation complexity, and adapt to resource-constrained devices.
It significantly improves the communication efficiency of privacy computing, reduces communication complexity and computing overhead, and is suitable for resource-constrained devices to ensure the accuracy of results and avoid misjudgment.
Smart Images

Figure CN120342743A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of cryptography and information security, and particularly relates to a two-party private data security query method based on a hash proof system, which is particularly applicable to privacy-preserving queries in an unbalanced scenario where the client has limited computing power and the server-side dataset is large, and can be applied to resource-constrained environments such as mobile devices and Internet of Things terminals. Background Art
[0002] The goal of two-party private data security query is to allow the client to obtain the query result without revealing the data it queries, and the server not only does not know the client's query content and query result, but also cannot disclose the content of its dataset other than the query result. This technology is used to solve the data collaboration problem in privacy-sensitive scenarios. With the advancement of the digitalization process, it has demonstrated important application value in many fields. For example: social platforms use this technology to find private contacts to avoid complete exposure of user address book information; security companies verify whether a user's password already exists in a leaked database; medical institutions compare patients' genetic data, and users share activity area location information, etc.
[0003] Existing private data query methods are mainly implemented through technologies such as key exchange, fully homomorphic encryption, oblivious transfer extension, and oblivious pseudorandom functions. For the scheme based on key exchange technology, the communication complexity between the two parties not only increases linearly with the size of the query request data, but also increases linearly with the size of the server-side dataset. When the data size of one server side is extremely large (such as the server stores millions of data), such a scheme has too much load during the communication process, resulting in resource waste; for the scheme based on fully homomorphic encryption technology, although the communication overhead is reduced, it is necessary to construct a complex circuit, and the circuit calculation depth increases exponentially with the size of the server-side data, making it difficult to be deployed in resource-constrained devices; for the scheme based on oblivious pseudorandom function technology, the preprocessing stage still depends on the size of the server data set, and the compression technology may introduce errors, affecting the reliability of high-precision scenarios such as medical or financial fields. In addition, some schemes rely on trusted third parties or complex cryptography technologies, which not only increase the deployment complexity, but also may introduce single-point failure or collusion risks. These problems do not fully consider the requirements of the unbalanced scenario of the computing power and data size between the client and the server, resulting in a very high communication complexity of the existing computing methods in this scenario.
[0004] Therefore, how to reduce the communication complexity of the privacy data security query method between the client and the server in the scenario of unbalanced computing power and data scale, so that the communication overhead of both parties is only related to the data scale of the client's query request and shows a linear growth relationship, eliminating the dependence of the original scheme's communication complexity on the server's data scale; how to reduce the computational complexity of the server so that it is only linearly related to the smaller set scale and linearly related to the logarithm of the larger set scale; how to reduce the dependence on complex cryptographic tools (such as oblivious transfer extension, fully homomorphic encryption) to adapt to resource-constrained devices; and how to eliminate errors to ensure the complete accuracy of the intersection calculation result and avoid the misjudgment problem introduced by existing compression technologies are the technical problems that urgently need to be solved at present. Summary of the Invention
[0005] By providing a two-party privacy data security query method based on a hash proof system, this application solves the problem of high communication complexity of the privacy data security query method between the client and the server in the scenario of unbalanced computing power and data scale, and greatly improves the communication efficiency of privacy computing.
[0006] This application provides a two-party privacy data security query method based on a hash proof system, which is characterized by including:
[0007] Step 1, initialize and generate the key and global parameters of the hash proof system;
[0008] Step 2, the client calculates the hash value of each element in its query data set, and then sends the hash values of all element query data to the server;
[0009] Step 3, the server calculates the hash value of its own data set, matches the calculation result with the hash value of the received query data, and calculates the matching difference of each pair of elements;
[0010] Step 4, the server aggregates all the matching differences of each query data to generate a final query result, and returns the final query result to the client;
[0011] Step 5, the client verifies the returned final query result and obtains the corresponding privacy query data.
[0012] Further, the step 1 includes:
[0013] Step 1.1, generate the private key of the hash proof system, and randomly select an s-dimensional private vector where N is a large prime number and s is a parameter related to the security parameter;
[0014] Step 1.2, generate the public key of the hash proof system according to the private key, and select a random matrix Calculate the public key vector hpk = hsk·A mod N;
[0015] Step 1.3: Then randomly select t random vectors and L is the set of languages defined in the hash proof system;
[0016] Step 1.4: Publicize the parameter set {hpk, A, a1, a2, …, a t , N}.
[0017] Furthermore, Step 1 further includes:
[0018] Step 1.5: For each query data y j , randomly select two prime numbers p j and q j , satisfying p j ≡ q j ≡ 3 mod 4;
[0019] Step 1.6: Calculate the RSA modulus N j =(2p j +1)(2q j +1), ensuring that N j is a 2λ + 2-bit number;
[0020] Step 1.7: Then randomly generate a generator g of the group j , and the order of g j is φ(N j ) = 4p j q j ;
[0021] Step 1.8: Send the parameter set {N1, N2, … N n , g1, g2, … g n} to the server.
[0022] Furthermore, Step 2 includes:
[0023] Step 2.1: Decompose each query data y j into t-bit binary numbers, i.e., y j = y j1 y j2 … y jt , where y jk ∈ {0, 1};
[0024] Step 2.2: Randomly select a t-dimensional vector j for each y satisfying the constraint condition hpk·w j = p jmod N, where p j is the RSA prime number generated for the client;
[0025] Step 2.3, calculate the hash value vector of each query data
[0026] Step 2.4, send the hash value of the query data to the server.
[0027] Furthermore, the said Step 3 includes:
[0028] Step 3.1, the server decomposes each x n in its own data set X = {x1, x2,... x i} into binary numbers x j = x j1 x j2 …x jt , and calculates the hash value of each element
[0029] Step 3.2, the server matches the query data of the client according to the hash value of its own data set. First, it uses the hash proof system to verify the query data of the client that is, calculates
[0030] where hpk[k] is the k-th component of the public key vector. Combining the constraint condition hpk·w j = p j mod N, the server can obtain the encoding of the query data of the client where e k is the unit vector whose k-th component is 1.
[0031] Furthermore, the said Step 4 includes:
[0032] Step 4.1, the server matches the hash value of each client query data, and calculates the matching difference of each pair of elements (x i , y j ) where x′ i is the encoding vector of x i ;
[0033] When the client query data matches the server-side data, that is, x i = y j , then the matching difference z ij = p jmod N; otherwise, match the difference value z ij is a pseudo-random number;
[0034] Step 4.2. The server aggregates all the matching difference values z j for each query data y ij (i ∈ [m]) to generate the final corresponding query result
[0035] Step 4.3. The server returns the encoded values {C1, C2, … C n} of the query result to the client.
[0036] Furthermore, in Step 4.2, the power of the generator g j is obtained by using modular arithmetic to calculate cyclically.
[0037] Furthermore, Step 5 includes:
[0038] Step 5.1. The client first calculates the Euler's totient function value φ(N j corresponding to each query data y j ) = 4pq j of the RSA modulus N j q j , and sets t j = φ(N j ) / p j ;
[0039] Step 5.2. Judge whether holds. If it holds, then y j is successfully matched; obtain the corresponding private query data, otherwise the match fails and the data query ends.
[0040] The present invention also discloses an electronic device, which is characterized by including: a memory and at least one processor; wherein, the memory stores computer-executable instructions; the at least one processor executes the computer-executable instructions stored in the memory, so that the at least one processor executes the above two-party private data security query method based on the hash proof system.
[0041] The present invention also discloses a computer-readable storage medium, which is characterized in that computer-executable instructions are stored in the computer-readable storage medium, and when the processor executes the computer-executable instructions, the above two-party private data security query method based on the hash proof system is implemented.
[0042] The technical solution provided by this application has at least the following advantages:
[0043] 1. The present invention combines a hash proof system with RSA parameters, significantly improving the communication efficiency of its privacy computing. It reduces the two-party communication complexity to be only related to the size of the data queried by the client and shows a linear relationship, reducing the communication bandwidth consumption of both parties. Moreover, the client only needs to perform simple modulo addition, modulo multiplication, and fast modular exponentiation operations in terms of calculation, without the complex circuit design, calculation, or preprocessing operations required in other solutions, and is applicable to devices with limited computing resources.
[0044] 2. The privacy data security query method of the present invention reduces the computational complexity of the server side, making it linearly related only to a relatively small set size and linearly related to the logarithm of a relatively large set size.
[0045] 3. The privacy data security query method of the present invention reduces the dependence on complex cryptographic tools and is suitable for resource-constrained devices.
[0046] 4. The privacy data security query method of the present invention can eliminate errors, ensure the complete accuracy of the intersection calculation result, and avoid the misjudgment problem introduced by existing compression technologies. BRIEF DESCRIPTION OF THE DRAWINGS
[0047] Figure 1 It is a schematic diagram of the privacy intersection calculation scenario in the embodiment.
[0048] Figure 2 It is a schematic diagram of the initialization process of privacy intersection calculation in the non-equilibrium scenario in the embodiment.
[0049] Figure 3 It is a schematic diagram of the calculation stage process of privacy intersection calculation in the non-equilibrium scenario in the embodiment. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0050] The core of the present invention lies in constructing a computational method for privacy data security query based on a hash proof system (Hash Proof System, HPS), which is applicable to scenarios where the computing capabilities and data scales of the client and the server are unbalanced. This technical solution is divided into an initialization stage and a calculation stage.
[0051] The initialization goal is to generate the global parameters required by the scheme and the key of the hash proof system to ensure the security and efficiency of subsequent calculations. It mainly includes the following two steps:
[0052] Step 1. The server generates the key of the hash proof system. As the party with rich resources in privacy computing, the server is responsible for initializing the public-private key pair of the hash proof system.
[0053] Step 1.1. The server generates the private key of the hash proof system: randomly select an s-dimensional private vector where N is a large prime number (for example, N≈2 512), where s is a parameter related to the security parameter;
[0054] Step 1.2: The server generates the public key of the hash proof system according to the private key: Select a random matrix Calculate the public key vector hpk = hsk · A mod N;
[0055] Step 1.3: The server randomly selects t random vectors and (L is the set of languages defined in the hash proof system);
[0056] Step 1.4: The server publishes the parameter set {hpk, A, a1, a2, …, a t , N}.
[0057] The client generates RSA system parameters. Due to limited computing resources, the client can generate multiple query requests at one time, and then needs to generate independent RSA system moduli for each query data y j ∈ Y (j ∈ [n]). The specific steps include:
[0058] Step 1.5: For each query data y j , the client randomly selects two prime numbers p j and q j with λ (security parameter) bits in length, satisfying p j ≡ q j ≡ 3 mod 4, and then calculates the RSA modulus N j = (2p j + 1)(2q j + 1), ensuring that N j is 2λ + 2 bits;
[0059] Step 1.6: The client then randomly generates a generator g of the group j , requiring that the order of g j is φ(N j ) = 4p j q j ;
[0060] Step 1.7: The client sends the parameter set {N1, N2, … N n , g1, g2, … g n} to the server.
[0061] In the calculation stage, the server calculates the required query results according to the client query request set Y, which is specifically divided into four sub-steps: client hash value calculation, server-side data matching, query result generation, and client verification.
[0062] Step 2. The client performs a hash calculation on each element in its query data set \(Y = \{y_1, y_2, \ldots, y\}\), and the specific steps are as follows: n} including:
[0063] Step 2.1. The client decomposes each \(y\) j into \(t\)-bit binary numbers, that is, \(y\) j = \(y\) j1 \(y\) j2 \(\ldots y\) jt where \(y\) jk \(\in \{0, 1\}\);
[0064] Step 2.2. The client randomly selects a \(t\)-dimensional vector j for each \(y\) that satisfies the constraint \(hpk \cdot w\) j = \(p\) j \(\bmod N\), where \(p\) j is the RSA prime number generated by the client;
[0065] Step 2.3. The client calculates the hash value vector of each query data
[0066] Step 2.4. The client sends the hash value of the query data to the server.
[0067] Step 3. The server uses the private key \(hsk\) of the hash proof system to calculate the hash value of its own data set \(X=\{x_1, x_2, \ldots, x\}\) n} and performs data matching. The specific steps are as follows:
[0068] Step 3.1. The server decomposes each \(x\) i \(\in X\) into binary numbers \(x\) j = \(x\) j1 \(x\) j2 \(\ldots x\) jt and calculates the hash value of each element
[0069] Step 3.2. The server starts to match the client's query data according to the hash value of its own data set. It first uses the hash proof system to verify the client's query data that is, calculates where \(hpk[k]\) is the \(k\)-th component of the public key vector. Combining the constraint \(hpk \cdot w\) j = \(p\) j \(\bmod N\), the server can obtain the encoding of the client's query data where \(e\) k is the unit vector whose \(k\)-th component is 1.
[0070] Step 4, Query result generation, the specific steps include:
[0071] Step 4.1, The server starts to match the hash values of the query data of each client, and calculates the matching difference of each pair of elements (x i , y j ). Where x' i is the encoded vector of x i . When the client query data matches the server - side data, that is, x i = y j , then the matching difference z ij = p j mod N; otherwise, the matching difference z ij is a pseudo - random number.
[0072] Step 4.2, The server aggregates the matching differences for each query data. That is, the server aggregates all its matching differences z j for each query data y ij (i ∈ [m]) to generate the final corresponding query result Here, the calculation uses modular arithmetic to calculate the powers of the generator g j to obtain the result, so that the server cannot identify which query data is matched and which query data fails to match the content, rather than calculating the product of multiple differences z ij first and then using it as an exponent.
[0073] Step 4.3, The server returns the encoded values {C1, C2, … C n} of the query result to the client.
[0074] Step 5, The client verifies and obtains the corresponding private query data. The client determines whether the query data y j is successfully matched through the following steps:
[0075] Step 5.1, The client first calculates the Euler's totient value φ(N j ) of the corresponding RSA modulus N j for each query data y j ) = 4p j q j , and sets t j = φ(N j ) / p j ;
[0076] Step 5.2, The client only needs to judge whether holds. If it holds, then y jMatch successful; obtain the corresponding privacy query data, otherwise match fails and end the data query.
[0077] To better understand the above technical solution, the above technical solution will be described in detail below in conjunction with the accompanying drawings of the specification and specific implementation manners.
[0078] This embodiment provides a privacy data security query calculation method based on a hash proof system, which is applicable to scenarios where the computing capabilities and data scales of the client and the server are unbalanced. Figure 1 The applicable scenario of the present invention is shown. For example, a certain medical institution uses this technology for privacy protection of gene data comparison. The client is a mobile medical terminal (such as a doctor's handheld device), which stores the patient gene fragment set Y = {y1, y2, …, y 1000}, and each y j is a gene marker (such as an SNP site). The server is a cloud gene database, which stores the gene data shared by global research institutions for matching potential genetic disease associated sites of patients.
[0079] As Figure 2 shown, the server and the client first perform parameter generation and transmission. In this embodiment, the server selects the security parameter λ = 512, generates the hash proof private key and the public key hpk = hsk·A mod N, where N = 2 512 +1657, the matrix A is randomly selected from , and then the random vector is randomly selected. Finally, the parameters {hpk, A, a1, a2, …, a t , N} are published.
[0080] The client generates 1024-bit long RSA moduli N1, N2, … N 1000 and the corresponding generators g1, g2, … g 1000 and sends them to the server.
[0081] As Figure 3 shown, the server matches the data and returns the matching result. The client performs privacy protection processing on the gene marker y j , calculates where y jk is the binary expansion of the gene marker data, w j is a randomly selected vector that satisfies the constraint conditions, and sends to the server.
[0082] The server performs matching processing on the received data, first calculates the hash value of its own data set Next, calculate the hash value of the data queried by the client Then calculate the difference between the hash value of its own data and the hash value of the data queried by the client Finally, aggregate all the matching differences corresponding to each queried data And return the query results C1, C2, … C 1000 to the client;
[0083] The client calculates and judges If they are equal, the match is successful; otherwise, the match is unsuccessful. Finally, based on the obtained matching results, judge whether the gene data has a genetic disease risk. The entire calculation process does not disclose the patient's gene data.
[0084] The above are only the preferred embodiments of the present invention. It should be pointed out that for those of ordinary skill in the art, without departing from the principle of the present invention, several improvements can be made, and these improvements should also be regarded as the protection scope of the present invention.
Claims
1. A two-party private data security query method based on a hash proof system, characterized in that Including: Step 1: Initialize the key and global parameters of the hash proof system; Step 2: The client calculates the hash value for each element of its query dataset, and then sends the hash values of all the element query data to the server; Step 3: The server calculates the hash value for its own dataset, and performs data matching between the calculation result and the hash value of the received query data, and calculates the matching difference for each pair of elements; Step 4: The server aggregates all the matching differences for each of the said query data, generates the final query result, and returns the final query result to the client; Step 5: The client verifies the returned final query result and obtains the corresponding private query data.
2. The two-party private data security query method based on the hash proof system according to claim 1, wherein The said Step 1 includes: Step 1.
1. Generate a private key of the hash proof system and randomly select an s-dimensional private vector where N is a large prime number and s is a parameter related to the security parameter; Step 1.
2. Generate the public key of the hash proof system based on the private key and select a random matrix Calculate the public key vector hpk = hsk · A mod N; Step 1.
3. Then randomly select t random vectors and L is the set of languages defined in the hash proof system; Step 1.4, disclose the parameter set {hpk, A, a1, a2, …, a t , N}.
3. The two-party private data security query method based on the hash proof system according to claim 1, wherein The said Step 1 further includes: Step 1.
5. For each query data y j , randomly select two prime numbers p j and q j such that p j ≡ q j ≡ 3 mod 4; Step 1.6, calculate the RSA modulus N j =(2p j +1)(2q j +1), ensure that N j is 2λ + 2 bits; Step 1.
7. Randomly generate a generator g of the group j , and the order of g j is φ(N j ) = 4pq j ; j ; Step 1.8, send the parameter set {N1, N2, … N n , g1, g2, … g n} to the server.
4. The two-party private data security query method based on the hash proof system according to claim 2, wherein The said Step 2 includes: Step 2.1: Decompose each query data y j into t-bit binary numbers, i.e., y j = y j1 y j2 …y jt , where y jk ∈{0,1}; Step 2.2: For each y j Randomly select a t-dimensional vector Satisfying the constraint condition hpk·w j = p j mod N, where p j Is the RSA prime number generated for the client; Step 2.3, calculate the hash value vector of each query data Step 2.4, send the hash value of the query data to the server.
5. The two-party private data security query method based on the hash proof system according to claim 4, characterized in that The said Step 3 includes: Step 3.
1. The server decomposes each \(x\) in its own data set \(X=\{x_1,x_2,\ldots,x\) n \}\) into binary numbers \(x\) i \(=x\) j \(x\) j1 \(\ldots x\) j2 \(\ldots\), and calculates the hash value of each element jt Step 3.2: The server matches the client's query data according to the hash value of its own dataset. First, it uses the hash proof system to verify the client's query data That is, calculate where hpk[k] is the k-th component of the public key vector, combined with the constraint hpk · w j = p j mod N, the server can obtain the encoding of the client's query data where e k is the unit vector with the k-th component being 1.
6. The two-party private data security query method based on the hash proof system according to claim 5, characterized in that The said Step 4 includes: Step 4.1: The server matches the hash values of the data queried by each client and calculates the matching difference of each pair of elements (x i , y j ). Where x' i is the encoded vector of x i . When the data queried by the client matches the server - side data, i.e., x i = y j , then the matching difference z ij = p j mod N; otherwise, the matching difference z ij is a pseudo - random number; Step 4.
2. The server aggregates all the matching differences z j for each query data y ij (i ∈ [m]) to generate the final corresponding query result Step 4.
3. The server returns the encoded values {C1, C2, … C n} of the query result to the client.
7. The two-party private data security query method based on the hash proof system according to claim 6, wherein Step 4.2 uses modular arithmetic to loop through the calculations of powers of the generator g j to obtain the result.
8. The two-party private data security query method based on the hash proof system according to claim 6, wherein, The said Step 5 includes: Step 5.1: The client first calculates each query data y j The corresponding RSA modulus N j The Euler value φ(N j )=4p j q j , and set t j =φ(N j ) / p j ; Step 5.2, determine whether holds. If it holds, then y j is successfully matched, and the corresponding privacy query data is obtained; otherwise, the match fails and the data query ends.
9. An electronic device, characterized in that, Including: A memory and at least one processor; wherein, the memory stores computer-executable instructions; the at least one processor executes the computer-executable instructions stored in the memory, so that the at least one processor executes the two-party private data security query method based on the hash proof system as described in any one of claims 1 to 8.
10. A computer-readable storage medium, characterized in that, Computer-executable instructions are stored in the said computer-readable storage medium, and when the processor executes the computer-execution, the two-party private data security query method based on the hash proof system as described in any one of claims 1 to 8 is implemented.