Implementation method for security verifiable outsourcing database with privacy protection
Through the combination of proxy re-encryption, vector commitment and Merkle tree, combined with homomorphic encryption and inadvertent pseudo-random functions, the problems of data privacy leakage, query results cannot be verified and user query privacy protection in outsourcing databases are solved, and the secure storage of data and the reliability of query results are achieved.
Patent Information
- Application Number
- CN202510215518.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-26
- Publication Date
- 2025-06-13
AI Technical Summary
The prior art has problems in outsourcing databases with data privacy leakage, inability to verify query results, and user query privacy protection, especially in an environment where cloud servers cannot be fully trusted.
Through the combination of proxy re-encryption mechanism, vector commitment and Merkle tree, secure encryption of data and verifiability of query results are achieved, while protecting user query privacy using homomorphic encryption and inadvertent pseudo-random functions.
It effectively solves the problems of data privacy leakage, inability to verify query results, and privacy protection for user query, realizes the secure storage of data and the reliability of query results, and reduces the overhead of computing and communications.
Smart Images

Figure BDA0005287246540000042 
Figure BDA0005287246540000046 
Figure BDA0005287246540000048
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of computer science and technology, and particularly relates to a method for implementing a secure and verifiable outsourced database with privacy protection. Background Art
[0002] With the rapid development of technologies such as big data and cloud computing, the amount of data collected and stored has increased exponentially, and many enterprises have obtained petabytes or even exabytes of massive data. With the rapid growth of the data volume, the cost of data storage has also risen sharply. A large number of newly added data require large-scale expansion of storage devices, resulting in a rapid increase in storage costs, and the development of many enterprises has been restricted accordingly. To solve the problem of excessively high storage costs, many enterprises choose to outsource the storage and computing tasks of the database to professional cloud storage service providers. Database as a Service (DBaaS), as a new outsourcing model, can reduce the complexity and cost of enterprise database management, achieve unified monitoring and management of the database, and further lower the threshold of enterprise database applications, thus attracting wide attention. Data owners can outsource their large databases to cloud service providers, so that users with limited resources can query the data in the database. However, although DBaaS has many advantages, there are still some disadvantages. For example, there may be a large amount of private data in the database of the data owner, and the cloud service provider cannot steal this data. In addition, users cannot ensure that the cloud server truly returns the query data. In fact, many studies have shown that cloud servers cannot be fully trusted. The server may attempt to obtain the content of the outsourced data, obtain the privacy of the query user, or even forge query results. Without a mechanism to verify the query results, the server can return forged data to the user to impersonate the query results, and the user will not be able to tell the difference. Therefore, researching a method that can verify the correctness and integrity of query results has become an urgent problem to be solved currently.
[0003] Currently, due to the importance of verifiable databases, a large number of researchers have started to study how to enable users to verify the correctness of query results and have proposed various solutions. A variety of solutions have been proposed for verifiable databases. The current solutions are mainly divided into two modes, namely zero-knowledge proof-based and Merkle tree-based. Although these solutions can ensure the verifiability of query results, most of them outsource the unencrypted database to the server, resulting in data leakage in the database. Therefore, many researchers are concerned about this security risk and consider encrypting the data and then outsourcing the encrypted database to the server. However, few researchers have paid attention to the privacy leakage problem of user access patterns. If users directly send real query indexes to the server during queries, adversaries can obtain a large amount of useful information from the query indexes. Since the data owner is not completely trustworthy, he may be curious about the user's query indexes (half-truths), so it is also crucial to hide the query indexes from the data owner so that the data owner cannot directly decrypt the encrypted data. In short, decrypting the encrypted data without the participation of the data owner and a trusted third party and how to prevent the data owner and the server from obtaining the privacy of the user's query indexes are two challenges faced by outsourced encrypted databases. It is necessary to propose a secure and efficient outsourced database solution that can solve all the above problems. Summary of the Invention
[0004] The purpose of the present invention is to solve the problems of data privacy leakage, unverifiable query results, and user query privacy protection in outsourced databases, and at the same time ensure that users can securely decrypt the data through proxy re-encryption.
[0005] To achieve the above object, the present invention adopts the following technical solutions:
[0006] The present invention provides a method for implementing a privacy-protected secure verifiable outsourced database, including the following steps:
[0007] System initialization: Generate system public parameters, public-private key pairs for the data owner and users, and proxy re-encryption keys;
[0008] Preprocessing stage: The data owner rearranges and encrypts the database to generate an encrypted database;
[0009] Outsourcing stage: The data owner constructs a Merkle tree based on the encrypted database and calculates vector commitments, and sends the encrypted database and commitment values to the cloud server;
[0010] Query encoding stage: The user generates a query matrix to protect query privacy and sends the query matrix to the cloud server;
[0011] Query phase: The cloud server generates query results and proofs based on the query matrix, and sends the query results, proofs, and key node matrix to the user;
[0012] Verification phase: The user verifies the correctness and integrity of the query results based on the received proofs and key node matrix;
[0013] Decryption phase: The user and the cloud server complete data decryption through the proxy re-encryption mechanism to ensure that the user can obtain the plaintext of the query data.
[0014] In the above scheme, the system initialization includes the following steps:
[0015] Step 1.1: Randomly select two groups and of prime order, and the generator g of group and satisfy the bilinear mapping
[0016] Step 1.2: Randomly select and denotes randomly selecting numbers from the set of integers modulo p, and e(·) denotes the bilinear mapping;
[0017] Step 1.3: Given the security parameter λ, the data owner DO and the user U respectively generate the private key sk = (x 1 , x 2 ) and the public key where x 1 , x 2 ∈ R Z p , ∈ R denotes randomly selected from Z p ;
[0018] Step 1.4: The data owner side DO calculates and Then the public parameter of the system is PP:
[0019]
[0020] Step 1.5: Given the public key of user U and the private key of data owner DO,
[0021]
[0022] rk DO→UDenotes the proxy re-encryption key from the data owner to user U, Denotes a part of the private key of the data owner, Denotes a part of the private key of user U.
[0023] In the above scheme, the preprocessing stage includes the following steps:
[0024] Step 2.1: For a database DB = {(key, value)} consisting of n data pairs, the data owner DO rearranges it into a data matrix with rows and columns
[0025] Step 2.2: The data owner DO randomly selects s ∈ R Z p and encrypts each piece of data according to the following formula to obtain the encrypted database:
[0026]
[0027] where ∈ R Denotes random selection, x 1 DO s represents the s-th power of;
[0028] Step 2.3: Finally, the data owner DO sets the maximum number of user queries Max U , and sets the query counter to 0.
[0029] In the above scheme, the outsourcing stage includes the following steps:
[0030] Step 3.1: For the i-th row of the encrypted database Finally, the data owner DO takes this piece of data as the leaf node of the Merkle tree and calculates the corresponding root node root i , obtaining the root node set
[0031] Step 3.2: Next, the data owner DO calculates the digest dig i = root i ||h(root||i), where where h(·) represents the hash function, and root||i represents concatenating root and i into a string;
[0032] Step 3.3: The data owner DO calculates the commitment:
[0033] Step 3.4: Finally, the data owner DO publicly discloses the commitment C and sends the encrypted database to the server S.
[0034] In the above scheme, during the query encoding phase, when the user needs to query the data in the i 0 -th row and j 0 -th column, the user needs to encode the query request to generate a query vector V to protect their query privacy. The specific steps are as follows:
[0035] Step 4.1: For the first row of the query matrix V[i][j], the user U will set the values according to the following formula:
[0036]
[0037] where E(0) is the homomorphic ciphertext of the number 0 after homomorphic encryption, and E(1) is the homomorphic ciphertext of the number 1 after homomorphic encryption;
[0038] Step 4.2: The user U needs to query the data at the position (i 0 , j 0 ). First, construct a query matrix according to the critical path of the queried node. For the value V[i][j] in the i-th row and j-th column of the query matrix, the formula is as follows:
[0039]
[0040] where E(0) is the ciphertext of the number 0 after homomorphic encryption, and E(1) is the ciphertext of the number 1 after homomorphic encryption;
[0041] Step 4.3: Finally, the user U sends the constructed query matrix to the server.
[0042] In the above scheme, during the query encoding phase, after the server S receives the query matrix sent by the user U, it obtains the query result according to the query matrix V and constructs a proof π, and sends the query result, the proof, and i 0 to the user U together. The specific steps are as follows:
[0043] Step 5.1: The server S checks whether the query counter meets the following formula:
[0044] q U ≤Max U
[0045] where q U represents the query counter of the user U, and Max U is the maximum query quantity of the user U. If it is satisfied, continue with the subsequent query steps; otherwise, reject the query.
[0046] Step 5.2: The server S multiplies the first row of the query matrix by the \(i\)th 0 row of the database, and multiplies the \(j\)th row of the query matrix by the \(j\)th 0 layer of the Merkle tree corresponding to the \(i\)th row of the database. Thus, the node matrix NV is as follows:
[0047]
[0048] Step 5.3: The server S constructs a Merkle tree based on the data 0 in the \(i\)th row of the database to obtain the root node Connect it with the row serial number \(i\) 0 and calculate the hash value. Finally, obtain the digest
[0049] Step 5.4: The server S finally calculates the following proof
[0050]
[0051] and sends the proof \(\pi\) and the key node matrix NV to the user U.
[0052] In the above scheme, in the verification phase, after receiving the proof \(\pi\) and the key node matrix NV sent by the server, the user U verifies the query result, which specifically includes the following steps:
[0053] Step 6.1: The user U extracts the corresponding values E (key node values) of the Merkle tree key nodes in NV, and the query result ciphertext and decrypts it using the key to obtain all the key node values and
[0054] Step 6.2: Construct a Merkle tree based on the key node values and obtain the root node value
[0055] Step 6.3: Connect the obtained root node value with the serial number \(i\) 0 and calculate the hash value to obtain the digest value
[0056] Step 6.4: U calculates whether the following formula holds:
[0057]
[0058] If it holds, it means that the query result returned by the server is true and complete; if it does not hold, it means that the query result has been tampered with, and the result is rejected.
[0059] In the above solution, in the decryption phase, user U and server S jointly complete the decryption operation through one round of communication, ensuring that server S does not know the data that user U needs to decrypt, while user U can complete the data decryption operation. The specific steps are as follows:
[0060] Step 7.1: For the database ciphertext that passes the verification User U randomly selects a random number and calculates the processing value a = (g s ) r , and sends a to server S;
[0061] Step 7.2: After receiving a, server S calculates the re-encrypted ciphertext EDB[i 0 [j 0 DO→U , and sends the re-encrypted ciphertext to user U:
[0062]
[0063] rk DO→U represents the proxy re-encryption key generated by data owner DO for user U, represents a part of user U's private key;
[0064] Step 7.3: After obtaining the re-encrypted ciphertext, user U first processes the data to obtain
[0065]
[0066] Step 7.4: Finally, user U decrypts using its own private key as follows:
[0067]
[0068] represents the encrypted data item in the i 0 th row and j 0 th column of the encrypted database EDB.
[0069] Since the present invention adopts the above technical means, it has the following beneficial effects:
[0070] 1. Through the combination of system initialization, preprocessing, and outsourcing phases, the problems of data privacy protection and the security of the outsourced database are solved, achieving the effects of data encryption and secure storage.
[0071] In the system initialization stage, by generating system public parameters, public-private key pairs, and proxy re-encryption keys, the security of data and the authentication of user identities are ensured. In the preprocessing stage, the data owner rearranges and encrypts the database to generate an encrypted database, further protecting the privacy of the data. In the outsourcing stage, the data owner constructs a Merkle tree based on the encrypted database and calculates vector commitments to ensure the integrity and verifiability of the data. The combination of these steps fully protects the data during storage and transmission, preventing data leakage and tampering.
[0072] 2. Through the combination of query encoding and the query stage, the problem of user query privacy protection is solved, achieving the effect of hiding the user's query intention.
[0073] In the query encoding stage, the user generates a query matrix and protects the query privacy through homomorphic encryption technology, ensuring that the cloud server cannot directly obtain the user's query intention. In the query stage, the cloud server generates query results and proofs based on the query matrix and returns the results to the user. In this way, the user's query privacy is effectively protected, and the cloud server cannot infer the specific query content of the user, thus preventing the leakage of query privacy.
[0074] 3. Through the mechanism of the verification stage, the problem of query result verifiability is solved, achieving the effect of ensuring the correctness and integrity of the query results.
[0075] In the verification stage, the user reconstructs the Merkle tree and verifies the hash value of the root node based on the received proofs and the critical node matrix to ensure the correctness and integrity of the query results. If the verification passes, the user can confirm that the query results have not been tampered with; otherwise, the user will reject the results. This mechanism effectively prevents the cloud server from returning forged query results and ensures the reliability of the query results.
[0076] 4. Through the proxy re-encryption mechanism in the decryption stage, the problem of users decrypting data securely is solved, achieving the effect that users can obtain plaintext data securely.
[0077] In the decryption stage, the user and the cloud server conduct a round of communication through the proxy re-encryption mechanism, ensuring that the cloud server does not know the specific data that the user needs to decrypt, while the user can decrypt the data securely. Through this mechanism, the user can obtain plaintext data without exposing the decryption key, ensuring the security of the data and the privacy of the user.
[0078] 5. By optimizing the combination of vector commitments and Merkle trees, the efficiency problem of large-scale database queries is solved, achieving the effect of efficient querying and verification.
[0079] In the preprocessing and outsourcing phases, the data owner optimizes the verification process of query results by constructing a Merkle tree and calculating vector commitments. Compared with traditional methods such as Lagrange interpolation, this scheme only requires modular exponentiation and modular multiplication operations, significantly reducing the computational overhead. Especially in large-scale database queries, the computational efficiency of this scheme is significantly better than other schemes, making it suitable for handling the query and verification requirements of large-scale data.
[0080] 6. By combining homomorphic encryption and proxy re-encryption, the potential threats to user query privacy from the data owner and the cloud server are addressed, achieving the effect of dual privacy protection.
[0081] In the query encoding phase, the user hides the query intention through homomorphic encryption technology to prevent the cloud server from obtaining the user's query privacy. In the decryption phase, through the proxy re-encryption mechanism, it is ensured that the data owner cannot directly decrypt the user's data. This dual privacy protection mechanism effectively prevents the potential threats to user query privacy from the data owner and the cloud server, ensuring the security of user data.
[0082] 7. Through experimental comparison, the advantages of this scheme in terms of computational overhead and communication overhead are verified, achieving the effect of efficiently processing large-scale database queries.
[0083] The experimental results show that the computational overhead of this scheme in the offline and online phases is significantly lower than other schemes. Especially in large-scale database queries, the computational efficiency advantage is more obvious. In addition, this scheme also shows high efficiency in terms of communication overhead. Especially when querying large-scale databases, the communication overhead is only 1.37% of other schemes. These experimental results prove the efficiency and practicality of this scheme in processing large-scale database queries. Description of the Drawings
[0084] Figure 1 is the overall process diagram;
[0085] Figure 2 is the Merkle tree;
[0086] Figure 3 is the PPSVDB system model;
[0087] Figure 4 is the computational overhead diagram for the initialization phase;
[0088] Figure 5 is the computational overhead diagram for the query encoding phase;
[0089] Figure 6 is the computational overhead diagram for the verification and decryption phases;
[0090] Figure 7 is the overall computational overhead diagram for the online phase;
[0091] Figure 8 It is a comparison chart of communication overhead for three schemes. Detailed implementation manners
[0092] The embodiments of the present invention will be described in detail below. Although the present invention will be described and explained in conjunction with some specific implementation manners, it should be noted that the present invention is not limited to these implementation manners only. On the contrary, any modifications or equivalent replacements made to the present invention should be covered within the scope of the claims of the present invention.
[0093] In addition, for better illustration of the present invention, numerous specific details are given in the following detailed implementation manners. Those skilled in the art will understand that the present invention can also be implemented without these specific details.
[0094] Aiming at the limitations of the above-verifiable database, the present invention proposes a privacy-preserving secure verifiable outsourced database. Based on vector commitment and Merkle tree, this method proposes an optimized vector commitment verification scheme (abbreviated as Mer-VectorCommitment) to further reduce the verification and initialization efficiency. Based on this optimized vector commitment verification scheme, by introducing oblivious pseudorandom function (OPRF) and homomorphic encryption to protect user query privacy, and finally through proxy re-encryption to ensure that users can decrypt the encrypted data by themselves, a complete privacy-preserving secure verifiable outsourced database (abbreviated as PPSVDB) is realized.
[0095] Optimized vector commitment verification scheme
[0096] The optimized vector commitment verification scheme can provide verification of query results to users when they obtain the query results. The optimized vector commitment verification scheme includes the following definitions:
[0097] Definition 1: The scheme involves two parties, namely the committer and the verifier. The committer makes a commitment to the data and opens it later. The verifier verifies the data after the committer opens the commitment.
[0098] Definition 2: A Merkle tree is a binary tree composed of a root node, a set of internal nodes, and a set of leaf nodes. The leaf nodes store the hash values of data blocks, the internal nodes store the hash values of their two child nodes, and the root node also stores the hash values of the contents of its two child nodes. For example, the value stored in internal node n is: where n l and n 2is a child node of n, and H is a one-way, collision-resistant hash function. The values stored in the entire tree are calculated iteratively from the bottom up. Based on the value stored in the Merkle tree root node, any tampering with the recorded data can be detected, thus effectively ensuring data integrity.
[0099] Figure 2 shows a typical Merkle tree. To verify the integrity of data m and m7 in the figure, a verification path containing VP = h(m 1 ), N 10 , N 11 , h(m 8 ) is required. Using h(m 2 ), h(m 7 ) and the verification path VP, the root node can be reconstructed. Then, it is compared with the previously saved root node. If they are equal, it means that m 2 and m 7 have not been tampered with.
[0100] Next, the present invention will further introduce the technical concept of the present invention in detail:
[0101] Secure Verifiable Outsourced Database Scheme for Privacy Protection
[0102] Based on the optimized vector commitment verification scheme, the present invention designs a secure verifiable outsourced database scheme for privacy protection. The scheme is mainly divided into 7 stages, namely system initialization, preprocessing, data outsourcing, query encoding, query, verification, and decryption. The secure verifiable outsourced database scheme for privacy protection includes the following definitions:
[0103] Definition 1: As Figure 3 shown, the PPSVDB system includes three entities: the data owner (DO), the cloud server (S), and the user (U).
[0104] Definition 2: Data owner (DO): The DO has a database containing a large amount of data. Due to limitations in computing and storage resources, the DO hopes to outsource the database to the cloud server. The DO first generates corresponding re-encryption keys for U based on the public key. Then, the DO rearranges the database into a matrix and encrypts the database to obtain EDB. For each row of EDB, the DO constructs a Merkle tree and stores its root node. Next, the DO calculates a vector commitment for the root node and makes it public. Finally, the DO sends the vector commitment C, the encrypted database EDB[i][j], the re-encryption key for U, and the query counter to U.
[0105] Definition 3: Server (S): S has a large amount of storage and computing resources and can provide data outsourcing storage services. After U submits a query request to S, S will respond to the request and return the query result and the proof π to U. In addition, when the data owner submits a proxy re-encryption request for the data, S will re-encrypt the data and return the result to U.
[0106] Definition 4: User (U): U hopes to query the data in the i-th row and j-th column of the database. First, U sends a query request for the i-th row to S. When U receives the query result and the proof π, U will verify whether the result is valid. If the result is valid, U will then submit a proxy re-encryption request to S through an oblivious pseudo-random function (OPRF). After receiving the re-encrypted ciphertext, U uses the private key to decrypt and obtain the query data.
[0107] Next, the secure and verifiable outsourced database scheme for privacy protection will be described in detail.
[0108] Step 1: System initialization. The algorithm in the system initialization phase will generate the public parameters of the scheme, the public and private keys of the user and the data owner, and the proxy re-encryption key.
[0109] Step 1.1: Randomly select two groups and as well as the generator g of the group .
[0110] The groups and satisfy the bilinear mapping
[0111] Step 1.2: Randomly select and
[0112] Step 1.3: Given the security parameter λ, the data owner DO and the user U respectively generate the private key sk = (x 1 , x 2 ) and the public key where x 1 , x 2 ∈ R Z p (∈ R means randomly selected from Z p ).
[0113] Step 1.4: DO calculates and Then the public parameters of the system are
[0114] Step 1.5: Given the public key of the user U and the private key of the data owner DO The DO calculates the following re-encryption key:
[0115]
[0116] Step 2: Preprocessing stage. In the preprocessing stage, the DO rearranges and encrypts the database.
[0117] Step 2.1: For a database DB = {(key, value)} consisting of n data pairs, the data owner DO rearranges it into a data matrix with rows and
[0118] columns R Z p and encrypts each piece of data according to the following formula to obtain the encrypted database:
[0119]
[0120] Step 2.3: Finally, the DO sets the maximum number of user queries Max U , and sets the query counter to 0.
[0121] Step 3: Outsourcing stage. In the data outsourcing stage, the DO constructs a Merkle tree based on the database and calculates the commitment.
[0122] Step 3.1: For the i-th row of the encrypted database the DO takes this data as the leaf node of the Merkle tree and calculates the corresponding root node root i . Obtain the root node set
[0123] Step 3.2: Next, the DO calculates the digest dig i = root i ||h(root||i), where
[0124] Step 3.3: The DO calculates the commitment
[0125]
[0126] Step 3.4: Finally, the DO makes the commitment C public and sends the encrypted database to the server S.
[0127] Step 4: Query encoding stage. In the query encoding stage, the user needs to query the i-th 0 row j 0When querying the data in a column, it is necessary to encode the query request to generate a query vector V to protect the query privacy of oneself.
[0128] Step 4.1: For the first row of the query matrix V[i][j], U will set the values according to the following formula:
[0129]
[0130] Among them, E(0) is the homomorphic ciphertext after homomorphic encryption of the number 0, and E(1) is the homomorphic ciphertext after homomorphic encryption of the number 1.
[0131] Step 4.2: U needs to query the data at the position (i 0 , j 0 ). First, construct a query matrix according to the critical path of the queried node. For the value V[i][j] in the i-th row and j-th column of the query matrix, its formula is as follows:
[0132]
[0133] Among them, E(0) is the ciphertext after homomorphic encryption of the number 0, and E(1) is the ciphertext after homomorphic encryption of the number 1. Here, take a database containing 64 elements as an example for illustration. Suppose you want to query the data EDB in the 3rd column of the i-th row i3 , then the query matrix is:
[0134]
[0135] Step 4.3: Finally, U sends the constructed query matrix to the server.
[0136] Step 5: In the query phase, the server S receives the query matrix sent by U, obtains the query result according to the query matrix V, constructs a proof π, and sends the query result, the proof and i 0 to U together.
[0137] Step 5.1: First, the server S checks whether the query counter meets the following formula:
[0138] q U ≤ Max U
[0139] Among them, q U represents the query counter of user U, and Max U is the maximum query quantity of user U. If the above formula is satisfied, continue with the subsequent query steps, otherwise reject the query.
[0140] Step 5.2: S will compare the first row of the query matrix with the i-th 0 row Multiply the $i$-th row of the query matrix by the $i$-th layer of the Merkle tree corresponding to the $i$-th row in the database. Thus, the node matrix $NV$ is as follows: 0
[0141]
[0142] Here, a simple proof is given. Similar to the assumption in Step 4.2, assume that the query $EDB$ 1,2 At this time, the node matrix $NV$ is:
[0143]
[0144] Step 5.3: $S$ constructs a Merkle tree based on the data in the $i$-th row of the database 0 to obtain the root node Connect it with the row serial number $i$ 0 and calculate the hash value, and finally obtain the digest
[0145] Step 5.4: $S$ finally calculates the following proof
[0146]
[0147] and sends the proof $\pi$ and the key node matrix $NV$ to the user $U$.
[0148] Step 6: Verification phase. After receiving the proof $\pi$ and the key node matrix $NV$ sent by the server, $U$ verifies the query result.
[0149] Step 6.1: $U$ extracts the corresponding values $E$ (key node values) of the Merkle tree key nodes in $NV$, and the query result ciphertext Decrypts it using the key to obtain all key node values and
[0150]
[0151] Step 6.2: Constructs a Merkle tree based on the key node values and obtains the root node value
[0152]
[0153]
[0154] Step 6.4: $U$ calculates whether the following formula holds:
[0153]
[0154] If it holds, it means that the query result returned by the server is true and complete. If it does not hold, it means that the query result has been tampered with, and the result is rejected.
[0155] Step 7: Decryption phase. In the decryption phase, the user U and the server complete the decryption operation through a round of communication, ensuring that the server S cannot know the data that the user U needs to decrypt, while the user U can complete the data decryption operation.
[0156] Step 7.1: For the verified database ciphertext The user U randomly selects a random number and calculates the processed value a = (g s ) r , and sends a to S.
[0157] Step 7.2: After receiving a, S calculates the re-encrypted ciphertext:
[0158]
[0159] and sends the re-encrypted ciphertext to the user U.
[0160] Step 7.3: After obtaining the re-encrypted ciphertext, the user U first processes the data to obtain
[0161]
[0162] Step 7.4: Finally, the user U decrypts using its own private key as follows:
[0163]
[0164] Experimental results:
[0165] To more accurately test the performance of the scheme, based on the PBC library (pairing-based cryptography library), code was written using the C++ programming language and experiments were conducted in a simulated environment, and comparisons were made with PPVDB[1| and APIR[2]. The schemes were all run on a machine with an Intel(R) Xeon(R) Gold CPU@2.10GHZ processor and 192GB of RAM, and the operating system was Ubuntu 20.04.6. In addition, the scheme was divided into an offline and an online phase. The offline phase included system initialization, data preprocessing, and data outsourcing. The online phase included two sub-phases: the query and response phase and the verification and decryption phase. The user U and the server S collaborated to obtain the query result and perform verification and decryption.
[0166] Since APIR[2] stores the plaintext database, there is no computational overhead during the data outsourcing process. Therefore, in this phase, only the overheads of PSVDB and PPVDB[1] were compared. Figure 4Shows the comparison of the computational overhead in the offline phase. Database encryption is the most time-consuming task in the offline phase. As can be seen from the figure, the computational overhead of PPVDB[1] and the proposed scheme PSVDB of this application in the offline phase increases with the increase of the database size n. Specifically, for a database of size 2 20 in PPVDB|1|, the data owner DO needs 323005.78s to generate the encrypted database, while PSVDB only needs 2647.15s, which is 0.82% of PPVDB[1]. This is because in addition to both needing to perform the same encryption operation, PPVDB[1] also needs to perform Lagrange interpolation on n data points, while PSVDB only involves exponentiation modulo operations and multiplication modulo operations. When n is very large, the operations of PSVDB are more efficient than Lagrange interpolation.
[0167] For the online phase, this phase is divided into the query and response phase and the verification and decryption phase.
[0168] Figure 5 Shows the computational overhead of the three schemes in generating query vectors and response vectors when the database size n ranges from 2 8 to 2 20 . As can be seen from the figure, both PPVDB[1] and APIR[2] show a steep growth trend, while the growth slope of PSVDB is significantly lower. Especially as n increases, the computational efficiency of PSVDB is higher than that of PPVDB[1] and APIR[2]. Taking n = 2 8 as an example, the computational costs of PPVDB[1] and APIR[2] are 0.31s and 16.50s respectively. In contrast, PSVDB takes 0.16s, which is only 51.6% and 0.097% of PPVDB[1] and APIR[2]. When n = 2 20 , the computational costs of PPVDB[1] and APIR[2] rise to 1271.37s and 64654.92s respectively, while PSVDB only needs 10.20s, which is only 0.80% and 0.016% of PPVDB[1] and APIR[2]. As n increases, the efficiency advantage of PSVDB over APIR[2] and PPVDB[1] becomes more obvious, and it is more suitable for efficiently processing large-scale database queries. This is because APIR[2] needs to perform n exponentiation modulo operations and n multiplication modulo operations to generate query vectors and query responses respectively. PPVDB[1] needs to calculate an n-degree polynomial through Lagrange interpolation to find an (n - 1)-degree polynomial. In contrast, the PSVDB scheme only needs to perform homomorphic multiplication to obtain the response vector, which has higher efficiency.
[0169] Figure 6 shows the comparison of the computational overhead of the three schemes in the verification and decryption phases. As can be seen from the figure, when the database size n changes from 2 8 to 2 20 , the computational overhead of APIR[2] and PPVDB[1] remains approximately constant at about 5 ms. When n = 2 8 , the computational overhead of PSVDB is 15.1 ms. When n = 2 20 , the computational overhead of PSVDB is 27.4 ms and increases linearly with the size of n. This is because the scheme of this application requires an additional proxy re-encryption process to obtain the plaintext of the data. However, compared with the query and response phases, the computational time of this phase is in the millisecond level and can be ignored. As can also be seen from Figure 7 , the total computational overhead in the online phase is mainly related to the query and response phases, and Figure 7 is Figure 5 almost the same.
[0170] (2) Communication overhead
[0171] PSVDB and PPVDB[1] protect the security of data through encryption. Both involve 5 rounds of communication, while APIR[2] only contains 2 rounds of communication because it stores data in plaintext. Since the encryption and outsourcing of the database are only completed once in the offline phase, this application does not analyze the communication overhead in the encryption and outsourcing phases of the database, but only analyzes the communication overhead in the frequently executed online phase.
[0172] The figure shows the comparison of the three schemes in terms of communication overhead. Figure 8 (a) in depicts the communication overhead from the client U to the server S. PPVDB[1] only sends the obfuscated index in this phase. APIR[2] sends a vector containing n group elements. PSVDB sends homomorphic ciphertexts to the server S. Therefore, the communication overhead of PPVDB[1] is relatively small, while the communication overhead of PSVDB is between APIR[2] and PPVDB[1]. Figure 7 (b) in illustrates the communication overhead from the server S to the client U. Both APIR[2] and PPVDB[1] return the proof π and the data ciphertext. Therefore, the communication overhead of the two schemes is similar. In contrast, PSVDB returns a node vector NV containing all the key verification nodes. Therefore, its overhead increases with the growth of n. Figure 7Figure (c) shows the total communication overhead in the query phase. It can be seen that PPVDB[1] has a very low communication overhead of about 0.0024KB because its query request only contains one index, which is significantly lower than that of APIR[2] and PSVDB. Among the remaining two schemes, compared with APIR[2], the proposed scheme PSVDB has higher communication efficiency. For example, when n = 2 20 , the communication overhead of APIR[2] is about 65536KB, while that of PSVDB is 900KB, only 1.37% of APIR[2]. Therefore, it can be considered that the proposed scheme PSVDB has a certain advantage in communication overhead.
Claims
1. A privacy-preserving, secure and verifiable outsourcing database implementation method, characterized in that: The following steps are involved: System initialization: Generate system public parameters, public and private key pairs of data owners and users, and proxy re-encryption keys; Preprocessing stage: The data owner rearranges and encrypts the database to generate an encrypted database; Outsourcing stage: The data owner builds a Merkle tree based on the encrypted database, calculates the vector commitment, and sends the encrypted database and commitment value to the cloud server; Query encoding phase: The user generates a query matrix to protect query privacy and sends the query matrix to the cloud server; Query phase: The cloud server generates query results and proofs based on the query matrix, and sends the query results, proofs, and key node matrix to the user; Verification phase: The user verifies the correctness and completeness of the query results based on the received proof and key node matrix; Decryption phase: The user and the cloud server complete data decryption through a proxy re-encryption mechanism to ensure that the user can obtain the plaintext of the query data.
2. A privacy-protected secure and verifiable outsourcing database implementation method according to claim 1, characterized in that: System initialization includes the following steps: Step 1.1: Randomly select two groups of order prime, and And group The generator g of the group and Satisfies the bilinear map e: Step 1.2: Random Selection and represents a random selection from the set of integers modulo p number, e(·) represents a bilinear mapping; Step 1.3: Given the security parameter λ, the data owner DO and user U generate a private key sk = (x1, x2) and a public key in ∈ R Indicates that from Z p Randomly selected from Step 1.4: DO calculation on the data owner side and Then the public parameter of the system is PP: Step 1.5: Given the public key of user U and the private key of the data owner DO The data owner DO calculates the following re-encryption key: rk DO→U represents the proxy re-encryption key from the data owner to user U, represents a portion of the data owner's private key, Represents a portion of the private key of user U.
3. A privacy-protected secure and verifiable outsourcing database implementation method according to claim 2, characterized in that: The preprocessing phase includes the following steps: Step 2.1: For a database DB = {(key, value)} consisting of n data pairs, the data owner DO rearranges them into a database with OK Data matrix of columns Step 2.2: The data owner DO randomly selects s∈ R Z p And encrypt each piece of data according to the following formula to obtain an encrypted database: where ∈ R represents random selection, express s to the power of; Step 2.3: Finally, the data owner DO sets the maximum number of user queries Max U , and set the query counter to 0.
4. A privacy-protected secure and verifiable outsourcing database implementation method according to claim 3, characterized in that: The outsourcing phase includes the following steps: Step 3.1: Encrypt the row of the database Finally, the data owner DO will The data is used as the leaf node of the Merkle tree, and the corresponding root node root is calculated i , get the root node set Step 3.2: Next, the data owner DO calculates the summary dig i =root i ||h(root||i), where Where h(·) represents a hash function, and root||i represents concatenating root and i into a string; Step 3.3: The data owner DO calculates the commitment: Step 3.4: Finally, the data owner DO makes the commitment C public and sends the encrypted database to the server S.
5. A privacy-protected secure and verifiable outsourcing database implementation method according to claim 4, characterized in that: In the query encoding stage, the user needs to query the data in row i0 and column j0. The query request needs to be encoded to generate a query vector V to protect the privacy of the query. The specific steps include: Step 4.1: For the first row of the query matrix V[i][j], user U will set the value according to the following formula: Among them, E(0) is the homomorphic ciphertext of the number 0 after homomorphic encryption, and E(1) is the homomorphic ciphertext of the number 1 after homomorphic encryption; Step 4.2: User U needs to query the data at the location (i0, j0) First, construct a query matrix based on the critical path of the node to be queried. For the value V[i][j] in the i-th row and j-th column of the query matrix, the formula is as follows: Among them, E(0) is the ciphertext of the number 0 after homomorphic encryption, and E(1) is the ciphertext of the number 1 after homomorphic encryption; Step 4.3: Finally, user U sends the constructed query matrix to the server.
6. A privacy-protected secure and verifiable outsourcing database implementation method according to claim 5, characterized in that: In the query encoding phase, server S receives the query matrix sent by user U, obtains the query result according to the query matrix V, constructs proof π, and sends the query result, proof and i0 to user U. The specific steps include: Step 5.1: Server S checks whether the query counter meets the following formula: q U ≤Max U Among them, q U Indicates the query counter of user U, Max U The maximum number of queries for user U. If it is satisfied, the query will continue to the subsequent query steps, otherwise the query will be rejected; Step 5.2: Server S compares the first row of the query matrix with the i0th row of the database Multiply, the ,th row of the query matrix is multiplied by the ,th layer of the Merkle tree corresponding to the database row i0, so the node matrix NV is as follows: Step 5.3: Server S calculates the data in row i0 of the database. Construct a Merkle tree and get the root node Connect it with the row sequence number i0 and calculate the hash value to get the summary Step 5.4: Server S finally calculates the following proof, And send the proof π and key node matrix NV to user U.
7. A privacy-protected secure and verifiable outsourcing database implementation method according to claim 6, characterized in that: In the verification phase, user U verifies the query result after receiving the proof π and key node matrix NV sent by the server, which specifically includes the following steps: Step 6.1: User U extracts the corresponding value E (key node value) of the Merkle tree key node in NV, and the query result ciphertext After decryption using the key, all key node values and Step 6.2: Based on the key node values and Construct a Merkle tree and get the root node value Step 6.3: Get the root node value Connect with the serial number i0 and calculate the hash value to obtain the summary value Step 6.4: U calculates whether the following formula is true: If true, it means that the query result returned by the server is true and complete. If false, it means that the query result has been tampered with and the result is rejected.
8. A privacy-protected secure and verifiable outsourcing database implementation method according to claim 7, characterized in that: In the decryption phase, user U and server S complete the decryption operation through a round of communication, ensuring that server S cannot know the data that user U needs to decrypt, while user U can complete the data decryption operation. The specific steps are as follows: Step 7.1: For the verified database ciphertext User U randomly selects a random number And calculate the processing value a=(g s ) r , send a to server S; Step 7.2: After receiving a, server S calculates the re-encrypted ciphertext EDB[i0][j0] DO→U , and send the re-encrypted ciphertext to user U: rk DO→U represents the proxy re-encryption key generated by the data owner DO for user U, Represents a part of the private key of user U; Step 7.3: After user U obtains the re-encrypted ciphertext, he first processes the data to obtain Step 7.4: Finally, user U uses his own private key to decrypt as follows: Represents the encrypted data item at row i0 and column j0 in the encrypted database EDB.