A method for deduplicating similar data ciphertexts
By designing similar ciphertext deduplication methods in cloud storage systems, using multiplication loop group and hash function to perform file label matching and error correction information calculation, the problem of similar file deduplication failure in the prior art is solved, efficient storage and transmission is achieved, and data security and integrity verification are enhanced.
Patent Information
- Application Number
- CN202410392012.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-04-02
- Publication Date
- 2025-06-10
- Estimated Expiration
- 2044-04-02
AI Technical Summary
The existing ciphertext deduplication technology fails when processing similar but not exactly the same files, cannot effectively deduplicate, and there are problems with security and integrity verification.
A similar ciphertext deduplication method is designed, through collaboration between cloud storage server, auxiliary verification server and data user, multiplication loop group and hash function to perform file label matching and error correction information calculation, to realize file deduplication and proof of ownership, and ensure data security through integrity verification.
On the premise of ensuring data security, the deduplication of similar ciphertext data is achieved, which reduces storage space usage and data transmission volume, improves storage and transmission efficiency, and enhances the privacy protection and integrity verification capabilities of data.
Smart Images

Figure CN118364486B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of information security where ciphertext needs to be processed, optimized for storage and transmission, and particularly relates to a method for deduplicating similar data ciphertexts. Background Art
[0002] In the context of the ever-growing data volume, the efficient utilization of storage resources has become an important challenge. As an optimization means, ciphertext deduplication technology can effectively reduce the storage space occupied by duplicate encrypted data while maintaining data security, thereby achieving effective management and savings of storage resources. At the same time, during data transmission, by eliminating redundant information in the ciphertext, ciphertext deduplication helps improve the efficiency of data transmission, reduce the amount of data transmitted, and reduce transmission costs and bandwidth pressure, which is particularly significant for resource-constrained Internet of Things devices and edge computing environments. In addition, the research on ciphertext deduplication technology is also crucial for data privacy protection. While protecting data privacy, it optimizes encrypted data to avoid the risk of plaintext data leakage and ensures data security and privacy. Furthermore, this technology also provides support for the research and verification in the field of cryptography. By analyzing and comparing ciphertexts, the security of encryption algorithms can be tested and evaluated, promoting further research and improvement in the field of cryptography. The research on ciphertext deduplication technology not only concerns the improvement of storage and transmission efficiency, but also involves data privacy protection and the development of the cryptography field, which has profound and important significance for the current information security and data management fields.
[0003] Currently, although the field of ciphertext deduplication can provide relatively reliable ciphertext deduplication technology, the comparison it uses is exact comparison. Therefore, it can only perform deduplication on exactly the same files or file blocks. Even if two files are very similar, or there are very few differences in different parts of the file, the deduplication of files and file blocks using exact comparison will lose its deduplication effect. In the fields of images and videos, some differences in files are not visible to the naked eye, but their file data is different during the comparison process, and the previous ciphertext deduplication technology cannot be used to achieve the deduplication function. Therefore, the field of using similar data deduplication requires further research. There are some limitations in the current research in this field. The encryption method has low security to support integrity verification, or there may be attacks that prevent it from achieving the expected integrity verification method.
[0004] Therefore, in order to simultaneously achieve ownership authentication and integrity verification, it is necessary to design an overall ciphertext deduplication architecture, study how to implement the function of ciphertext deduplication under similar plaintext data, and ensure the confidentiality of the ciphertext. Summary of the Invention
[0005] To achieve the above object, the present application provides a method for deduplicating similar ciphertexts, which is used to optimize the storage efficiency while ensuring data security.
[0006] The system to which the method is applied includes a cloud storage server, an auxiliary verification server, and a data client; the data client is divided into a first-time data uploading client and a subsequent data uploading client; the method includes the following processes: system initialization process, duplicate data detection process, file first-time uploading process, and file subsequent uploading process;
[0007] In the system initialization process, the cloud storage server generates a multiplicative cyclic group and two hash functions;
[0008] In the duplicate data detection process, the cloud storage server detects the file tags sent by the data client and performs matching calculations to confirm whether the current data client is a first-time data uploading client or a subsequent data uploading client;
[0009] In the file first-time uploading process, the first-time data uploading client generates error correction information and a convergence key for the file tag and block tag, a sub-secret ciphertext, an inner-layer key ciphertext of the convergence key, encrypts the file twice, generates an outer-layer secret key ciphertext, and uploads it to the cloud storage server;
[0010] In the file subsequent uploading process, the cloud storage server proves the ownership of the ciphertext file that matches the file tag uploaded by the subsequent data uploading client. The subsequent data uploading client generates an ownership proof and sends it to the cloud storage server for verification. When the verification passes, the cloud storage server distributes part of the ciphertext to the subsequent data uploading client, which calculates the convergence key and decrypts the ciphertext; the subsequent data uploading client also performs data integrity verification. After passing, the cloud storage server re-encrypts the ciphertext.
[0011] The specific system initialization process is as follows: The cloud storage server generates three multiplicative cyclic groups with a prime order p for initialization, which are Among them, g is a generator of the multiplicative cyclic group and the mapping relationship holds; g 1 is a generator of the multiplicative cyclic group ;
[0012] The two hash functions provided by the cloud storage server are:
[0013]
[0014]
[0015] The duplicate data detection process specifically includes the following steps:
[0016] Step 2.1: The data client of the data to be uploaded calculates the file tag Then send the file tag to the cloud storage server;
[0017] Step 2.2: After receiving the file tag, the cloud storage server performs a matching calculation in the pre-stored file tag set {T}, and tries to match a similar file tag. The matching calculation formula is as follows:
[0018]
[0019] Among them, the dis() function is the Hamming distance calculation function, and t is the threshold for determining file tag similarity;
[0020] If the above formula holds, it means that the cloud storage server may have similar ciphertext data, and the corresponding data client is the subsequent data upload client; if the above formula does not hold, it means that no similar file tag is detected, and the current data user is the first data upload client.
[0021] The initial value of t is 6.
[0022] The first upload process of the file includes the following steps:
[0023] Step 3.1: The first data upload client first calculates the error correction information of the file tag
[0024] Then divide the file F into n equal parts to form {B i}, i = 1, 2,..., n, and calculate And the error correction information corresponding to the file block B i Finally, the data client randomly selects t for the file block B i To generate comparable file block tags The formula is as follows:
[0025]
[0026] Step 3.2: Select a random value And calculate the secret value S, and determine the file block tag threshold k required for the file, and generate random numbers a 0 、a 1 、…a k-1 ,where a 0 = S, and construct a polynomial f(x) = a 0 +a 1 *x+a 2 *x 2 +…+a k-1 *x k-1 ,Use the file block codeword information to generate sub-secrets Finally, encrypt the corresponding sub-secret information using the file block codeword information:
[0027]
[0028] Step 3.3: Select a random value K and use this random value to encrypt the file F to obtain the inner ciphertext And calculate the hash value of the ciphertext file And encrypt the ciphertext hash. Select a random value r and calculate the ciphertext of the ciphertext hash;
[0029]
[0030] Step 3.4: Generate the inner key ciphertext corresponding to the convergence key;
[0031]
[0032] Step 3.5: Select a random value x again and use this random value to encrypt the inner ciphertext file C F_K Obtain the outer ciphertext C F_K_x ;
[0033]
[0034] Step 3.6: Select a random value r 1 Calculate the outer random key ciphertext;
[0035]
[0036] Upload the outer random key ciphertext to the cloud storage server.
[0037] The subsequent file upload process includes the following steps:
[0038] Step 4.1: The cloud storage server performs ownership proof on the ciphertext file that matches the file tag uploaded by the subsequent upload data client: The cloud storage server randomly selects from the file block identifiers of similar files Extract the corresponding file block numbers T = {k 1 ,…,k x}, and the corresponding parity bit information Send the corresponding parity bit information to the subsequent upload data client;
[0039] Step 4.2: After the subsequent upload data client receives the file block numbers and the corresponding parity bit information, calculate the block tag corresponding to each number where B' iFile blocks for similar files, and calculate the codeword information by fusing the locally calculated block tags and error correction bit information Select a random number t' i Generate comparable block tags
[0040]
[0041] and send them to the cloud storage server;
[0042] Step 4.3: After the cloud storage server receives the user's proof of ownership, verify the proof generated by the current data user, that is, verify whether the following formula holds:
[0043]
[0044] If the above formula holds, the file ownership verification is passed;
[0045] Step 4.4: After passing the file ownership verification, the cloud storage server will send the file tag error correction information Ciphertext hash ciphertext Inner layer key ciphertext C K and outer layer key ciphertext C x to the subsequent data upload user side;
[0046] Step 4.5: After the subsequent data upload user side receives the ciphertext message, it first generates the file tag codeword and then calculates the corresponding sub-secret information Perform the secret recovery process calculation S through the decrypted sub-secret information, and finally calculate the convergence key
[0047]
[0048]
[0049]
[0050] Step 4.6: Through the convergence key Parse out the corresponding inner layer random key information Ciphertext hash information Outer layer random key information x, and the calculation formula is as follows:
[0051]
[0052]
[0053] Then, the data user terminal uploads the data again and initiates a data integrity verification request to the auxiliary verification server, and sends the outer-layer random key information x to the auxiliary verification server;
[0054] Step 4.7: The auxiliary verification server initiates a ciphertext data request to the cloud storage server to obtain the outer-layer ciphertext data C F_K_x After that, it can perform outer-layer decryption using the outer-layer random key information x provided by the data user to obtain the inner-layer ciphertext And calculate the hash value of the inner-layer ciphertext file C' F_K of Return the calculated hash value to the data user terminal that uploads the data again;
[0055] Step 4.8: The subsequent data user terminal verifies whether the ciphertext data is still complete by comparing the hash values obtained from the auxiliary verification server and the cloud storage server. The comparison formula is
[0056] If the above formula holds, the integrity verification passes. The data user terminal selects a random value x' and calculates the re-encryption key and the outer-layer random key ciphertext C x , and uploads the re-encryption key rk and the outer-layer random key ciphertext C x to the cloud storage server; C x The calculation formula is as follows:
[0057]
[0058] Step 4.9: Ciphertext re-encryption: After receiving the re-encryption key rk, the cloud storage server performs the re-encryption process and replaces the outer-layer random key ciphertext C uploaded by the data uploaded again x .
[0059] The method further includes the download and decryption processes of the ciphertext file, and this process includes the following steps:
[0060] Step 5.1: The data user terminal initiates a ciphertext data request to the cloud storage server, and the cloud storage server returns the ciphertext data and its related information to the corresponding data user terminal;
[0061] Step 5.2: The data user terminal calculates other information, the inner-layer random key information ciphertext hash information outer-layer random key information x through the convergence key , removes the outer-layer ciphertext, and the removal formula is:
[0062]
[0063] By judging the equation Whether it holds or not is used to verify again whether the ciphertext data is complete. If it holds, it proves completeness. Finally, the inner-layer ciphertext is decrypted to obtain the plaintext file, and the decryption formula is:
[0064] The beneficial effects of this application are as follows:
[0065] The present invention deduplicates ciphertext data with similar plaintexts, and proposes an ownership proof and integrity verification method to ensure the user's verification of data ownership and the integrity audit of ciphertext data. Only by passing the ownership authentication can the encryption key be obtained with the help of the cloud storage server.
[0066] By the above means, the present invention can identify and eliminate duplicate ciphertexts generated by files with similar plaintexts in encrypted data, minimize the storage space occupied by duplicate encrypted data, facilitate the effective management of storage resources, and improve the efficiency of data storage.
[0067] By deduplicating similar ciphertexts, redundant information in data transmission can be reduced, the data transmission efficiency can be improved, and the transmission cost and bandwidth pressure can be reduced. The present invention improves the efficiency of data management while ensuring data privacy, and provides a novel and efficient solution for the field of information security. Brief Description of the Drawings
[0068] Figure 1 It is a schematic diagram of the system model applied in the present invention and the interaction process between various parts of the system. Detailed Embodiment
[0069] As Figure 1 shown, the system applied in the present invention includes a cloud storage server, an auxiliary verification server, and a data user terminal. In the process, according to whether the file data upload is a first upload, the data user terminal is divided into a first-upload data user terminal and a subsequent-upload data user terminal. Among them, the cloud storage server is a cloud storage server that provides storage services for the data user terminal and may deceive users for its own interests. The data user terminals are independent of each other, have a need for data storage but the local storage conditions cannot meet the requirements, and need the cloud storage server to provide cloud data storage services. The auxiliary verification server only needs to receive the data provided by the cloud storage server when a verification request is initiated by the data user terminal, calculate according to the user's rules, and send the calculation result to the user. As a first-upload user, only the ciphertext needs to be uploaded to the cloud storage server, and there is no need to consider other things. As a subsequent uploader, after the tag duplicate verification, it is known that there may be duplicate data, and then the ownership proof of the possible duplicate data needs to be carried out. If the verification passes, it means that the cloud storage server already has the duplicate data and there is no need to upload it again. At this time, only the integrity of the ciphertext stored on the cloud storage server needs to be verified again.
[0070] The duplicate data detection method of the above system will be described in detail below:
[0071] This method includes a system initialization process, a duplicate data detection process, a first file upload process, a subsequent file upload process, and a ciphertext file download and decryption process.
[0072] I. System initialization process
[0073] This process is the first stage of the method described in the present invention and is executed by the cloud storage server. First, the cloud storage server provides a multiplicative cyclic group of prime number p for three initialization stages, which are respectively where g is a generator of the multiplicative cyclic group and the mapping relationship holds; g 1 is a generator of the multiplicative cyclic group
[0074] In addition, the cloud storage server also provides two hash functions h 1 , h 2 :
[0075]
[0076]
[0077] 1 is a special similarity hash (mapping similar data to similar bit strings).
[0078] II. Duplicate data detection process
[0079] This process is jointly completed by the cloud storage server and the data user side, and the result is returned to the current data user side after the process ends.
[0080] This process is divided into the following steps:
[0081] Step 2.1: When the data user side wants to use the cloud storage service to store file F, the data user side first calculates the hash tag corresponding to the plaintext file Then the file tag is sent to the cloud storage server.
[0082] Step 2.2: After receiving the file tag, the cloud storage server performs a matching calculation in the pre-stored file tag set {T}. The matching calculation formula is as shown in formula [1]. Try to see if a similar file tag can be matched.
[0083]
[0084] dis() is a Hamming distance calculation function, which is an existing function; t is a threshold for determining the similarity of file tags at will, and its initial value is set by the user during encryption. In this embodiment, the default value is set to 6.
[0085] If the above formula holds, it indicates that there may be similar ciphertext data in the cloud storage server. The corresponding data user side is the subsequent data upload user side, and it is necessary to perform file ownership verification (the ownership of the similar data file by the current data user side). If the above formula does not hold, it means that no similar file tags are detected, and the current data user is the first data upload user side, and the file can be directly uploaded.
[0086] III. File First Upload Process
[0087] This stage needs to be executed locally by the first data upload user side, including the following steps:
[0088] Step 3.1: Generate error correction information for file tags and block tags: The first data upload user side first calculates the error correction information of the file tag (code_word is a function for calculating the error correction information bits of the error correction code. Just input the information that needs to be error-corrected, and the output result is the codeword). Then, the file F is divided into n equal parts to form {B i}, i = 1, 2,..., n, and calculate and the error correction information corresponding to the file block B i Finally, the data user side randomly selects t for the file block B to generate comparable file block tags i The calculation formula is as follows: (Hide the actual value of the file block tag and implement equality calculation) (Hide the actual value of the file block tag and implement equality calculation)
[0089]
[0090] Step 3.2: Generate a convergence key: First, select a random value ( represents the plaintext file F), and calculate the secret value S, as shown in formula [3], and determine the threshold k of the file block tags required for this file, generate random numbers a 0 、a 1 、…a k-1 , where a 0 = S), and construct a polynomial f(x) = a 0 + a 1 * x + a 2 * x 2 + … + a k-1 * x k-1 , and use the file block codeword information to generate sub-secrets Finally, encrypt the corresponding sub-secret information using the file block codeword information, as shown in Equation [4].
[0091]
[0092]
[0093] Step 3.3: Encrypt the file: The first-time data uploading client first selects a random value K (file key, used to encrypt the plaintext file). K is the inner-layer key for encrypting the plaintext file, and uses this random value to encrypt the file F to obtain the inner-layer ciphertext of the file (Enc() is the encryption algorithm), and calculate the hash value of the ciphertext file And encrypt the ciphertext hash. Select a random value r and calculate the ciphertext of the ciphertext hash The calculation formula is as follows:
[0094]
[0095] Step 3.4: Generate the random inner-layer key ciphertext C corresponding to the convergence key K .
[0096]
[0097] The above formula uses the convergence key K F to blind the file key K.
[0098] Step 3.5: Re-encrypt the file: Select another random value x (outer-layer key, the file key corresponding to the outer-layer ciphertext in double-layer encryption), and use this random value to encrypt the inner-layer ciphertext C F_K to obtain the outer-layer encrypted file C F_K_x .
[0099]
[0100] Step 3.6: Generate the outer-layer random key ciphertext: Select a random value r 1 Calculate the outer-layer random key ciphertext. The calculation formula is as follows:
[0101]
[0102]
[0103] Upload the outer-layer random key ciphertext to the cloud storage server.
[0104] C x is the ciphertext corresponding to the outer-layer key x, and the content inside the curly brackets is all the ciphertexts.
[0105] IV: Subsequent file uploading process
[0106] This process is the interaction process between the subsequent data-uploading client, the cloud storage server, and the auxiliary verification server.
[0107] Step 4.1: Ownership proof: The cloud storage server conducts an ownership proof for the ciphertext files that match the file tags uploaded by the subsequent data-uploading client. The cloud storage server randomly selects from the file block identifiers of similar files Extract the corresponding file block numbers T = {k 1 , …, k κ}, and the corresponding parity bit information Send the ones that need to be and the corresponding parity bit information to the subsequent data-uploading client.
[0108] Step 4.2: Generate ownership evidence: After receiving the file block numbers T and the corresponding parity bit information, the subsequent data-uploading client calculates the block tag corresponding to each number B' i is the similar file block of the subsequent uploader; and fuses and calculates the codeword information from the locally calculated block tags and the parity bit information (correct is the error correction function in the () error correction code, and can correct the errors of the current data within the error correction threshold through the error correction information). Select a random number t' i Generate a comparable block tag Send the comparable tag to the cloud storage server;
[0109]
[0110] Step 4.3: Verify ownership evidence: After receiving the user's ownership evidence, the cloud storage server verifies the evidence generated by the current data client, that is, verifies whether formula
[10] holds. If the equation holds, the file ownership verification passes; otherwise, it fails.
[0111]
[0112] Step 4.4: Distribute partial ciphertext: After passing the file ownership verification, the cloud storage server distributes the file tag error correction information Ciphertext hash ciphertext Random key ciphertext C K 、Random key ciphertext C x (C K is the inner layer key ciphertext, C x is the outer layer key ciphertext) to the data client for re-uploading.
[0113] Step 4.5: Calculate the convergent key: After receiving the ciphertext message, the data client for re-uploading first generates the file tag codeword (The correct function is the error correction function of the error correction code, which is an existing function. The previous article uses the matching error correction function of the error correction code), and then calculates the corresponding sub-secret information The secret recovery process is performed through the decrypted sub-secret information to calculate S, and finally the convergent key is calculated
[0114]
[0115]
[0116]
[0117] Step 4.6: Decrypting the ciphertext: by converging the key Parse the corresponding inner random key information Ciphertext hash information The outer random key information x is calculated as follows:
[0118]
[0119]
[0120] Then, the user end of uploading data again initiates a data integrity verification request to the auxiliary verification server and sends the outer random key information x to the auxiliary verification server.
[0121] Step 4.7: Data integrity challenge: The auxiliary verification server initiates a ciphertext data request to the cloud storage server to obtain the outer ciphertext data C F_K_x Then, the outer random key information x provided by the data user can be used to perform outer decryption to obtain the inner ciphertext And calculate the inner ciphertext file C' F_K Hash value Return the calculated hash value to the user who uploads the data again;
[0122] Step 4.8: Data integrity verification: The user uploads the data again and compares it with the data from the auxiliary verification server. (generated here by the auxiliary verification server after decrypting the outer ciphertext) and the hash value obtained on the cloud storage server (Here is the integrity verification evidence left by the first uploader) to verify whether the ciphertext data is still intact. The comparison formula is
[0123] If the above equation holds, the integrity verification is passed, and the data user selects a random value x' and calculates the re-encryption key and the outer random key ciphertext C x, and upload the re-encryption key rk and the outer-layer random key ciphertext C x to the cloud storage server; C x The calculation formula is as follows:
[0124]
[0125] Step 4.9: Ciphertext re-encryption: After receiving the re-encryption key rk, the cloud storage server performs the re-encryption process and replaces the outer-layer random key ciphertext C uploaded by the data user who uploads the data again x .
[0126] V. Download and decryption process of ciphertext files
[0127] In this process, the data user who owns the ciphertext file downloads and decrypts the ciphertext. In this process, the first uploader will have all the information, while the re-uploader can perform local calculations after proving ownership and only needs to have the key to recover the plaintext. All the following are calculated through and ciphertext information.
[0128] Step 5.1: Ciphertext file download: The data user sends a ciphertext data request to the cloud storage server, and the cloud storage server returns the ciphertext data and its related information to the corresponding data user.
[0129] Step 5.2: File decryption: The data user calculates other information, the inner-layer random key information through the convergence key stored locally ciphertext hash information outer-layer random key information x, removes the outer-layer ciphertext, and the removal formula is:
[0130]
[0131] C F_K_x is the outer-layer ciphertext, C' F_K is the inner-layer ciphertext. By judging whether the equation holds, the integrity of the ciphertext data is verified again. If it holds, it means it is complete. Finally, the inner-layer ciphertext is decrypted to obtain the plaintext file Dec() is the decryption function corresponding to the Enc() encryption function, which is an existing function. The encryption method is not limited, as long as the encryption and decryption functions correspond.
Claims
1. A method for deduplicating similar data ciphertext, characterized in that: The method is used to achieve similar deduplication. The system used by the method includes a cloud storage server, an auxiliary verification server, and a data user terminal; the data user terminal is divided into a first-time data uploading user terminal and a subsequent data uploading user terminal; the method includes the following processes: a system initialization process, a duplicate data detection process, a first file uploading process, and a subsequent file uploading process; During the system initialization process, the cloud storage server generates a multiplication cycle group and two hash functions; In the duplicate data detection process, the cloud storage server detects the file tag sent by the data user and performs matching calculations to confirm whether the current data user is the first data uploader or the subsequent data uploader; In the first file upload process, the first uploading data user generates error correction information and convergence key, sub-secret ciphertext, inner key ciphertext of convergence key for file tag and block tag, encrypts the file twice, generates outer key ciphertext, and uploads it to the cloud storage server; In the subsequent file upload process, the cloud storage server proves the ownership of the ciphertext file that matches the file tag uploaded by the subsequent data upload user terminal. The subsequent data upload user terminal generates ownership evidence and sends it to the cloud storage server for verification. When the verification is passed, the cloud storage server issues part of the ciphertext to the subsequent data upload user terminal, which calculates the convergence key and decrypts the ciphertext; the subsequent data upload user terminal also performs data integrity verification, and if it passes, the cloud storage server will re-encrypt the ciphertext.
2. The method for deduplicating similar data ciphertext according to claim 1, characterized in that: The system initialization process is specifically as follows: the cloud storage server generates and initializes three multiplication cyclic groups of prime number p, which are respectively Among them, g is a multiplicative cyclic group A generator of , and the mapping relationship e: Established; h1 is a multiplicative cyclic group A generator of ; The two hash functions provided by the cloud storage server are:
3. The method for deduplicating similar data ciphertext according to claim 2, characterized in that: The duplicate data detection process specifically includes the following steps: Step 2.1: The data user who wants to upload data calculates the file tag Then send the file tags to the cloud storage server; Step 2.2: After receiving the file tag, the cloud storage server performs a matching calculation in the pre-stored file tag set {T} to try to find a similar file tag. The matching calculation formula is as follows: Among them, the dis() function is the Hamming distance calculation function, and t is the threshold for determining the similarity of file labels; If the above formula is true, it means that similar ciphertext data may exist in the cloud storage server, and the corresponding data user end is the subsequent data uploading user end; if the above formula is not true, it means that no similar file tags are detected, and the current data user is the first data uploading user end.
4. The method for deduplicating similar data ciphertext according to claim 3, characterized in that: The initial value of t is 6.
5. The method for deduplicating similar data ciphertext according to claim 3 or 4, characterized in that: The first upload process includes the following steps: Step 3.1: When uploading data for the first time, the user first calculates the error correction information of the file tag Then divide the file F into n equal parts to form {B i }, i = 1, 2, ..., n, and calculate i=1,2,…,n and file block B i Corresponding error correction information Finally, the data user is file block B i Randomly select t to generate comparable file block labels The formula is as follows: Step 3.2: Choose a random value And calculate the secret value S, determine the file block label threshold k required for the file, and generate random numbers a0, a1, ...a k-1 , where a0 = S, and construct a polynomial f(x) = a0 + a1*x + a2*x 2 +…+a k-1 *x k-1 , the file block codeword information is used to generate the sub-secret Finally, the file block codeword information is used to encrypt the corresponding sub-secret information: Step 3.3: Select a random value K and use it to encrypt file F to obtain the inner ciphertext And calculate the hash value of the ciphertext file And encrypt the ciphertext hash, select a random value r, and calculate the ciphertext of the ciphertext hash; Step 3.4: Generate the inner key ciphertext corresponding to the convergence key; Step 3.5: Select a random value x again and use it to encrypt the ciphertext inner file C F_K Get the outer ciphertext C F_K_x ; Step 3.6: Select a random value r1 to calculate the outer random key ciphertext; Upload the outer random key ciphertext to the cloud storage server.
6. The method for deduplicating similar data ciphertext according to claim 3, characterized in that: The subsequent file upload process includes the following steps: Step 4.1: The cloud storage server performs ownership proof for the ciphertext file whose file tag matches the file uploaded by the subsequent data upload client: The cloud storage server randomly selects the file block identifier of similar files. Extract the corresponding file block number T = {k1,…,k κ }, and the corresponding error correction bit information Send the corresponding error correction bit information to the subsequent data uploading user end; Step 4.2: After the subsequent data upload user receives the file block sequence number and the corresponding error correction bit information, it calculates the block label corresponding to each sequence number. i=k1,k2,…,k κ , where B′ i For file blocks of similar files, the locally calculated block labels and error correction bit information are combined to calculate the codeword information i=k1,k2,…,k κ , choose a random number t′ i Generate comparable block tags Sent to a cloud storage server; Step 4.3: After receiving the user's ownership evidence, the cloud storage server verifies the evidence generated by the current data user, that is, verifies whether the following formula is true: If the above formula is established, the file ownership verification is passed; Step 4.4: After the file ownership is verified, the cloud storage server will correct the file tag information. Ciphertext Hash Ciphertext Inner key ciphertext C K 、Outer key ciphertext C x Issued to subsequent data uploading users; Step 4.5: After receiving the ciphertext message, the user end of subsequent data upload first generates the file label codeword Then calculate the corresponding sub-secret information The secret recovery process is performed through the decrypted sub-secret information to calculate S, and finally the convergent key is calculated Step 4.6: By converging the key Parse the corresponding inner random key information Ciphertext hash information The outer random key information x is calculated as follows: Then, the user end of uploading data again initiates a data integrity verification request to the auxiliary verification server and sends the outer random key information x to the auxiliary verification server; Step 4.7: The auxiliary verification server initiates a ciphertext data request to the cloud storage server to obtain the outer ciphertext data C F_K_x Then, the outer random key information x provided by the data user can be used to perform outer decryption to obtain the inner ciphertext x, and calculate the inner ciphertext file C′ F_K Hash value Return the calculated hash value to the user who uploads the data again; Step 4.8: The subsequent uploaded data client verifies whether the ciphertext data is still intact by comparing the hash values obtained from the auxiliary verification server and the cloud storage server. The comparison formula is: If the above equation holds, the integrity verification is passed, and the data user selects a random value x′ and calculates the re-encryption key and the outer random key ciphertext C x , and the re-encryption key rk and the outer random key ciphertext C x Upload to cloud storage server; C x The calculation formula is as follows: Step 4.9: Ciphertext re-encryption: The cloud storage server performs the re-encryption process after receiving the re-encryption key rk And replace the outer random key ciphertext C of the uploaded data again x .
7. The method for deduplicating similar data ciphertext according to claim 5 or 6, characterized in that: The method also includes a process of downloading and decrypting the ciphertext file, which includes the following steps: Step 5.1: The data user terminal initiates a ciphertext data request to the cloud storage server, and the cloud storage server returns the ciphertext data and related information to the corresponding data user terminal; Step 5.2: The data user uses the converged key Calculate other information, inner random key information Ciphertext hash information Outer random key information x, remove the outer ciphertext, the removal formula is: By judging the equation If it is established, it verifies whether the ciphertext data is complete. If it is established, it proves that it is complete. Finally, the inner ciphertext is decrypted to obtain the plaintext file. The decryption formula is:
Citation Information
Patent Citations
Outsourcing data deduplication cloud storage method supporting privacy and integrity protection
CN110677487A