Data Quality Verification Method Based on Efficient Hashing Algorithm in Cloud Environment
By decomposing data into redundant subfunctions in a cloud environment and merging hash codes using Markov hash trees, combined with the challenge-response protocol, the problem of difficulty in recovering and verifying integrity after data deletion is solved, fast and accurate data recovery and integrity verification are achieved, computing and communication overhead is reduced, and cloud fraud is effectively dealt with.
Patent Information
- Application Number
- CN202111043848.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-09-07
- Publication Date
- 2025-07-18
- Estimated Expiration
- 2041-09-07
AI Technical Summary
It is difficult to recover through cloud after deletion in existing cloud environments, and the verification method is poorly scientific, it is difficult to verify data integrity, it is difficult to effectively deal with cloud fraud, and it is difficult to accurately repair data abnormalities.
The data blocking processing and incremental update method is used to decompose the data into n mutually redundant subfunctions, and the hash code is merged using the Markov hash tree to periodically verify data integrity through the challenge-response protocol, and accurately repair data blocks when abnormalities are found.
It realizes fast and accurate data recovery and integrity verification, reduces computing and communication overhead, effectively deals with cloud fraud, and ensures the scientificity and integrity of data quality verification.
Smart Images

Figure CN113919001B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the processing of cloud data, and is a method for data quality verification based on an efficient hash algorithm in a cloud environment. Background Art
[0002] Cloud technology refers to a hosting technology that unifies a series of resources such as hardware, software, and networks within a wide area network or a local area network to achieve data computing, storage, processing, and sharing. It is a general term for network technology, information technology, integration technology, management platform technology, application technology, etc. based on the cloud computing business model, forming a resource pool, which can be used as needed, flexibly and conveniently. The simplest cloud computing technology can already be seen everywhere in network services, such as search engines, webmail, etc. Users only need to enter simple instructions to obtain a large amount of information. In the future, devices such as mobile phones and GPS will use cloud computing technology for data search and analysis functions through cloud computing technology. In addition, the Provable Data Possession (PDP) mechanism is often used to detect the integrity of big data, while the Proofs of Retrievability (POR) mechanism is often used to detect the integrity of important data. The POR mechanism can not only determine whether the data has been damaged, but also restore the damaged data through error correction codes to ensure the quality of the data. Some existing data quality verifications are based on hash comparison methods, such as the application number 201910242474.X disclosed in the Chinese patent literature, the application publication date is May 28, 2019, and the invention name is "A File Upload Method and System Based on Data Stream and Hash Comparison"; some other cloud data uses integrity detection methods for data analysis, such as the application number 201811264304.3 disclosed in the Chinese patent literature, the application publication date is January 11, 2019, and the invention name is "A Cloud Data Integrity Detection Method and System Based on Blockchain". However, it is difficult to recover the above-mentioned methods and similar methods through the cloud after data deletion, the scientific nature of the verification method is poor, it is difficult to verify the integrity of the data, it is difficult to effectively deal with cloud deception behaviors, and it is difficult to accurately repair data anomalies. Summary of the Invention
[0003] To overcome the above deficiencies, the purpose of the present invention is to provide a method for data quality verification based on an efficient hash algorithm in a cloud environment to the art, mainly solving the technical problems that the data of existing similar methods is difficult to be restored through the cloud after deletion, the scientific nature of the verification method is poor, it is difficult to verify the integrity of the data, it is difficult to effectively deal with cloud deception behaviors, and it is difficult to accurately repair data anomalies. Its purpose is achieved through the following technical solutions.
[0004] A data quality verification method based on an efficient hash algorithm in a cloud environment. Under the guidance of a data recoverable proof mechanism and based on a data quality verification algorithm of a hash function, the method decomposes data into n mutually redundant sub-functions, and the sub-functions are used as data blocks, so that any k data blocks can be safely restored; and the n data blocks are respectively delivered to n clouds; it is characterized in that when the data needs to be updated, a method of using data blocks as granularity, data block calculation and incremental update is adopted to deal with the problem of too much data block evidence when frequently updating data, and then the hash code of the same group of data blocks is merged by using a Markov hash tree; finally, based on the "challenge"-"response" protocol mode, a hash code consistency verification method is adopted to periodically implement data integrity verification and repair data blocks. That is, under the guidance of a data recoverable proof mechanism, the method proposes a data quality verification algorithm based on a hash function for the problem of excessive calculation and communication overhead when verifying data quality in the cloud; at the same time, the hash algorithm based on block calculation and incremental update effectively deals with frequent data updates, excessive evidence update overhead, and reduces the problem of excessive calculation and communication overhead in data integrity verification. In addition, once data corruption is found, any k data blocks can restore the data itself, so this method can effectively protect the integrity status check and data recovery of cloud data. From the perspective of data security, cloud storage servers are not completely trustworthy. Due to technical or cloud storage reasons, there may be inadvertent erasure of stored data or deletion of infrequently used data, and private reduction of the number of backup data, which may lead to behaviors that are harmful to data integrity. In order to ensure the integrity of the data stored in the cloud, users usually use a reliable data recoverability proof mechanism to verify the integrity of the data, thereby ensuring the data integrity of the cloud data.
[0005] The specific steps of the method are as follows: Step 1, the establishment phase, the data is divided into blocks, and a hash function is used to assign a value to each data block, and the Markov hash tree merges the hash codes of the same group of data blocks; Step 2, the challenge phase, the user periodically runs the challenge generation algorithm and generates challenge information, and sends it to the cloud server to start data integrity verification; Step 3, the storage phase, the cloud server performs corresponding calculation operations after receiving the challenge information, and then returns the calculated value to the user. The user runs the verification algorithm based on the returned information to determine whether the data is completely stored in the cloud server.
[0006] The specific steps in the establishment stage of Step 1 are as follows: 1. Perform fixed block processing on the data. This algorithm decomposes the data into n mutually redundant data blocks so that any k data blocks can be used to safely recover the data; obtain n data blocks, and use a hash function to generate a check value for each data block; 2. Transmission, synchronize the data block numbers and the hash values of each data block to the cloud, and deliver the n data blocks to n clouds respectively; 3. Local synchronization. The cloud obtains the data block numbers and hash codes, and combines with data encoding, and uses a Markov hash tree to merge the hash values of the same group of data blocks; 4. Once the data is updated, the user only needs to upload the data blocks with updated data; if the user deletes a certain data block, the user only needs to send the evidence of the deleted data block to the cloud. The cloud saves the evidence of these block deletions and performs the deletion process; if the user has a problem of a large amount of evidence for data blocks during the increase / decrease process, so assume the original data is F, F’ is the updated data, and ΔF is the differential data block of the updated data compared with the original data; 5. Before the user sends the differential data block ΔF, the user will send all the data block hash value information to the cloud. After receiving it, the cloud will use a filter to compare and analyze the data block hash values of the cloud itself to determine whether there are differences and record the difference information; 6. Send a message to the user requesting the user to send the differential data block ΔF. Each time the user updates, only ΔF needs to be uploaded and stored in the (n + 1)-th cloud, that is, a different file storage location from before. The user uses their respective F and ΔF for verification. If the verification passes, use F and ΔF to calculate F’; 7. Recombine the differential data block information and the original data block information to complete incremental synchronization.
[0007] When the data blocks in the establishment stage of Step 1 are merged through hash codes, since usually a group here consists of k data blocks, if any k data can be used to safely recover the data, the k data blocks are considered as the same group, and the local records the hash values of each group of data blocks.
[0008] In Step 2, since the data is decomposed into n mutually redundant data blocks so that any k data blocks can be used to safely recover the data, therefore, randomly select any k data blocks to calculate the challenge value c, and then send the challenge value c to the cloud server through a detection command to start data integrity verification.
[0009] The specific steps in the saving stage of Step 3 are as follows:
[0010] 1. After receiving the request, the cloud server uses a Markov hash tree to combine the hash values of any k data blocks among the n data blocks and generates multiple groups of hash values of data blocks. Assume that m groups of hash values r are obtained;
[0011] 2. Return m r values to the client. The client performs a numerical search based on the m r values to verify the data integrity. If c = r is found, the data is considered complete; otherwise, the data is incomplete.
[0012] If there is incomplete data, the client randomly generates a new challenge value c and sends it to the cloud again. After receiving the request, the cloud uses the Markov hash tree to combine the hash values of any k data blocks out of the n data blocks and generates multiple sets of hash values for the data blocks. Suppose m sets of hash values r are obtained and the m r values are returned to the client. The client performs a numerical search based on the m r values. If c = r cannot be found, continue to execute Step 3 until the traversal is complete. After the traversal is complete, if c = r cannot be found, it means the data cannot be restored from the existing data blocks in the cloud. On the contrary, if c = r can be found during several traversals, it means that although some data blocks are damaged, the data can still be restored. In this case, the client notifies the cloud of the data blocks for which the verification is performed this time and updates the abnormal data blocks to accurately locate the error position and repair the data blocks.
[0013] The data recovery of the present invention is accurate, fast, and relatively complete, and can effectively cope with the behavior of cloud deception and accurately repair data anomalies. It is suitable for data quality verification, recovery, and maintenance in a cloud environment and technical improvements of similar methods. BRIEF DESCRIPTION OF THE DRAWINGS
[0014] Figure 1 is a block diagram of the working principle of the Markov hash tree of the present invention for merging the same set of data blocks. EMBODIMENT
[0015] Now, in combination with the drawings, the specific implementation steps of the present invention will be further described in detail.
[0016] The specific steps of the method are as follows:
[0017] Step 1, Establishment stage: Perform block processing on the data, assign a value to each data block using a hash function, and the Markov hash tree merges the hash codes of the same set of data blocks. The specific steps of the establishment stage are as follows:
[0018] 1. Perform fixed block processing on the data. This algorithm decomposes the data into n mutually redundant data blocks so that any k data blocks can safely restore the data. Obtain n data blocks and generate a check value for each data block using a hash function.
[0019] 2. Transmission: Synchronize the data block numbers and the hash values of each data block to the cloud, and hand over the n data blocks to n clouds respectively.
[0020] 3. Locally synchronize the data block numbers and hash codes obtained from the cloud. Combine with data encoding and use Markov hash tree to merge the hash values of the same set of data blocks. Here, the same set usually consists of k data blocks. If any k data can safely restore the data, then these k data blocks are considered as the same set. Locally record the hash value of each set of data blocks. The specific method is as follows: As Figure 1 shown Figure 1 in, Markov hash tree is used to merge the hash codes of the same set of data blocks.
[0021] 4. Once the data is updated, the user only needs to upload the data blocks with updated data. Thus, the user only needs to generate a new hash code for the newly added data blocks, reducing the communication overhead of the system. A proof update technology combining block calculation and incremental update is proposed to achieve the update of data blocks. Therefore, it is divided into two types: deleting and updating evidence. Taking deleting data as an example, if the user deletes a certain data block, the user only needs to send the evidence of the deleted data block to the cloud. The cloud saves the evidence of deleting these blocks and performs the deletion process. Taking updating as an example, when the user adds / deletes data, it will cause the problem of a large amount of evidence for data blocks. Therefore, assume the original data is F, the updated data is F’, and ΔF is the differential data blocks between the updated data and the original data.
[0022] 5. Before sending the differential data blocks ΔF, the user will send all the data block hash value information to the cloud. After receiving it, the cloud will use a filter to compare and analyze the data block hash values on the cloud itself to determine whether there are differences and record the difference information.
[0023] 6. Send a message to the user to request the user to send the differential data blocks ΔF. Each time the user updates, only ΔF needs to be uploaded and stored on the (n + 1)-th cloud, that is, different from the previous file storage location. The user uses their respective F and ΔF for verification. If the verification passes, F and ΔF are used to calculate F’.
[0024] 7. Recombine the differential data block information and the original data block information to complete incremental synchronization.
[0025] Step 2: Challenge stage. The user periodically runs the challenge generation algorithm and generates challenge information, and sends it to the cloud server to start data integrity verification. Since the data is decomposed into n mutually redundant data blocks, any k data blocks can safely restore the data. Therefore, randomly select any k data blocks to calculate the challenge value c, and then send the challenge value c to the cloud server through a detection command to start data integrity verification.
[0026] Step 3: Preservation stage. After receiving the challenge information, the cloud server performs corresponding calculation operations and then returns the calculated value to the user. The user runs a verification algorithm based on the returned information to determine whether the data is completely preserved in the cloud server. The specific steps of the preservation stage are as follows:
[0027] 1. After receiving the request, the cloud server randomly combines the hash values of k out of n data blocks using a Markov hash tree and generates multiple sets of hash values for the data blocks. Suppose m sets of hash values r are obtained.
[0028] 2. The cloud server returns the m r values to the client. The client performs a numerical search based on the m r values to verify the integrity of the data. If c = r is found, the data is considered to be complete; otherwise, the data is incomplete.
[0029] If the data is incomplete, it is generally due to damage to the cloud data. The client randomly generates a new challenge value c and sends it to the cloud again. After receiving the request, the cloud server randomly combines the hash values of k out of n data blocks using a Markov hash tree and generates multiple sets of hash values for the data blocks. Suppose m sets of hash values r are obtained, and the cloud server returns the m r values to the client. The client performs a numerical search based on the m r values. If c = r cannot be found, Step 3 is continued until the traversal is complete. After the traversal is complete, if c = r cannot be found, it means that the data cannot be restored using the existing data blocks in the cloud. On the contrary, if c = r can be found during several traversals, it means that although some data blocks are damaged, the data can still be restored. In this case, the client notifies the cloud of the data blocks to be verified this time and updates the abnormal data blocks to accurately locate the error position and repair the data blocks.
[0030] The innovations of this method are as follows: 1. The data is updated by using data blocks and incremental updates. This method can flexibly handle dynamic data quality verification; taking data blocks as the granularity, the data is calculated and updated incrementally in blocks to address the problem of excessive data block evidence in frequently updated data, effectively reducing the communication overhead of the system. 2. The Markov hash tree is used to calculate the combined hash values of any k data. It can effectively deal with cloud deception because the cloud does not know which combination method the user uses to verify the data integrity. Calculate the challenge value c using any k data blocks. Among the m r values returned to the user when there is no missing cloud data block, there must be a matching one; on the contrary, if there are missing or damaged data blocks, the client reports the numbers of any k randomly selected data blocks for accurate update and repair of abnormal cloud data blocks. Since this method periodically verifies the data quality, once an abnormal data block is found, it repairs the abnormal data block without scanning and repairing all data blocks, greatly reducing the computational complexity of data quality verification and effectively protecting the integrity status check and data recovery of cloud data. 3. After receiving the request, the server uses the Markov hash tree to combine the hash values of any k data blocks among n data blocks and generates multiple sets of hash values of data blocks. Suppose m sets of hash values r are obtained.
Claims
1. A data quality verification method based on an efficient hash algorithm in a cloud environment. Under the guidance of a data recoverability proof mechanism and based on a data quality verification algorithm of a hash function, the data is decomposed into n mutually redundant sub-functions. The sub-functions are used as data blocks, so that any k data blocks can safely recover the data. And the n data blocks are respectively sent to n clouds. It is characterized in that When the data needs to be updated, a method of taking data blocks as granularity, calculating data in blocks and performing incremental updates is adopted to solve the problem of excessive data block evidence when frequently updating data. Then, a Markov hash tree is used to merge the hash codes of the same group of data blocks. Finally, based on the "challenge"-"response" protocol mode, a verification method of hash code consistency is used to periodically verify data integrity and repair data blocks. The specific steps of this method are as follows: Step 1, establishment phase: The data is block-processed, and a hash function is used to assign values to each data block. The Markov hash tree is used to merge the hash codes of the same group of data blocks. Step 2, challenge phase: The user periodically runs the challenge generation algorithm to generate challenge information and sends it to the cloud server to start data integrity verification. Step 3, storage phase: After receiving the challenge information, the cloud server performs corresponding calculation operations and then returns the calculated value to the user. The user runs the verification algorithm based on the returned information to determine whether the data is completely stored in the cloud server. The steps in the establishment phase of Step 1 are specifically as follows:
1. Perform fixed block processing on the data. This algorithm decomposes the data into n mutually redundant data blocks so that any k data blocks can safely restore the data. After obtaining n data blocks, a hash function is used to generate check values for each data block.
2. Transmission: Synchronize the data block numbers and the hash values of each data block to the cloud, and hand over the n data blocks to n clouds respectively.
3. Local synchronization: The cloud obtains the data block numbers and hash codes, and combines with the data encoding. The Markov hash tree is used to merge the hash values of the same group of data blocks.
4. Once the data is updated, the user only needs to upload the data blocks with updated data. If the user deletes a certain data block, the user only needs to send the evidence of the deleted data block to the cloud. The cloud saves the evidence of deleting these blocks and performs deletion processing. Assume the original data is F, F’ is the updated data, and ΔF is the differential data block of the updated data compared with the original data.
5. Before sending the differential data block ΔF, the user sends all the data block hash value information to the cloud. After receiving it, the cloud uses a filter to compare and analyze the data block hash values on the cloud itself to determine whether there are differences and records the difference information.
6. Send a message to the user to request the user to send the differential data block ΔF. Each time the user updates, only ΔF needs to be uploaded and stored in the (n + 1)-th cloud, that is, a different file storage location from before. The user uses their respective F and ΔF for verification. If the verification passes, F and ΔF are used to calculate F’.
7. Recombine the differential data block information and the original data block information to complete incremental synchronization. When the data blocks in the establishment phase of Step 1 are merged through hash codes, since the same group usually consists of k data blocks here, if any k data can safely restore the data, the k data blocks are considered the same group, and the hash values of each group of data blocks are locally recorded. In the second step, since the data is decomposed into n mutually redundant data blocks, any k data blocks can be used to safely recover the data. Therefore, any k data blocks are randomly selected to calculate the challenge value c, and then the challenge value c is sent to the cloud server through a detection command to start data integrity verification; The specific steps in the saving stage of the third step are as follows:
1. After receiving the request, the cloud server uses a Markov hash tree to combine the hash values of any k data blocks among the n data blocks and generates multiple sets of hash values of the data blocks. Suppose m sets of hash values r are obtained; 2. And the m r values are returned to the client, and the client performs a numerical search based on the m r values to verify the data integrity; If c = r is found, the data is considered to be complete; otherwise, the data is incomplete.
2. The data quality verification method based on an efficient hashing algorithm in a cloud environment according to claim 1, characterized in that In the third step, if there is an incomplete data situation, the client randomly generates a new challenge value c and then sends it to the cloud again; After receiving the request, the cloud uses a Markov hash tree to combine the hash values of any k data blocks among the n data blocks and generates multiple sets of hash values of the data blocks; Suppose m sets of hash values r are obtained, and the m r values are returned to the client. The client performs a numerical search based on the m r values. If the situation of c = r cannot be found, step three is continued until the traversal is completed; After the traversal is completed, if the situation of c = r cannot be found, it means that the data cannot be recovered through the existing data blocks in the cloud; On the contrary, if the situation of c = r can be found during several traversals, it means that although the data blocks are damaged, the data can still be recovered; In this case, the client notifies the cloud of the data blocks to be verified this time and updates the abnormal data blocks to accurately locate the error position and repair the data blocks.
Citation Information
Patent Citations
Cloud data integrity detection method and system based on block chain
CN109194466A
file uploading method and system based on data flow and Hash comparison
CN109819056A
Method of two-dimensional control and data integrity assurance
RU2696425C1
Lattice-based cloud storage data security audit method supporting uploading of data via proxy
WO2018201730A1