Cloud storage data integrity auditing method

By adopting a dynamic hash tree structure in the cloud storage system, the hash tree structure is dynamically adjusted according to the access situation of data blocks, the performance loss problem of traditional hash trees in skewed data access scenarios is solved, and efficient data integrity audit and system performance improvement is achieved.

CN119996430AActive Publication Date: 2025-05-13XIAN THERMAL POWER RES INST CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510113121.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-24
Publication Date
2025-05-13
Estimated Expiration
2045-01-24

AI Technical Summary

Technical Problem

In the large-scale data scenario of cloud storage, the traditional hash tree structure will result in performance losses when facing skewed data access scenarios, and cannot efficiently ensure data integrity.

Method used

The dynamic hash tree structure is adopted, and the hash value of the data block is used as the leaf node. The hash tree structure is dynamically adjusted according to the access or update of the data block, so that the parent nodes of the access or updated data block are used as the root nodes, while keeping the precedent sequence and node attributes unchanged.

Benefits of technology

By dynamically adjusting the hash tree structure, the performance waste of data verification is reduced, the efficiency and reliability of the cloud storage system are improved, and it is suitable for common access imbalance scenarios in cloud storage scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119996430A_ABST
    Figure CN119996430A_ABST
Patent Text Reader

Abstract

The embodiment of the invention relates to the technical field of network security, provides a cloud storage data integrity auditing method, designs an extensible dynamic hash tree structure based on an extensible tree thought in the field of load balancing, can be suitable for common access imbalance scenes in cloud storage scenes, and improves the auditing efficiency. The dynamic Hash tree structure is beneficial to reducing performance waste caused by data verification, and on the basis of the dynamic Hash tree structure, a data integrity auditing scheme is provided by utilizing path proof, so that the cloud storage system has high efficiency and reliability at the same time.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of network security technology, and in particular to a cloud storage data integrity audit method. Background Art

[0002] With the continuous development of the Internet, people's digital needs are increasing, cloud storage scenarios are becoming more and more popular, and network security attacks on cloud services are becoming increasingly serious, which makes people pay more and more attention to the data security issues of cloud storage. For the integrity of data storage, some special scenarios can use hardware to enhance the data integrity protection mechanism, while general scenarios mostly use integrity verification mechanisms such as hash tree structures such as Merkle trees. However, in large-scale data scenarios of cloud storage, the introduction of traditional hash tree structures may cause serious performance losses. This is because traditional hash tree structures such as Merkle trees usually assume that data access scenarios are average, that is, the frequency of access to each data node is randomized and approximate, but in reality, data access in most cloud storage services is tilted, that is, a small number of data blocks may be accessed frequently. If the frequently accessed data blocks are far from the root of the hash tree, integrity verification and data acquisition will cause a lot of performance waste. Therefore, in the face of tilted access scenarios in cloud service scenarios, how to build an efficient data storage structure and propose a dynamic audit mechanism is an urgent problem to be solved. Summary of the invention

[0003] The present disclosure aims to solve at least one of the problems existing in the prior art and provide a cloud storage data integrity audit method.

[0004] In one aspect of the present disclosure, a cloud storage data integrity audit method is provided, the method comprising:

[0005] The client divides the file to be stored into multiple data blocks, constructs a corresponding dynamic hash tree structure with the hash value of each data block as a leaf node and stores it locally, and sends the multiple data blocks and their corresponding dynamic hash tree structures to the cloud storage server;

[0006] The cloud storage server receives and stores the plurality of data blocks and the corresponding dynamic hash tree structures;

[0007] The client and the cloud storage server dynamically adjust the dynamic hash tree structures stored in each of them according to the access or update status of each of the data blocks, so that the parent node corresponding to the accessed or updated data block is used as the root node of the dynamic hash tree structure, while keeping the pre-order sequence of the dynamic hash tree structure unchanged and the node attributes of each node unchanged; wherein the pre-order sequence only includes the leaf nodes of the dynamic hash tree structure;

[0008] In response to the data integrity audit requirement, the cloud storage server sends the latest flag of the root node of the dynamic hash tree structure and the corresponding critical path proof to the client; the critical path proof is used to indicate that the hash value of the data block to be verified is the hash value of the brother node of the leaf node;

[0009] The client performs integrity audit verification on the data block to be verified based on the latest mark of the root node and the corresponding critical path proof sent by the cloud storage server.

[0010] Optionally, the client and the cloud storage server dynamically adjust the dynamic hash tree structures stored in each according to the access or update status of each data block, so that the parent node corresponding to the accessed or updated data block is used as the root node of the dynamic hash tree structure, while keeping the pre-order sequence of the dynamic hash tree structure unchanged and the node attributes of each node unchanged, including:

[0011] After each data access or data update, if the leaf node where the hash value of the accessed or updated data block is located belongs to the non-root child node of the dynamic hash tree structure stored by the client and the cloud storage server respectively, the client and the cloud storage server respectively adjust the parent node of the leaf node where the hash value of the accessed or updated data block is located to the root node of the dynamic hash tree stored by each of them.

[0012] Optionally, the client and the cloud storage server respectively adjust the parent node of the leaf node where the hash value of the accessed or updated data block is located to the root node of the dynamic hash tree stored by each client, including:

[0013] The client and the cloud storage server respectively use operations similar to the stretching operation of the splay tree to adjust the parent node of the leaf node where the hash value of the accessed or updated data block is located to the root node of the dynamic hash tree stored in each of them.

[0014] Optionally, the client and the cloud storage server dynamically adjust the dynamic hash tree structure stored in each of them after each data access and / or each data update, while avoiding the period of data integrity audit verification and data processing period.

[0015] Optionally, the responding to the existence of a data integrity audit requirement includes:

[0016] After accessing the data block stored in the cloud storage server or updating the data block stored in the cloud storage server, the client actively sends an integrity audit request for the data block to be verified to the cloud storage server.

[0017] Optionally, the responding to the existence of a data integrity audit requirement includes:

[0018] After the client accesses the data block stored in the cloud storage server or updates the data block stored in the cloud storage server, the cloud storage server actively initiates an integrity audit request for the data block to be verified.

[0019] Optionally, the client performs integrity audit verification on the data block to be verified based on the latest mark of the root node and the corresponding critical path proof sent by the cloud storage server, including:

[0020] The client performs hash calculation on the critical path proof sent by the cloud storage server and the hash value of the data block to be verified, obtains the root node calculation mark corresponding to the dynamic hash tree in the cloud storage server, and uses the root node calculation mark to verify the latest root node mark sent by the cloud storage server.

[0021] Optionally, the client performs integrity audit verification on the data block to be verified based on the latest mark of the root node and the corresponding critical path proof sent by the cloud storage server, including:

[0022] After each data update, the client uses the latest mark of the root node and the corresponding critical path proof of the dynamic hash tree structure stored locally to verify the latest mark of the root node and the corresponding critical path proof sent by the cloud storage server.

[0023] Optionally, before the client and the cloud storage server dynamically adjust the dynamic hash tree structures stored in each according to the access or update status of each data block, the method further includes:

[0024] The cloud storage server sends the pre-order sequence and the root node initial mark of the dynamic hash tree structure stored by it to the client.

[0025] Another aspect of the present disclosure provides a cloud storage data integrity audit system, the system comprising a client and a cloud storage server, wherein:

[0026] The client is used to divide the file to be stored into multiple data blocks, construct a corresponding dynamic hash tree structure with the hash value of each data block as a leaf node and store it locally, and send the multiple data blocks and their corresponding dynamic hash tree structures to the cloud storage server;

[0027] The cloud storage server is used to receive and store the plurality of data blocks and the corresponding dynamic hash tree structures;

[0028] The client and the cloud storage server are further used to dynamically adjust the dynamic hash tree structures stored in each according to the access or update status of each data block, so that the parent node corresponding to the accessed or updated data block is used as the root node of the dynamic hash tree structure, while keeping the pre-order sequence of the dynamic hash tree structure unchanged and the node attributes of each node unchanged; wherein the pre-order sequence only includes the leaf nodes of the dynamic hash tree structure;

[0029] The cloud storage server is also used to send the latest mark of the root node of the dynamic hash tree structure and the corresponding critical path proof to the client when there is a need for data integrity audit; the critical path proof is used to indicate that the hash value of the data block to be verified is the hash value of the brother node of the leaf node;

[0030] The client is also used to perform integrity audit verification on the data block to be verified based on the latest mark of the root node and the corresponding critical path proof sent by the cloud storage server.

[0031] Compared with the prior art, the present disclosure has the following beneficial effects: a scalable dynamic hash tree structure is designed, which can be applied to the common access imbalance scenario in cloud storage scenarios, and the structure helps to reduce the performance waste caused by data verification; based on the dynamic hash tree structure, a data integrity audit scheme is proposed using path proof, so that the cloud storage system has both high efficiency and reliability. BRIEF DESCRIPTION OF THE DRAWINGS

[0032] One or more embodiments are exemplarily described by pictures in the corresponding drawings, and these exemplifications do not constitute limitations on the embodiments. Elements with the same reference numerals in the drawings represent similar elements, and unless otherwise stated, the figures in the drawings do not constitute proportional limitations.

[0033] Figure 1 A flowchart of a cloud storage data integrity audit method provided by one embodiment of the present disclosure;

[0034] Figure 2 A schematic diagram of a cloud storage application scenario provided by another embodiment of the present disclosure;

[0035] Figure 3 A schematic diagram of an adjustment process of a dynamic hash tree provided in another embodiment of the present disclosure;

[0036] Figure 4 A schematic diagram of the application process of a cloud storage data integrity audit method provided by another embodiment of the present disclosure. DETAILED DESCRIPTION

[0037] In order to make the purpose, technical scheme and advantages of the embodiments of the present disclosure clearer, the embodiments of the present disclosure will be described in detail below in conjunction with the accompanying drawings. However, it can be understood by those skilled in the art that in each embodiment of the present disclosure, many technical details are proposed in order to enable readers to better understand the present disclosure. However, even without these technical details and various changes and modifications based on the following embodiments, the technical scheme claimed for protection in the present disclosure can also be implemented. The division of the following embodiments is for the convenience of description and should not constitute any limitation on the specific implementation of the present disclosure. The various embodiments can be combined and referenced with each other without contradiction.

[0038] In cloud storage scenarios, users tend to access specific data, which makes some of the cloud storage data accessed frequently, while the rest of the data is accessed infrequently, which leads to the phenomenon of data access skewness. Users often need to verify the integrity of data after storing data and before the next data access. Existing data integrity verification methods usually use balanced binary hash trees such as Merkle trees, while traditional hash trees usually use a fixed structure, and the data access path is from the root node of the hash tree to a certain intermediate node or leaf node. In the face of the skewed data access phenomenon in cloud storage scenarios, if the nodes accessed frequently are far away from the root node, and the nodes accessed infrequently are close to the root node, each data access query will have a large performance overhead. Moreover, if the security scenario of data integrity verification is further considered, the path from the root node to the node far away from the root node will increase the performance burden of data integrity verification.

[0039] In order to solve the problem caused by the inclination of data access in the above cloud storage scenario, an embodiment of the present disclosure provides a cloud storage data integrity audit method, the process of which is as follows: Figure 1 As shown, it includes steps S110 to S150.

[0040] The cloud storage data integrity audit method provided in this embodiment can be applied to Figure 2 The cloud storage scenario shown in Figure 1. Figure 2 As shown, in this scenario, the client may include applications such as a browser, a database management system (DBMS), a web service, and a mail service, and may use a file system to store the relevant data files of these applications. The client may also divide the relevant data files of each application into several data blocks and store them in the data block layer, use the cloud storage data integrity audit method provided in this embodiment to construct a dynamic hash tree structure of the files to be stored in the cloud storage server, store the dynamic hash tree in the cloud storage server, and use the dynamic hash tree structure to verify the integrity of the stored files.

[0041] In step S110, the client divides the file to be stored into multiple data blocks, constructs a corresponding dynamic hash tree structure for the leaf node with the hash value of each data block and stores it locally, and sends the multiple data blocks and their corresponding dynamic hash tree structures to the cloud storage server.

[0042] Specifically, the file to be stored refers to the file in the client that needs to be stored in the cloud storage server. For example, the client can slice the data of the file to be stored F, divide the file to be stored F into data blocks B0-Bn and store them, where n represents the maximum number of data blocks, and the total number of data blocks included in the file to be stored F is n+1. The hash values ​​of data blocks B0-Bn can be expressed as H(B0)-H(Bn), respectively, H represents the hash function, and the client can use H(B0)-H(Bn) as leaf nodes to construct the corresponding hash tree as a dynamic hash tree structure, and send the data blocks B0-Bn and the corresponding hash tree structure to the cloud storage server. In order to distinguish the dynamic hash tree structures stored by the client and the cloud storage server respectively, the dynamic hash tree structure in the client can be recorded as Tree_L, and the dynamic hash tree structure in the cloud storage server can be recorded as Tree_G.

[0043] It should be noted that the dynamic hash tree structure in this embodiment is a structure similar to a stretch tree. The stretch tree is mainly used in scenarios such as garbage collection mechanism and Internet Protocol (IP) routing query. Since the stretch tree has the characteristic of real-time stretching adjustment according to the load, this embodiment constructs a dynamic hash tree structure based on the stretch tree, and implements data query and integrity security verification in cloud storage scenarios based on the dynamic hash tree structure.

[0044] Step S120: The cloud storage server receives and stores multiple data blocks and their corresponding dynamic hash tree structures.

[0045] Specifically, when the client sends the data blocks B0-Bn and the corresponding hash tree structure Tree_G to the cloud storage server, the cloud storage server receives and stores the data blocks B0-Bn and the hash tree structure Tree_G.

[0046] In step S130, the client and the cloud storage server dynamically adjust the dynamic hash tree structures stored in each according to the access or update status of each data block, so that the parent node corresponding to the accessed or updated data block is used as the root node of the dynamic hash tree structure, while keeping the pre-order sequence of the dynamic hash tree structure unchanged and the node attributes of each node unchanged. The pre-order sequence only includes the leaf nodes of the dynamic hash tree structure.

[0047] It should be noted that the dynamic adjustment of the dynamic hash tree structure stored by the client and the cloud storage server occurs after each data access and / or each data update, while avoiding the period of data integrity audit verification and data processing.

[0048] Specifically, during the data integrity audit verification period, data consistency is usually required, so the dynamic hash tree structure should not be adjusted during this period to avoid deviations in data integrity audit verification. During the data processing period, such as when background storage tasks require privacy-oriented data processing, the data structure tree also needs to be strongly consistent, so the dynamic hash tree structure should not be adjusted at this time.

[0049] That is to say, the client and the cloud storage server adjust the dynamic hash tree structure stored by each other, which can be after each data access or after each data update, but the adjustment process of the dynamic hash tree structure needs to avoid the period of data integrity audit verification and the data processing period to maintain the strong consistency of the data structure tree. To this end, an adjustment position bit f can be set, and the adjustment position bit f is used to indicate whether the dynamic hash tree structure is allowed to be adjusted. For example, the adjustment position bit f can take a value of 0 or 1. Among them, the adjustment position bit f takes a value of 0, which can indicate that the dynamic hash tree structure is not allowed to be adjusted at this time, and correspondingly, the adjustment position bit f takes a value of 1, which indicates that the dynamic hash tree structure is allowed to be adjusted at this time. Alternatively, the adjustment position bit f takes a value of 0, which can also indicate that the dynamic hash tree structure is allowed to be adjusted at this time, and correspondingly, the adjustment position bit f takes a value of 1, which indicates that the dynamic hash tree structure is not allowed to be adjusted at this time.

[0050] In the process of adjusting the dynamic hash tree structure, the node attributes of each node remain unchanged, which means that if a node is a leaf node before the adjustment, it will still be a leaf node after the adjustment. Correspondingly, if a node is a non-leaf node before the adjustment, it will still be a non-leaf node after the adjustment. By keeping the pre-order sequence unchanged and the node attributes of each node unchanged during the adjustment of the dynamic hash tree structure, it is helpful for subsequent data integrity audit verification.

[0051] By adjusting the dynamic hash tree structure, the nodes that are frequently accessed can be adjusted to child nodes of the root node to shorten the data access path and improve data access efficiency.

[0052] Exemplarily, in step S130, the client and the cloud storage server dynamically adjust the dynamic hash tree structures they store respectively according to the access or update status of each data block, so that the parent node corresponding to the accessed or updated data block serves as the root node of the dynamic hash tree structure, while keeping the pre-order sequence of the dynamic hash tree structure unchanged and the node attributes of each node unchanged, including: after each data access or data update, if the leaf node where the hash value of the accessed or updated data block is located belongs to the non-root child node of the dynamic hash tree structure stored by the client and the cloud storage server respectively, then the client and the cloud storage server respectively adjust the parent node of the leaf node where the hash value of the accessed or updated data block is located to the root node of the dynamic hash tree stored by each.

[0053] Specifically, if the leaf node where the hash value of the accessed or updated data block is located is already a child node of the root node of the dynamic hash tree structure, then the dynamic hash tree structure at this time can already reflect that the leaf node is a frequently accessed node. At this time, there is no need to adjust the dynamic hash tree structure.

[0054] Exemplarily, the client and the cloud storage server respectively adjust the parent node of the leaf node where the hash value of the accessed or updated data block is located to the root node of the dynamic hash tree stored by each of them, including: the client and the cloud storage server respectively adopt an operation similar to the stretch operation of the stretch tree to adjust the parent node of the leaf node where the hash value of the accessed or updated data block is located to the root node of the dynamic hash tree stored by each of them.

[0055] Specifically, a stretch tree is a self-adjusting binary search tree that maintains the balance of the tree through a series of processes called stretch operations. Among them, the stretch operation refers to moving a node to the root of the tree through a series of rotation operations after accessing a node (whether it is search, insert or delete), so as to effectively shorten the access path and help redistribute the nodes in the tree. The present embodiment makes improvements on this basis, and uses the stretch operation to adjust the parent node of the visited or updated leaf node to the root node, so as to reduce the performance overhead caused by the adjustment process to a certain extent.

[0056] For example, combining Figure 3 Assuming that the leaf node where the hash value of the accessed or updated data block is located is B, its parent node is T, and T's parent node is P, then according to the specific position of T in the dynamic hash tree structure, the adjustment process of the dynamic hash tree structure can be divided into different situations.

[0057] For example, Figure 3As shown in (1), when the dynamic hash tree structure before adjustment includes B, T, P and leaf nodes L1 and L2, and B and T are both left child nodes, only one right rotation is required to turn T into a root node. Since the level of B is increased by one level during the adjustment of the dynamic hash tree structure, the level fq of B is also changed from 0 to 1, that is, from fq=0 to fq=1. Figure 3 The adjustment process of the dynamic hash tree structure shown in (1) belongs to the case of single adjustment.

[0058] For example, Figure 3 As shown in (2), when the dynamic hash tree structure before adjustment includes B, T, P, leaf nodes L1, L2, L3 and root node N1, and B, T, and P are all left child nodes, two right rotations are required, that is, first rotating P to the right and then rotating T to the right to turn T into a root node. Since each right rotation during the adjustment of the dynamic hash tree structure raises the level of B by one level, the level fq of B also changes from 0 to 1 and then to 2, that is, from fq=0 to fq=1 and then to fq=2. Figure 3 The adjustment process of the dynamic hash tree structure shown in (2) belongs to the case of multiple one-way adjustments.

[0059] For example, Figure 3 As shown in (3), when the dynamic hash tree structure before adjustment includes B, T, P, leaf nodes L1, L2, L3 and root node N1, B and T are right child nodes, and P is a left child node, two rotations are required, that is, first rotating T to the left and then rotating P to the right to turn T into a root node. Since the left rotation raises the level of B by one level during the adjustment of the dynamic hash tree structure, while the right rotation does not change the level of B, the level fq of B only changes from 0 to 1, that is, from fq=0 to fq=1. Figure 3 The adjustment process of the dynamic hash tree structure shown in (3) belongs to the case of multiple bidirectional adjustments.

[0060] Exemplarily, before the client and the cloud storage server dynamically adjust the dynamic hash tree structures stored by each according to the access or update status of each data block, the cloud storage data integrity audit method also includes: the cloud storage server sends the pre-sequence and root node initial mark of the dynamic hash tree structure stored by it to the client, so that before the client dynamically adjusts the dynamic hash tree structure, it ensures that the dynamic hash tree structure stored by itself is consistent with the dynamic hash tree structure stored by the cloud storage server according to the pre-sequence and root node initial mark. The root node initial mark here refers to the location identifier of the root node of the dynamic hash tree structure sent by the client and received by the cloud storage server.

[0061] Step S140: In response to the need for data integrity audit, the cloud storage server sends the latest flag of the root node of the dynamic hash tree structure and the corresponding critical path certificate to the client. The critical path certificate is used to indicate that the hash value of the data block to be verified is the hash value of the brother node of the leaf node.

[0062] Specifically, the data integrity audit requirement can be proposed by the client or by the cloud storage server.

[0063] When a data integrity audit requirement is raised by the client, in step S140, in response to the existence of a data integrity audit requirement, the client actively sends an integrity audit request for the data block to be verified to the cloud storage server after accessing the data block stored on the cloud storage server or updating the data block stored on the cloud storage server.

[0064] When a data integrity audit requirement is raised by the cloud storage server, in step S140, in response to the existence of the data integrity audit requirement, the cloud storage server proactively initiates an integrity audit request for the data block to be verified after the client accesses the data block stored in the cloud storage server or updates the data block stored in the cloud storage server.

[0065] Step S150: The client performs integrity audit verification on the data block to be verified based on the latest mark of the root node and the corresponding critical path proof sent by the cloud storage server.

[0066] Specifically, when the data integrity audit requirement is raised by the client, the client actively performs data integrity audit verification. When the data integrity audit requirement is raised by the cloud storage server, the client passively performs data integrity audit verification.

[0067] The data integrity audit verification performed by the client includes regular verification and strict verification. Among them, regular verification is used for each data integrity audit verification. Strict verification is used for data integrity audit verification after each data update.

[0068] Exemplarily, during routine verification, step S150 includes: the client performs hash calculation on the key path proof sent by the cloud storage server and the hash value of the data block to be verified, obtains the root node calculation mark corresponding to the dynamic hash tree in the cloud storage server, and uses the root node calculation mark to verify the latest root node mark sent by the cloud storage server.

[0069] Specifically, the root node calculation mark here refers to the current position mark of the root node of the dynamic hash tree in the cloud storage server calculated by the client based on the key path proof sent by the cloud storage server and the hash value of the data block to be verified. The root node latest mark sent by the cloud storage server refers to the current position mark of the root node of the dynamic hash tree in the cloud storage server. If the root node calculation mark is consistent with the root node latest mark sent by the cloud storage server, the cloud storage data is considered complete, otherwise the cloud storage data is considered incomplete.

[0070] Exemplarily, during strict verification, step S150 includes: after each data update, the client uses the latest root node mark and the corresponding critical path proof of the locally stored dynamic hash tree structure to verify the latest root node mark and the corresponding critical path proof sent by the cloud storage server.

[0071] Specifically, in order to prevent the insecurity and privacy risks of the cloud storage server, during strict verification, the client uses its locally stored latest root node mark and the corresponding critical path proof to verify the latest root node mark and the corresponding critical path proof sent by the cloud storage server. If the latest root node mark stored locally by the client is consistent with the latest root node mark sent by the cloud storage server, and the critical path proof stored locally by the client is consistent with the critical path proof sent by the cloud storage server, the cloud storage data is considered complete, otherwise the cloud storage data is considered incomplete.

[0072] In order to enable those skilled in the art to better understand the above implementation, a specific example is provided below for illustration.

[0073] like Figure 4 As shown, a cloud storage data integrity audit method is applied to a cloud storage scenario and includes steps ① to ⑨.

[0074] Step ①: The client cuts the file F into data blocks B0~Bn, and builds a dynamic hash tree based on the data blocks B0~Bn, recorded as Tree_L, and stores the data blocks B0~Bn and the dynamic hash tree Tree_L locally.

[0075] Step ②: The client sends the data blocks B0-Bn and the dynamic hash tree Tree_G which is the same as the dynamic hash tree Tree_L to the server, which can be a cloud storage server. The server receives the data blocks B0-Bn and the dynamic hash tree Tree_G, and sends the initial root mark r of the root node of the dynamic hash tree Tree_G and the hash pre-sequence L to the client. Assume that the client divides the file F into four data blocks, namely B0, B1, B2, and B3. The hash values ​​of B0, B1, B2, and B3 are H(B0), H(B1), H(B2), and H(B3), respectively. H(B0), H(B1), H(B2), and H(B3) correspond to the leaf nodes B, L1, L2, and L3 of the dynamic hash tree Tree_G, respectively. The parent nodes of B and L1 are both T, the parent nodes of T and L2 are both P, and the parent nodes of P and L3 are both root nodes N1. Then the initial root flag r of the dynamic hash tree Tree_G can be expressed as r=H(H(H(H(B0, H(B1)), H(B2)), H(B3)), and the hash pre-order sequence L can be expressed as L={H(B0), H(B1), H(B2), H(B3)}.

[0076] Step 3: The client and the server perform multiple data access, update and other interactions.

[0077] Step ④: Assuming that the last data interaction between the client and the server was to update the data block B0 corresponding to the leaf node B, the server adjusts the dynamic hash tree Tree_G according to the interaction situation, adjusts the parent node of the leaf node B to the root node, and the client synchronously adjusts the dynamic hash tree Tree_L. After the adjustment is completed, the new root mark of the dynamic hash tree Tree_G is r', and r' is expressed as r'=H(H(B0),H(H(B1),H(H(B2),(B3)))).

[0078] Step ⑤: After completing the update of data block B0, the client sends a data block B0 integrity audit request Rv(H(B0)) to the server.

[0079] Step ⑥: The server sends the new root mark r'=H(H(B0),H(H(B1),H(H(B2),(B3)))) and the key path proof proof=H(P)=H(H(B1),H(H(B2),H(B3))) of the dynamic hash tree Tree_G to the client.

[0080] Step 7: The client uses the critical path proof proof = H(P) = H(H(B1), H(H(B2), H(B3))) and H(B0) to perform hash calculation and obtain the root node calculation mark r^ = H(H(B0), proof).

[0081] Step ⑧: The client performs a routine integrity verification to verify whether r^ is equal to r', that is, r^? = r'. If r^ is equal to r', the data block B0 stored on the server is considered complete, otherwise the data block B0 stored on the server is considered incomplete.

[0082] Step ⑨: The client performs strict integrity verification and uses the locally stored dynamic hash tree Tree_L to verify r' and proof. If the current root flag of the dynamic hash tree Tree_L is consistent with r' and the key path proof obtained based on the dynamic hash tree Tree_L is consistent with the proof, the data block B0 stored on the server is considered complete, otherwise the data block B0 stored on the server is considered incomplete.

[0083] The cloud storage data integrity audit method provided by the embodiments of the present disclosure has the following beneficial effects compared with the prior art: a scalable dynamic hash tree structure is designed, which can be applied to the common access imbalance scenarios in cloud storage scenarios, and the structure helps to reduce the performance waste caused by data verification; based on the dynamic hash tree structure, a data integrity audit scheme is proposed using path proof, so that the cloud storage system has both high efficiency and reliability.

[0084] Another embodiment of the present disclosure relates to a cloud storage data integrity audit system, including a client and a cloud storage server.

[0085] The client is used to divide the file to be stored into multiple data blocks, build a corresponding dynamic hash tree structure for the leaf node with the hash value of each data block and store it locally, and send the multiple data blocks and their corresponding dynamic hash tree structure to the cloud storage server.

[0086] The cloud storage server is used to receive and store multiple data blocks and their corresponding dynamic hash tree structures.

[0087] The client and the cloud storage server are also used to dynamically adjust the dynamic hash tree structures stored in each of them according to the access or update status of each data block, so that the parent node corresponding to the accessed or updated data block serves as the root node of the dynamic hash tree structure, while keeping the pre-order sequence of the dynamic hash tree structure unchanged and the node attributes of each node unchanged; wherein the pre-order sequence only includes the leaf nodes of the dynamic hash tree structure.

[0088] The cloud storage server is also used to send the latest flag of the root node of the dynamic hash tree structure and the corresponding key path proof to the client when there is a need for data integrity audit. The key path proof is used to indicate that the hash value of the data block to be verified is the hash value of the brother node of the leaf node.

[0089] The client is also used to perform integrity audit verification on the data block to be verified based on the latest root node flag and corresponding critical path proof sent by the cloud storage server.

[0090] The specific implementation method of the cloud storage data integrity audit system provided by the embodiment of the present disclosure can be found in the cloud storage data integrity audit method provided by the embodiment of the present disclosure, and will not be repeated here.

[0091] The cloud storage data integrity audit system provided by the embodiments of the present disclosure has the following beneficial effects compared with the prior art: a scalable dynamic hash tree structure is designed, which can be applied to the common access imbalance scenarios in cloud storage scenarios, and this structure helps to reduce the performance waste caused by data verification; based on the dynamic hash tree structure, path proof is used to implement data integrity audit, so that the cloud storage system has both high efficiency and reliability.

[0092] Those skilled in the art will appreciate that the above-mentioned embodiments are specific embodiments for implementing the present disclosure, and in actual applications, various changes may be made thereto in form and detail without departing from the spirit and scope of the present disclosure.

Claims

1. A cloud storage data integrity audit method, characterized in that: The method comprises: The client divides the file to be stored into multiple data blocks, constructs a corresponding dynamic hash tree structure with the hash value of each data block as a leaf node and stores it locally, and sends the multiple data blocks and their corresponding dynamic hash tree structures to the cloud storage server; The cloud storage server receives and stores the plurality of data blocks and the corresponding dynamic hash tree structures; The client and the cloud storage server dynamically adjust the dynamic hash tree structures stored in each of them according to the access or update status of each of the data blocks, so that the parent node corresponding to the accessed or updated data block is used as the root node of the dynamic hash tree structure, while keeping the pre-order sequence of the dynamic hash tree structure unchanged and the node attributes of each node unchanged; wherein the pre-order sequence only includes the leaf nodes of the dynamic hash tree structure; In response to the data integrity audit requirement, the cloud storage server sends the latest flag of the root node of the dynamic hash tree structure and the corresponding critical path proof to the client; the critical path proof is used to indicate that the hash value of the data block to be verified is the hash value of the brother node of the leaf node; The client performs integrity audit verification on the data block to be verified based on the latest mark of the root node and the corresponding critical path proof sent by the cloud storage server.

2. The method according to claim 1, characterized in that The client and the cloud storage server dynamically adjust the dynamic hash tree structures stored in each according to the access or update status of each data block, so that the parent node corresponding to the accessed or updated data block is used as the root node of the dynamic hash tree structure, while keeping the pre-order sequence of the dynamic hash tree structure unchanged and the node attributes of each node unchanged, including: After each data access or data update, if the leaf node where the hash value of the accessed or updated data block is located belongs to the non-root child node of the dynamic hash tree structure stored by the client and the cloud storage server respectively, the client and the cloud storage server respectively adjust the parent node of the leaf node where the hash value of the accessed or updated data block is located to the root node of the dynamic hash tree stored by each of them.

3. The method according to claim 2, characterized in that The client and the cloud storage server respectively adjust the parent node of the leaf node where the hash value of the accessed or updated data block is located to the root node of the dynamic hash tree stored by each of them, including: The client and the cloud storage server respectively use operations similar to the stretching operation of the splay tree to adjust the parent node of the leaf node where the hash value of the accessed or updated data block is located to the root node of the dynamic hash tree stored in each of them.

4. The method according to any one of claims 1 to 3, characterized in that: The client and the cloud storage server dynamically adjust the dynamic hash tree structure stored in each of them after each data access and / or each data update, while avoiding the period of data integrity audit verification and data processing period.

5. The method according to claim 1, characterized in that The response to the existence of a data integrity audit requirement includes: After accessing the data block stored in the cloud storage server or updating the data block stored in the cloud storage server, the client actively sends an integrity audit request for the data block to be verified to the cloud storage server.

6. The method according to claim 1, characterized in that The response to the existence of a data integrity audit requirement includes: After the client accesses the data block stored in the cloud storage server or updates the data block stored in the cloud storage server, the cloud storage server actively initiates an integrity audit request for the data block to be verified.

7. The method according to claim 1, characterized in that The client performs integrity audit verification on the data block to be verified based on the latest root node flag and the corresponding critical path proof sent by the cloud storage server, including: The client performs hash calculation on the critical path proof sent by the cloud storage server and the hash value of the data block to be verified, obtains the root node calculation mark corresponding to the dynamic hash tree in the cloud storage server, and uses the root node calculation mark to verify the latest root node mark sent by the cloud storage server.

8. The method according to claim 1, characterized in that The client performs integrity audit verification on the data block to be verified based on the latest root node flag and the corresponding critical path proof sent by the cloud storage server, including: After each data update, the client uses the latest mark of the root node and the corresponding critical path proof of the dynamic hash tree structure stored locally to verify the latest mark of the root node and the corresponding critical path proof sent by the cloud storage server.

9. The method according to claim 1, characterized in that: Before the client and the cloud storage server dynamically adjust the dynamic hash tree structures stored in each according to the access or update status of each data block, the method further includes: The cloud storage server sends the pre-order sequence and the root node initial mark of the dynamic hash tree structure stored by it to the client.

10. A cloud storage data integrity audit system, characterized in that: The system includes a client and a cloud storage server; The client is used to divide the file to be stored into multiple data blocks, construct a corresponding dynamic hash tree structure with the hash value of each data block as a leaf node and store it locally, and send the multiple data blocks and their corresponding dynamic hash tree structures to the cloud storage server; The cloud storage server is used to receive and store the plurality of data blocks and the corresponding dynamic hash tree structures; The client and the cloud storage server are further used to dynamically adjust the dynamic hash tree structures stored in each according to the access or update status of each data block, so that the parent node corresponding to the accessed or updated data block is used as the root node of the dynamic hash tree structure, while keeping the pre-order sequence of the dynamic hash tree structure unchanged and the node attributes of each node unchanged; wherein the pre-order sequence only includes the leaf nodes of the dynamic hash tree structure; The cloud storage server is also used to send the latest mark of the root node of the dynamic hash tree structure and the corresponding critical path proof to the client when there is a need for data integrity audit; the critical path proof is used to indicate that the hash value of the data block to be verified is the hash value of the brother node of the leaf node; The client is also used to perform integrity audit verification on the data block to be verified based on the latest mark of the root node and the corresponding critical path proof sent by the cloud storage server.

Citation Information

Patent Citations

  • Data authentication method and system

    CN115328875A

  • Integrity auditing for multi-copy storage

    WO2021007863A1