Data processing method, electronic equipment, storage medium and program
By generating local Merkle trees on data processing nodes and building global Merkle trees, the problem of Merkle tree calculation taking time and difficulty in meeting dynamic update requirements is solved, and efficient verification and dynamic update of data sharding is achieved.
Patent Information
- Application Number
- CN202510222376.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-27
- Publication Date
- 2025-05-30
AI Technical Summary
When processing data sharding, the calculation of Merkle tree takes a long time, the structure is complex, it is difficult to meet the dynamic update requirements, and the verification process is lengthy.
Generate a global Merkle tree by generating a local Merkle tree at the data processing node and sending the local root to the central coordination point. This method improves the verification efficiency and dynamic update capability of data sharding through the construction of distributed multi-level Merkle trees.
It solves the problem that Merkle tree computing takes time and is difficult to meet the needs of dynamic updates, improves the verification efficiency of data sharding, and can meet the needs of dynamic updates of data sharding.
Smart Images

Figure CN120067122A_ABST
Abstract
Description
Technical Field
[0001] The embodiments of the present invention relate to the technical field of data processing, and in particular, to a data processing method, apparatus, electronic device, storage medium, and program. Background Art
[0002] Data sharding is the technical foundation of distributed storage. It divides data into multiple parts, and each part is stored on different nodes. Data sharding is often applied in technical scenarios such as distributed systems, large-scale data storage and processing, cloud computing, database management, and ETL (Extract, Transform, Load).
[0003] A Merkle Tree (also known as a hash tree) is a binary tree structure that is widely used in data verification and integrity checking. The application of Merkle Tree in data sharding mainly ensures the integrity, traceability, and cross-shard cooperation efficiency of data sharding through its efficient hash verification mechanism.
[0004] In the process of implementing the present invention, the inventors found that the prior art has the following defects: Since the data sharding scenario usually involves a relatively large amount of data, the calculation of the Merkle Tree takes a relatively long time and the structure is relatively complex, thereby reducing the verification efficiency of data sharding. When some data shards are updated, it is necessary to traverse and calculate the entire tree structure of the Merkle Tree again, which is difficult to meet the dynamic update requirements of data sharding. At the same time, when verifying the data integrity and consistency of data shards by using a Merkle Tree with a complex structure through a hash calculation and verification mechanism, there will also be a problem of a long verification process. Summary of the Invention
[0005] The embodiments of the present invention provide a data processing method, apparatus, electronic device, storage medium, and program, which can meet the dynamic update requirements of data sharding and improve the verification efficiency of data sharding.
[0006] According to one aspect of the present invention, there is provided a data processing method, which is applied to a central coordination node and includes:
[0007] Obtain the local roots of local Merkle Trees sent by each data processing node; wherein, the local Merkle Trees are generated by the data processing nodes according to local data shards;
[0008] Use the local roots of the local Merkle Trees as the bottom-layer leaf nodes of a global Merkle Tree to generate the global Merkle Tree.
[0009] According to another aspect of the present invention, there is provided a data processing apparatus, which is configured in a central coordination node and includes:
[0010] A local root acquisition module, configured to acquire local roots of local Merkle trees sent by each data processing node; wherein, the local Merkle trees are generated by the data processing nodes according to local data shards;
[0011] A global Merkle tree generation module, configured to use the local roots of the local Merkle trees as the bottom-level leaf nodes of a global Merkle tree to generate the global Merkle tree.
[0012] According to another aspect of the present invention, there is provided an electronic device, which includes:
[0013] At least one processor; and
[0014] A memory communicatively connected to the at least one processor; wherein,
[0015] The memory stores a computer program executable by the at least one processor, and when the computer program is executed by the at least one processor, the at least one processor can execute the data processing method according to any embodiment of the present invention.
[0016] According to another aspect of the present invention, there is provided a computer-readable storage medium storing computer instructions, and when the computer instructions are used to be executed by a processor, the data processing method according to any embodiment of the present invention is implemented.
[0017] According to another aspect of the present invention, there is also provided a computer program product, including a computer program, and when the computer program is executed by a processor, the data processing method according to any embodiment of the present invention is implemented.
[0018] In the embodiments of the present invention, after each data processing node generates a local Merkle tree according to local data shards, each data processing node sends the local root of the generated local Merkle tree to the central coordination node. After the central coordination node acquires the local roots of the local Merkle trees sent by each data processing node, it can use the local roots of the local Merkle trees as the bottom-level leaf nodes of the global Merkle tree to generate the global Merkle tree, which can solve the problems of long calculation time, difficulty in meeting dynamic update requirements, and low verification efficiency existing in the prior art when processing data shards through Merkle trees, can meet the dynamic update requirements of data shards, and improve the efficiency of data shard verification.
[0019] It should be understood that the content described in this part is not intended to identify the key or important features of the embodiments of the present invention, nor is it used to limit the scope of the present invention. Other features of the present invention will become easily understood through the following description. Description of the Drawings
[0020] To more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the accompanying drawings required for the description of the embodiments. Obviously, the accompanying drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other accompanying drawings can be obtained based on these drawings.
[0021] Figure 1 is a flowchart of a data processing method provided in the first embodiment of the present invention;
[0022] Figure 2 is a flowchart of a data processing method provided in the second embodiment of the present invention;
[0023] Figure 3 is a schematic structural diagram of a Merkle tree provided in the second embodiment of the present invention;
[0024] Figure 4 is a schematic diagram of a data processing device provided in the third embodiment of the present invention;
[0025] Figure 5 is a schematic structural diagram of an electronic device provided in the fourth embodiment of the present invention. Detailed Embodiments
[0026] In order to enable those skilled in the art to better understand the solutions of the present invention, the following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only some of the embodiments of the present invention, rather than all of them. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0027] It should be noted that the terms "first", "second", etc. in the specification and claims of the present invention and the above accompanying drawings are used to distinguish similar objects, and do not necessarily need to describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances so that the embodiments of the present invention described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device that includes a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or are inherent to these processes, methods, products, or devices.
[0028] Embodiment 1
[0029] Figure 1 FIG. 0 is a flowchart of a data processing method provided in the first embodiment of the present invention. This embodiment is applicable to the situation of managing data shards according to a local Merkle tree and a global Merkle tree. This method can be executed by a data processing device, which can be implemented in a software and / or hardware manner and is generally integrated in an electronic device. The electronic device can be a terminal device or a server device, as long as it can execute the data processing method. The specific device type of the electronic device is not limited in the embodiments of the present invention. Correspondingly, as Figure 1 shown, the method includes the following operations:
[0030] S110. Obtain the local roots of the local Merkle trees sent by each data processing node; wherein, the local Merkle tree is generated by the data processing node according to the local data shards.
[0031] Among them, the data processing node can be a device node for storing data shards. The local Merkle tree can be a Merkle tree generated by the data processing node according to the local data shards. The local root can be the root of the local Merkle tree. The local data shards are the data shards stored locally by the data processing node.
[0032] First, the technical solution of the embodiment of the present invention is mainly applicable to such an application scenario: a central coordination node and each data processing node form a data processing system (or a data management system, hereinafter referred to as the system). The central coordination node first shards a large amount of raw data to obtain a plurality of data shards, and sends each data shard to the corresponding data processing node for storage. Among them, the type of the data shard corresponding to the raw data can be any data type, for example, it can include but is not limited to log data, transaction data, and background data of application programs, etc. The specific type and data content of the data shard corresponding to the raw data are not limited in the embodiments of the present invention.
[0033] Optionally, the data processing node can store one or more data shards. The number of data shards stored by the data processing node is not limited in the embodiments of the present invention. A data processing node can generate a local Merkle tree according to the local data shards, and each data processing node can send the root of the generated local Merkle tree as the local root to the central coordination node, and the central coordination node generates a global Merkle tree according to the local roots of each local Merkle tree. Optionally, the central coordination node can be one of the data processing nodes selected from each data processing node, or can also be a node with central control function. The node type of the central coordination node is not limited in the embodiments of the present invention.
[0034] S120. Use the local roots of the respective local Merkle trees as the bottom - layer leaf nodes of the global Merkle tree to generate the global Merkle tree.
[0035] It can be understood that the local roots of the local Merkle trees sent by each data - processing node are also hash values. After the central coordination node obtains the local roots of the local Merkle trees sent by each data - processing node, it can use the local roots of the respective local Merkle trees as the bottom - layer leaf nodes of the global Merkle tree, arrange the bottom - layer leaf nodes of the global Merkle tree, and calculate the hash values of the intermediate nodes layer by layer in the order from bottom to top of the bottom - layer leaf nodes until the root value of the global Merkle tree is generated, thus obtaining the global Merkle tree.
[0036] In the prior art, usually, the hash values calculated by sharding all the data stored in a data - processing node directly are used as the bottom - layer leaf nodes of the Merkle tree, and then the Merkle tree is generated. When the capacity of the original data is large, the data calculation takes a long time. The hash calculation needs to process more data, which will lead to an extended calculation time, thus increasing the generation time and verification time of the Merkle tree. When some data shards are updated, the entire tree structure of the Merkle tree needs to be recalculated.
[0037] In the embodiments of the present invention, at the data - processing node, a local Merkle tree corresponding to a data - processing node is first calculated for the data shards. The bottom - layer leaf nodes of the local Merkle tree can be one or a small number of data shards. The magnitude of the data shards is significantly reduced compared to all the data shards. When calculating the local Merkle tree with it as the bottom - layer leaf node, the time consumption is less, and the generation efficiency and verification efficiency are higher. Then, a distributed multi - level Merkle tree is formed by the local Merkle tree and the global Merkle tree, which can improve the generation efficiency of the global Merkle tree. When some data shards are updated, only the corresponding local Merkle tree needs to be recalculated, and then the global Merkle tree is updated dynamically and quickly.
[0038] In the embodiments of the present invention, after each data - processing node generates a local Merkle tree according to the local data shards, each data - processing node sends the local root of the generated local Merkle tree to the central coordination node. After the central coordination node obtains the local roots of the local Merkle trees sent by each data - processing node, it can use the local roots of the respective local Merkle trees as the bottom - layer leaf nodes of the global Merkle tree to generate the global Merkle tree, which can solve the problems existing in processing data shards by the Merkle tree in the prior art, such as long calculation time, difficulty in meeting the dynamic update requirements, and low verification efficiency, and can meet the dynamic update requirements of data shards and improve the efficiency of data - shard verification.
[0039] Embodiment 2
[0040] Figure 2 FIG. is a flowchart of a data processing method provided in Embodiment 2 of the present invention. This embodiment is a specific implementation based on the above embodiment. In this embodiment, various specific and optional implementation manners for generating a global Merkle tree are given. Correspondingly, as Figure 2 shown, the method of this embodiment may include:
[0041] S210. Obtain the local roots of the local Merkle trees sent by each data processing node; wherein, the local Merkle tree is generated by the data processing node according to the local data shard.
[0042] S220. Determine, according to the target reference factor, the target node arrangement rule for each of the local roots as the underlying leaf nodes of the global Merkle tree.
[0043] Among them, the target reference factor may be a factor for arranging the local roots at the positions of the underlying leaf nodes. The target node arrangement rule is a specific rule for arranging the local roots at the positions of the underlying leaf nodes.
[0044] S230. Arrange each of the local roots in sequence at the node positions of the underlying leaf nodes of the global Merkle tree according to the target node arrangement rule, obtain the underlying leaf nodes of the global Merkle tree, and generate the global Merkle tree.
[0045] When using the local roots of each local Merkle tree as the underlying leaf nodes of the global Merkle tree, some factors affecting data verification may be screened as the target reference factor according to the law of later data verification. The rule for arranging the local roots is determined through the determined target reference factor as the target node arrangement rule, and each local root is arranged in sequence at the node positions of the underlying leaf nodes of the global Merkle tree, and then the underlying leaf nodes of the global Merkle tree can be obtained.
[0046] Optionally, the target reference factor may include but is not limited to the generation time of the local root, the geographical location of the data processing node where the local root is located, or the data query frequency of the data processing node where the local root is located, etc. For the underlying leaf nodes of the global Merkle tree arranged according to the target node arrangement rule determined by the target reference factor, the correlation between adjacent nodes is relatively strong.
[0047] It can be understood that for a Merkle tree, the stronger the correlation between adjacent nodes, the higher the efficiency of data verification using the Merkle tree, and the higher the update efficiency of the Merkle tree. Figure 3It is a schematic structural diagram of a Merkle tree provided in the second embodiment of the present invention. In a specific example, as Figure 3 shown, assume that the global Merkle tree includes 4 bottom-level leaf nodes, namely node 1, node 2, node 3, and node 4. Every two bottom-level leaf nodes can be calculated to generate an intermediate node, obtaining 2 intermediate nodes of node 5 and node 6. The 2 intermediate nodes can be calculated to generate the root node of node 7. Figure 3 All nodes of the Merkle tree shown are hash values. If the two adjacent bottom-level leaf nodes, node 1 and node 2, have a strong correlation, when performing data verification or Merkle tree update operations, if it is necessary to verify or update some data in the data slices corresponding to node 1 and node 2, then only the hash values of node 1 and node 2 need to be recalculated, and the hash values of node 3 and node 4 do not need to be recalculated repeatedly, and the new root can be quickly calculated layer by layer; if the two non-adjacent bottom-level leaf nodes, node 1 and node 3, have a strong correlation, when performing data verification or Merkle tree update operations, if it is necessary to verify or update some data in the data slices corresponding to node 1 and node 3, then it is necessary to recalculate the hash values of all bottom-level leaf nodes, and calculate the root of the Merkle tree step by step according to the recalculated hash values of the bottom-level leaf nodes, and the efficiency is relatively low.
[0048] In an optional embodiment of the present invention, the target reference factor may include the generation time of the local root; the arranging each of the local roots in sequence at the node positions of the bottom-level leaf nodes of the global Merkle tree according to the target node arrangement rule to obtain the bottom-level leaf nodes of the global Merkle tree may include: determining the generation time of each of the local roots; determining the node positions of each of the local roots at the bottom-level leaf nodes of the global Merkle tree according to the chronological order of the generation times of each of the local roots; arranging each of the local roots in sequence at the node positions of the bottom-level leaf nodes of the global Merkle tree to obtain the bottom-level leaf nodes of the global Merkle tree.
[0049] Optionally, when the target reference factor includes the generation time of local roots, the time for each data processing node to generate the local root of the local Merkle tree can be determined. Generally, the closer the generation times of two local roots are, it indicates that the generation times of the data shards stored by the data processing nodes corresponding to these two local roots are closer, and further indicates that the correlation between the data shards stored by the data processing nodes corresponding to these two local roots is stronger. Therefore, when determining the node positions of each local root at the bottom-level leaf nodes of the global Merkle tree in the order of the generation times of each local root, local roots with stronger data correlation can be configured at adjacent bottom-level leaf nodes. In this way, further arranging each local root in sequence according to the node positions of each local root at the bottom-level leaf nodes of the global Merkle tree, the correlation between adjacent bottom-level leaf nodes in the bottom-level leaf nodes of the global Merkle tree is stronger.
[0050] In an optional embodiment of the present invention, the target reference factor includes the node geographical location of the data processing node; the arranging each of the local roots in sequence at the node positions of the bottom-level leaf nodes of the global Merkle tree according to the target node arrangement rule to obtain the bottom-level leaf nodes of the global Merkle tree may include: determining the node geographical location of each data processing node where the local root is located; determining the node positions of each local root at the bottom-level leaf nodes of the global Merkle tree according to the proximity order or node position weight of the corresponding node geographical locations of each local root; arranging each of the local roots in sequence according to the node positions of each local root at the bottom-level leaf nodes of the global Merkle tree to obtain the bottom-level leaf nodes of the global Merkle tree.
[0051] Among them, the node geographical location may be the geographical location where the data processing node is located. The node position weight may be a weight value assigned to each data processing node, which can be used to represent the importance degree of the geographical location where the data processing node is located.
[0052] Optionally, when the target reference factor includes the node geographical location of the data processing node, the node geographical locations of the data processing nodes where each local root is located can be determined. Generally, the closer the node geographical locations of the data processing nodes where two local roots are located, the stronger the correlation of the data shards stored by the data processing nodes corresponding to the two local roots. For example, the data shards stored by data processing nodes with close node geographical locations can be data shards from the same user or the same equipment manufacturer. Therefore, when determining the node positions of each local root at the bottom leaf nodes of the global Merkle tree in the order of proximity of the corresponding node geographical locations of each local root, local roots with stronger data correlation can be configured at adjacent bottom leaf nodes. In this way, further arranging each local root in sequence according to the node positions of each local root at the bottom leaf nodes of the global Merkle tree, the correlation between adjacent bottom leaf nodes in the bottom leaf nodes of the global Merkle tree is relatively strong.
[0053] It can be understood that some data processing nodes can also mark the node position weights for their corresponding node geographical locations according to the importance level. The larger the node position weight, the higher the importance level of the data processing node, and the higher the importance level of the data shards stored by it. For example, the node position weights corresponding to each data processing node arranged in City A are 0.5, and the node position weights corresponding to each data processing node arranged in City B are 0.2. Then, the data shards stored by each data processing node arranged in City A are generally more important than the data shards stored by each data processing node arranged in City B. At the same time, the closer the node position weights of the data processing nodes are, the stronger the correlation of the data shards they store. For example, when the node position weight corresponding to data processing node 1 is 0.5, the node position weight corresponding to data processing node 2 is 0.5, and the node position weight corresponding to data processing node 3 is 0.2, it indicates that the data shards stored by data processing node 1 and data processing node 2 are usually data shards of the same importance level. Therefore, according to the magnitude relationship of the node position weights of the corresponding node geographical locations of each local root, local roots with similar node position weights can be arranged as close as possible at adjacent node positions of the bottom leaf nodes of the global Merkle tree. In this way, further arranging each local root in sequence according to the node positions of each local root at the bottom leaf nodes of the global Merkle tree, the correlation between adjacent bottom leaf nodes in the bottom leaf nodes of the global Merkle tree is relatively strong.
[0054] In an alternative embodiment of the present invention, the target reference factor includes the data query frequency of the data processing node; the arranging each of the local roots in sequence at the node positions of the bottom-layer leaf nodes of the global Merkle tree according to the target node arrangement rule to obtain the bottom-layer leaf nodes of the global Merkle tree may include: determining the data query frequency of the data processing node where each of the local roots is located or the data query frequency corresponding to the combination of data processing nodes where the combination of local roots is located; determining the node positions of each of the local roots at the bottom-layer leaf nodes of the global Merkle tree according to the magnitude order of the data query frequencies corresponding to each of the local roots or the combination of local roots; and arranging each of the local roots in sequence according to the node positions of each of the local roots at the bottom-layer leaf nodes of the global Merkle tree to obtain the bottom-layer leaf nodes of the global Merkle tree.
[0055] The data query frequency may be the frequency of query verification of the stored data shards by the data processing node. The combination of local roots may be a combination composed of multiple local roots. The combination of data processing nodes may be a combination composed of multiple data processing nodes.
[0056] Optionally, when the target reference factor includes the data query frequency of the data processing node, the data query frequency of the data processing node where each local root is located or the data query frequency corresponding to the data processing node combination where the local root combination is located can be determined. Generally, the higher the data query frequency of a data processing node or data processing node combination, the higher the verification requirement of this type of data processing node for the locally stored data. The closer the data query frequencies of different data processing nodes are, the closer the data query verification requirements of this type of data processing node are. Exemplarily, when the data query frequency of data processing node 1 is 100 times per month, the data query frequency of data processing node 2 is 10 times per month, and the data query frequency of data processing node 3 is 110 times per month, data processing node 1 and data processing node 3 can be used as a group of adjacent bottom-level leaf nodes that can generate hash values. Correspondingly, when data processing node 1 and data processing node 3 perform data query verification later, in the global Merkle tree, only the hash values of data processing node 1 and data processing node 3 need to be recalculated, and the hash values of other nodes that have not been query-verified do not need to be recalculated repeatedly, and the new root can be quickly calculated layer by layer. Thus, the closer the data query frequencies of the data processing nodes where the local roots are located, the closer the probabilities of query-verifying the data shards later. Therefore, when determining the node positions of each local root in the bottom-level leaf nodes of the global Merkle tree according to the magnitude order of the data query frequencies corresponding to the data processing nodes where the local roots are located, the local roots with stronger data query frequency correlation can be configured at adjacent bottom-level leaf nodes. In this way, further arranging each local root in sequence according to the node positions of each local root in the bottom-level leaf nodes of the global Merkle tree, in the bottom-level leaf nodes of the global Merkle tree, the correlation of data query operations between adjacent bottom-level leaf nodes is stronger, and the calculation amount for data query verification based on this global Merkle tree is less and the efficiency is higher.
[0057] It can be understood that in some cases, users need to perform data query verification on different data processing nodes simultaneously. For example, at the same moment, the user needs to query and verify some data shards stored in data processing node 1 and data processing node 3. Therefore, when arranging the node positions of the local roots at the bottom leaf nodes of the global Merkle tree, the data query frequency corresponding to the data processing node combination where the local root combination is located can also be referred to, and the node positions of the local root combinations at the bottom leaf nodes of the global Merkle tree are determined in the order of the magnitudes of the data query frequencies corresponding to the data processing node combinations where the local root combinations are located. Exemplarily, the data query frequency for data query verification performed simultaneously by data processing node 1 and data processing node 3 is 90 times per month, the data query frequency for data query verification performed simultaneously by data processing node 2 and data processing node 4 is 80 times per month, and the data query frequency for data query verification performed simultaneously by data processing node 5 and data processing node 6 is 20 times per month. Then, the node positions of each local root at the bottom leaf nodes of the global Merkle tree can be: node 1 - node 3 - node 2 - node 4 - node 5 - node 6. Among them, the two local roots corresponding to node 1 and node 3 can generate a hash value as an intermediate node of the global Merkle tree; the two local roots corresponding to node 2 and node 4 can generate a hash value as an intermediate node of the global Merkle tree; the two local roots corresponding to node 5 and node 6 can generate a hash value as an intermediate node of the global Merkle tree. In this way, in the order of the magnitudes of the data query frequencies corresponding to the local root combinations, each local root combination is arranged in sequence at the node positions of the bottom leaf nodes of the global Merkle tree. In the bottom leaf nodes of the global Merkle tree, the correlation between data query operations between adjacent bottom leaf nodes is relatively strong. Based on this global Merkle tree, the computational amount for data query verification is less and the efficiency is higher.
[0058] In an alternative embodiment of the present invention, after generating the global Merkle tree, it may further include: obtaining a digital signature of the global root of the global Merkle tree; storing the digital signature of the global root, as well as the index information and path information of each data shard in the local Merkle tree and the global Merkle tree.
[0059] After obtaining the global root of the global Merkle tree, the central coordination node can submit the global root of the global Merkle tree to the CA (Certificate Authority). The CA uses its private key for digital signature, which can only sign 256-bit data, significantly reducing the signature calculation pressure. At the same time, the central coordination node can also pre-record the index information and path information of each data shard in the multi-level Merkle tree of the local Merkle tree and the global Merkle tree, so as to quickly extract the optimal verification path from the multi-level Merkle tree during data query verification, and complete the data integrity verification with the least number of hash value calculations.
[0060] In an optional embodiment of the present invention, after generating the global Merkle tree, it may further include: obtaining the updated local root of the updated Merkle tree generated by the shard update data processing node for the updated data shard; using the updated local root to replace the target bottom leaf node of the shard update data processing node in the global Merkle tree to generate an updated global Merkle tree.
[0061] Among them, the shard update data processing node may be a data processing node where the local stored data shard has an update operation. The updated data shard is also the data shard updated by the shard update data processing node. The updated Merkle tree may be a local Merkle tree regenerated by the shard update data processing node according to the updated local stored data shard. The updated local root is also the root of the updated Merkle tree. The target bottom leaf node may be the bottom leaf node corresponding to the local root of the shard update data processing node in the global Merkle tree. The updated global Merkle tree may be a global Merkle tree regenerated according to the updated local root.
[0062] Correspondingly, due to the adoption of the form of a distributed multi-level Merkle tree. Therefore, when some data processing nodes perform data update operations, it is not necessary to recalculate the hash values of the intermediate nodes for updating for all the bottom leaf nodes of the global Merkle tree. Specifically, it is possible to only obtain the updated local root of the updated Merkle tree generated by the shard update data processing node for the updated data shard, and use the updated local root to replace the original local root of the target bottom leaf node corresponding to the shard update data processing node in the global Merkle tree, and dynamically update the path affected by the updated local root to generate an updated global Merkle tree.
[0063] In a specific example, such as Figure 3As shown in the figure, assume that there is an update operation on the data shard of the data processing node corresponding to node 1. Then, the local Merkle tree corresponding to the updated data shard of the data processing node corresponding to node 1 can be regenerated to obtain the updated local root corresponding to node 1, and the original local root of node 1 can be updated according to the updated local root corresponding to node 1. Since the local roots of other nodes are not updated, only the hash value of node 5 can be dynamically updated. Further, based on the dynamically updated hash value of node 5 and combined with the original hash value of node 6, the hash value of node 7 can be regenerated to obtain the updated global Merkle tree. Thus, it can be seen that applying a distributed multi-level Merkle tree to the data processing process of data shards can improve the update efficiency of the data storage path.
[0064] In an optional embodiment of the present invention, after generating the global Merkle tree, it may further include: obtaining the global root to be verified of the data shard to be verified in the shard verification data processing node; obtaining the global root stored in the global Merkle tree; comparing and verifying the global root to be verified and the global root stored in the global Merkle tree; verifying the digital signature of the global root stored in the global Merkle tree.
[0065] Among them, the shard verification data processing node may be a data processing node with data query and verification requirements. The data shard to be verified may be a data shard that needs to query and verify data integrity and accuracy. The number of data shards to be verified may be one or more, and the embodiments of the present invention do not limit this. The global root to be verified may be the global root of the global Merkle tree calculated layer by layer in real time according to the shard identifier and index information of the data shard to be verified in accordance with the global Merkle tree path information. The shard identifier may uniquely identify the data shard.
[0066] When it is necessary to verify the data shard, the data processing node where the data shard to be verified is located can be used as the shard verification data processing node. The shard verification data processing node generates the local root of a new local Merkle tree in real time by combining the data shard to be verified with other locally stored data shards, and then calculates the global root to be verified in real time by combining the hash values of other underlying leaf nodes of the global Merkle tree. The shard verification data processing node sends the global root to be verified to the central coordination node. After obtaining the global root to be verified, the central coordination node can obtain the global root stored in the global Merkle tree, and compare and verify the globally calculated global root to be verified and the global root stored in the global Merkle tree. If the two global roots are consistent, it means that the data verification passes, and the data is complete and has not been tampered with. If the two global roots are inconsistent, it means that the data verification fails, and there may be problems such as incomplete data or data being tampered with.
[0067] Alternatively, the shard verification data processing node may also generate a local root of a new local Merkle tree in real time based on the data shard to be verified in combination with other locally stored data shards, and send the locally generated root calculated in real time to the central coordination node. After the central coordination node calculates and generates the global root to be verified in real time based on the received local root to be verified, it can obtain the global root stored in the global Merkle tree, and compare and verify the globally generated root to be verified and the global root stored in the global Merkle tree. If the two global roots are consistent, it means that the data verification is passed, and the data is complete and has not been tampered with. If the two global roots are inconsistent, it means that the data verification fails, and there may be problems such as incomplete data or data being tampered with.
[0068] That is, the global root to be verified may be generated by the shard verification data processing node or by the central coordination node. The embodiments of the present invention do not limit the generation method of the global root to be verified.
[0069] To more clearly describe the technical solutions provided by the embodiments of the present invention, the embodiments of the present invention are illustrated by taking the log data processing as a specific application scenario example. The traditional method of directly performing CA signature on large log files faces problems such as high computing time consumption, difficulty in dynamic update, and long verification process. Aiming at the defects of single-machine signature processing, such as low efficiency, the need for full-scale re-signature after data update, and the need to traverse a large amount of data during verification, there is an urgent need for a distributed, incremental processing, and fast verification method.
[0070] In a specific example, the embodiments of the present invention propose a method for performing CA signature and verification on log data based on a multi-level Merkle tree, which mainly includes the following steps:
[0071] (1) Data sharding and preprocessing
[0072] The log management system (which may be composed of each data processing node and the central coordination node) can divide the large log file into data shards according to the dual criteria of data size (for example, every 10 MB) and time stamp (for example, every hour) through the central coordination node, and logically organize the data shards into subsets. Only the corresponding data shards are updated for the newly added log data, avoiding overall reconstruction.
[0073] Exemplarily, the 500GB log files generated every day can be split into approximately 50,000 basic data shards at a size of 10MB. At the same time, according to the timestamp information in the log records, every 50,000 shards are further grouped by hour (or other business rules) to form multiple subsets of data shards, and each subset can be sent to the corresponding data processing node for storage. When new log data is added to the log file, only local hash calculations and Merkle tree updates are performed on the newly added or affected data shards, avoiding the reconstruction of all data. Data compression and redundancy checks are performed before data shard transmission to ensure that the subsequent distributed processing nodes receive the data correctly.
[0074] (2), Construction of Distributed Multi-level Merkle Tree
[0075] (2.1), Construction of Local Merkle Tree
[0076] Each data processing node calculates the hash values of the data shards it is responsible for in parallel. For example, secure hash algorithms such as SHA-3 (Secure Hash Algorithm 3, the third-generation secure hash algorithm) can be used. For the 1,000 data shards processed by a data processing node, pairwise merging calculations are performed for every two data shards in sequence until a local Merkle tree with a depth of up to 16 layers is constructed to obtain a local root.
[0077] (2.2), Aggregation of Local Roots across Nodes
[0078] Each data processing node uploads the local root to the central coordination node. The central coordination node groups the local roots according to preset rules, such as by time series or geographical location, to construct a second-level global Merkle tree. If a global Merkle tree with 2 - 3 layers is constructed after aggregating 50 local roots, the global Merkle root consists of only a few consecutive hash calculations.
[0079] (2.3), Storage of Dynamic Merkle Path
[0080] The log management system can pre-record the index information and path information of each data shard in each level of the Merkle tree, so as to quickly extract the optimal verification path during verification and quickly complete the integrity verification only by using a few hash value calculations.
[0081] (3), CA Signature and Verification
[0082] (3.1), Generation of CA Signature
[0083] The central coordination node can submit the global Merkle root to the CA, and the CA uses its private key for digital signature. It can sign only 256-bit data, significantly reducing the signature calculation pressure.
[0084] (3.2) Data verification process
[0085] When it is necessary to verify specific log records, the log management system can automatically retrieve the corresponding dynamic Merkle path according to the timestamps and shard identifiers of the data shard records corresponding to the specific logs, recalculate the hash value of the data shard, and calculate layer by layer along the retrieved Merkle path. After obtaining the globally calculated global root, it is compared with the global root of the pre-stored global Merkle tree. Finally, the digital signature of the global root is verified using the CA public key to ensure that the data has not been tampered with since its generation.
[0086] In a specific example, taking a certain enterprise that generates 500GB of log data per day as an example, during the data sharding stage, the 500GB of data can be divided into approximately 50,000 10MB data shards, and each data shard contains approximately 1,000 log records; after further aggregating by hour, 50 subsets are formed. In the local processing stage, each subset is responsible for an independent data processing node, and a 16-layer local Merkle tree can be constructed. For example, node A processes 1,000 shards and finally generates a local root with a Merkle tree. In the global construction stage, the central coordination node aggregates the local roots from 50 data processing nodes to construct a 2-3 layer global Merkle tree to obtain the global root of the global Merkle tree. In the signature and verification stage, the CA can only sign the global root; during verification, only the Merkle path of approximately 20 hash values needs to be retrieved to complete the local data verification and ensure the integrity and authenticity of the verified log data.
[0087] The data processing method provided by the above technical solution can achieve incremental data sharding and update, support dynamic append local updates according to data shards without full-scale reconstruction. By constructing a distributed multi-level Merkle tree for data shards, it can efficiently utilize the multi-node parallel computing and hierarchical aggregation methods to ensure data integrity. At the same time, using low-overhead CA signatures for the global roots of the distributed multi-level Merkle tree, that is, only signing the fixed and extremely small amount of global roots, can greatly reduce the computing and transmission burden. The application of the distributed multi-level Merkle tree can achieve fast dynamic verification of data, that is, by pre-storing the dynamic Merkle path, it can quickly locate and verify any data shard.
[0088] It should be noted that the relevant information (including but not limited to user device information, log information, etc.) and data (including but not limited to data for display, data for analysis, etc.) involved in this disclosure are all information and data authorized by users or fully authorized by all parties, and the collection, use, and processing of the relevant data comply with the relevant laws, regulations, and standards of the relevant regions.
[0089] It should be noted that any permutation and combination of the technical features in the above embodiments also fall within the protection scope of the present invention.
[0090] Embodiment III
[0091] Figure 4 is a schematic diagram of a data processing device provided in Embodiment III of the present invention. As Figure 4 shown, the device includes: a local root acquisition module 310 and a global Merkle tree generation module 320, where:
[0092] The local root acquisition module 310 is configured to acquire the local roots of the local Merkle trees sent by each data processing node; wherein, the local Merkle tree is generated by the data processing node according to the local data shard.
[0093] The global Merkle tree generation module 320 is configured to use the local roots of the local Merkle trees as the bottom-layer leaf nodes of the global Merkle tree to generate the global Merkle tree.
[0094] In the embodiment of the present invention, after each data processing node generates a local Merkle tree according to the local data shard, each data processing node sends the local root of the generated local Merkle tree to the central coordination node. After the central coordination node acquires the local roots of the local Merkle trees sent by each data processing node, it can use the local roots of the local Merkle trees as the bottom-layer leaf nodes of the global Merkle tree to generate the global Merkle tree, which can solve the problems of long calculation time, difficulty in meeting the dynamic update requirements, and low verification efficiency existing in the prior art when processing data shards through the Merkle tree, can meet the dynamic update requirements of data shards, and improve the verification efficiency of data shards.
[0095] Optionally, the global Merkle tree generation module 320 is further configured to: determine a target node arrangement rule for each of the local roots as the bottom-layer leaf nodes of the global Merkle tree according to a target reference factor; arrange each of the local roots in sequence at the node positions of the bottom-layer leaf nodes of the global Merkle tree according to the target node arrangement rule to obtain the bottom-layer leaf nodes of the global Merkle tree.
[0096] Optionally, the target reference factor includes the generation time of the local root; the global Merkle tree generation module 320 is further configured to: determine the generation time of each local root; determine the node positions of each local root at the bottom leaf nodes of the global Merkle tree in the order of the generation time of each local root; arrange each local root in sequence according to the node positions of each local root at the bottom leaf nodes of the global Merkle tree to obtain the bottom leaf nodes of the global Merkle tree.
[0097] Optionally, the target reference factor includes the node geographical location of the data processing node; the global Merkle tree generation module 320 is further configured to: determine the node geographical location of the data processing node where each local root is located; determine the node positions of each local root at the bottom leaf nodes of the global Merkle tree according to the proximity order or node position weight of the node geographical locations corresponding to each local root; arrange each local root in sequence according to the node positions of each local root at the bottom leaf nodes of the global Merkle tree to obtain the bottom leaf nodes of the global Merkle tree.
[0098] Optionally, the target reference factor includes the data query frequency of the data processing node; the global Merkle tree generation module 320 is further configured to: determine the data query frequency of the data processing node where each local root is located or the data query frequency corresponding to the combination of data processing nodes where the local root combination is located; determine the node positions of each local root at the bottom leaf nodes of the global Merkle tree according to the magnitude order of the data query frequencies corresponding to each local root or the local root combination; arrange each local root in sequence according to the node positions of each local root at the bottom leaf nodes of the global Merkle tree to obtain the bottom leaf nodes of the global Merkle tree.
[0099] Optionally, the above device further includes a signature index path storage module, configured to: obtain the digital signature of the global root of the global Merkle tree; store the digital signature of the global root, as well as the index information and path information of each data shard in the local Merkle tree and the global Merkle tree.
[0100] Optionally, the above device further includes a data verification module, configured to: obtain the to-be-verified global root of the to-be-verified data shard in the shard verification data processing node; obtain the global root stored in the global Merkle tree; compare and verify the to-be-verified global root and the global root stored in the global Merkle tree; verify the digital signature of the global root stored in the global Merkle tree.
[0101] The above data processing device can execute the data processing method provided by any embodiment of the present invention, and has the corresponding functional modules and beneficial effects for executing the method. For the technical details not described in detail in this embodiment, reference can be made to the data processing method provided by any embodiment of the present invention.
[0102] Since the above-introduced data processing device is a device that can execute the data processing method in the embodiments of the present invention, based on the data processing method introduced in the embodiments of the present invention, those skilled in the art can understand the specific implementation manners and various variations of the data processing device in this embodiment. Therefore, the details of how the data processing device implements the data processing method in the embodiments of the present invention will not be described in detail here. As long as the device adopted by those skilled in the art to implement the data processing method in the embodiments of the present invention belongs to the scope to be protected by this application.
[0103] Embodiment 4
[0104] Figure 5 FIG. 10 shows a schematic structural diagram of an electronic device 10 that can be used to implement the embodiments of the present invention. The electronic device is intended to represent various forms of digital computers, such as, laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as, personal digital processors, cellular phones, smart phones, wearable devices (such as helmets, glasses, watches, etc.) and other similar computing devices. The components shown herein, their connections and relationships, and their functions are only examples and are not intended to limit the implementation of the present invention described and / or claimed herein.
[0105] As Figure 5 shown, the electronic device 10 includes at least one processor 11, and a memory communicatively connected to the at least one processor 11, such as a read-only memory (ROM) 12, a random access memory (RAM) 13, etc. Among them, the memory stores a computer program executable by the at least one processor. The processor 11 can execute various appropriate actions and processes according to the computer program stored in the read-only memory (ROM) 12 or the computer program loaded from the storage unit 18 into the random access memory (RAM) 13. In the RAM 13, various programs and data required for the operation of the electronic device 10 can also be stored. The processor 11, the ROM 12, and the RAM 13 are connected to each other through a bus 14. The input / output (I / O) interface 15 is also connected to the bus 14.
[0106] Multiple components in the electronic device 10 are connected to the I / O interface 15, including: an input unit 16, such as a keyboard, a mouse, etc.; an output unit 17, such as various types of displays, speakers, etc.; a storage unit 18, such as a disk, an optical disc, etc.; and a communication unit 19, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 19 allows the electronic device 10 to exchange information / data with other devices through a computer network such as the Internet and / or various telecommunication networks.
[0107] The processor 11 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the processor 11 include but are not limited to a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The processor 11 executes the various methods and processes described above, such as the data processing method.
[0108] Optionally, when the data processing method is applied to the central coordination node, it may include: obtaining the local roots of the local Merkle trees sent by each data processing node; wherein, the local Merkle trees are generated by the data processing nodes according to local data shards; using the local roots of the local Merkle trees as the bottom-level leaf nodes of the global Merkle tree to generate the global Merkle tree.
[0109] In some embodiments, the data processing method can be implemented as a computer program, which is tangibly contained in a computer-readable storage medium, such as the storage unit 18. In some embodiments, part or all of the computer program can be loaded and / or installed onto the electronic device 10 via the ROM 12 and / or the communication unit 19. When the computer program is loaded into the RAM 13 and executed by the processor 11, one or more steps of the data processing method described above can be executed. Alternatively, in other embodiments, the processor 11 can be configured to execute the data processing method in any other suitable way (e.g., by means of firmware).
[0110] The various embodiments of the systems and techniques described above in this specification can be implemented in digital electronic circuitry, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems-on-chip (SOCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include: being implemented in one or more computer programs that are executable and / or interpretable on a programmable system including at least one programmable processor, which can be a special-purpose or general-purpose programmable processor that receives data and instructions from, and transmits data and instructions to, a storage system, at least one input device, and at least one output device.
[0111] The computer programs for implementing the methods of the present invention can be written in any combination of one or more programming languages. These computer programs can be provided to a processor of a general purpose computer, special purpose computer, or other programmable data processing apparatus, such that the computer programs, when executed by the processor, cause the functions / operations specified in the flowchart and / or block diagram to be implemented. The computer programs can be executed entirely on the machine, partly on the machine, as a stand-alone software package partly on the machine and partly on a remote machine or entirely on the remote machine or server.
[0112] In the context of the present invention, a computer-readable storage medium can be a tangible medium that can contain or store a computer program for use by or in connection with an instruction execution system, apparatus, or device. The computer-readable storage medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. Alternatively, the computer-readable storage medium can be a machine-readable signal medium. More specific examples of the machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0113] To provide interaction with a user, the systems and techniques described herein can be implemented on an electronic device having: a display device for displaying information to the user (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor); and a keyboard and a pointing device (e.g., a mouse or a trackball) through which the user can provide input to the electronic device. Other kinds of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).
[0114] The systems and techniques described herein can be implemented in a computing system including backend components (e.g., as a data server), or a computing system including middleware components (e.g., an application server), or a computing system including frontend components (e.g., a user computer having a graphical user interface or a web browser through which the user can interact with an implementation of the systems and techniques described herein), or a computing system including any combination of such backend components, middleware components, or frontend components. The components of the system can be interconnected to each other by digital data communication in any form or medium (e.g., a communication network). Examples of communication networks include: local area network (LAN), wide area network (WAN), blockchain network, and the Internet.
[0115] A computing system can include a client and a server. The client and the server are generally remote from each other and typically interact through a communication network. The client-server relationship is created by computer programs running on the respective computers and having a client-server relationship with each other. The server can be a cloud server, also known as a cloud computing server or a cloud host, which is a host product in the cloud computing service system and solves the defects of difficult management and weak business scalability existing in traditional physical hosts and VPS services.
[0116] It should be understood that various forms of the processes shown above can be used, with steps reordered, added, or deleted. For example, the steps recited in this disclosure can be executed in parallel, sequentially, or in a different order, as long as the desired results of the technical solutions disclosed in this disclosure can be achieved, and this is not limited herein.
[0117] The above specific embodiments do not constitute a limitation on the protection scope of this disclosure. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this disclosure shall be included within the protection scope of this disclosure.
Claims
1. A data processing method, characterized in that: Applied to the central coordination node, including: Obtaining a local root of a local Merkle tree sent by each data processing node; wherein the local Merkle tree is generated by the data processing node according to the local data sharding; The local roots of the local Merkle trees are used as bottom leaf nodes of the global Merkle tree to generate the global Merkle tree.
2. The method according to claim 1, characterized in that The method of using the local roots of the local Merkle trees as bottom leaf nodes of the global Merkle tree includes: Determine, according to the target reference factor, a target node arrangement rule for each of the local roots as a bottom leaf node of the global Merkle tree; The local roots are arranged in sequence at the node positions of the bottom leaf nodes of the global Merkle tree according to the target node arrangement rule to obtain the bottom leaf nodes of the global Merkle tree.
3. The method according to claim 2, characterized in that The target reference factor includes the generation time of the local root; the step of sequentially arranging the local roots at the node positions of the bottom leaf nodes of the global Merkle tree according to the target node arrangement rule to obtain the bottom leaf nodes of the global Merkle tree includes: Determining the generation time of each of the local roots; Determine the node position of each local root at the bottom leaf node of the global Merkle tree according to the order of generation time of each local root; The local roots are arranged in sequence according to the node positions of the bottom leaf nodes of the global Merkle tree to obtain the bottom leaf nodes of the global Merkle tree.
4. The method according to claim 2, characterized in that: The target reference factor includes the node geographical location of the data processing node; the local roots are arranged in sequence at the node positions of the bottom leaf nodes of the global Merkle tree according to the target node arrangement rule to obtain the bottom leaf nodes of the global Merkle tree, including: Determine the node geographical location of the data processing node where each of the local roots is located; Determine the node position of each local root at the bottom leaf node of the global Merkle tree according to the proximity order or node position weight of the geographical location of the nodes corresponding to each local root; The local roots are arranged in sequence according to the node positions of the bottom leaf nodes of the global Merkle tree to obtain the bottom leaf nodes of the global Merkle tree.
5. The method according to claim 2, characterized in that: The target reference factor includes a data query frequency of a data processing node; and the local roots are sequentially arranged at the node position of the bottom leaf node of the global Merkle tree according to the target node arrangement rule to obtain the bottom leaf node of the global Merkle tree, including: Determine the data query frequency of the data processing node where each local root is located or the data query frequency corresponding to the data processing node combination where the local root combination is located; Determine the node position of each local root at the bottom leaf node of the global Merkle tree according to the order of the data query frequency corresponding to each local root or the combination of local roots; The local roots are arranged in sequence according to the node positions of the bottom leaf nodes of the global Merkle tree to obtain the bottom leaf nodes of the global Merkle tree.
6. The method according to claim 1, characterized in that After generating the global Merkle tree, the method further includes: Obtaining a digital signature of the global root of the global Merkle tree; The digital signature of the global root, as well as the index information and path information of each data shard in the local Merkle tree and the global Merkle tree are stored.
7. The method according to claim 6, characterized in that After generating the global Merkle tree, the method further includes: Obtain the global root to be verified of the data shard to be verified in the shard verification data processing node; Obtaining a global root stored for the global Merkle tree; Compare and verify the global root to be verified with the global root stored in the global Merkle tree; The digital signature of the global root stored in the global Merkle tree is verified.
8. An electronic device, characterized in that: The electronic device comprises: at least one processor; and a memory communicatively connected to the at least one processor; wherein, The memory stores a computer program executed by the at least one processor, and the computer program is executed by the at least one processor so that the at least one processor can execute the data processing method according to any one of claims 1 to 7.
9. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores computer instructions, and the computer instructions are used to enable a processor to implement the data processing method described in any one of claims 1 to 7 when executed.
10. A computer program product comprising a computer program / instructions, wherein: When the computer program / instructions are executed by a processor, the data processing method described in any one of claims 1 to 7 is implemented.