Data processing method and device, electronic equipment, storage medium and program

By generating local Merkle trees in the TEE environment of the data processing node, generating global Merkle trees at the central coordination point and performing digital signatures, the problems of low data shard verification efficiency and difficulty in dynamic updates are solved, and efficient and secure data shard verification is achieved.

CN120067108APending Publication Date: 2025-05-30BEIJING YOUTEJIE INFORMATION TECH
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510222382.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-27
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

When processing data sharding, the calculation of Merkle tree is long and complex, and it is difficult to meet the dynamic update requirements. The verification process is long and ignores the security and reliability of the Merkle tree generation process.

Method used

By sharding the original data at the central coordinating point, and generating a local Merkle tree in the TEE environment built-in trusted execution environment in the data processing node, obtaining the local roots, generating a global Merkle tree at the central coordinating point, and digitally signing in the TEE environment, and finally writing the signature information to the blockchain storage.

Benefits of technology

It solves the problem that Merkle tree computing takes time and is difficult to meet the needs of dynamic updates, improves the efficiency and security of data shard verification, and ensures the security and reliability of the Merkle tree generation process.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120067108A_ABST
    Figure CN120067108A_ABST
Patent Text Reader

Abstract

The embodiment of the invention discloses a data processing method and device, electronic equipment, a storage medium and a program, and the method comprises the steps: carrying out the fragmentation of original data, obtaining data fragments, and transmitting the data fragments to a data processing node; acquiring a local root of a local Merkle tree sent by each data processing node; the local Merkle tree is generated by the data processing node in a built-in TEE environment according to local fragment data; taking the local root of each local Merkle tree as a bottom-layer leaf node of a global Merkle tree to generate the global Merkle tree; performing digital signature on the global root of the global Merkle tree in a local TEE environment to obtain a global root signature; and writing the global root signature, the target local root and the timestamp information of the global root signature into a block chain for storage. According to the technical scheme, the dynamic updating requirement of the data fragments can be met, and the verification efficiency of the data fragments and the safety of verification operation are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Embodiments of the present invention relate to the technical field of data processing, and in particular, to a data processing method, apparatus, electronic device, storage medium, and program. Background Art

[0002] Data sharding is the technical foundation of distributed storage. It divides data into multiple parts, and each part is stored on different nodes. Data sharding is often applied to distributed systems, large-scale data storage and processing, cloud computing, database management, and ETL (Extract, Transform, Load) technology scenarios.

[0003] A Merkle Tree (also known as a hash tree) is a binary tree structure widely used in data verification and integrity checking. The application of the Merkle Tree in data sharding mainly ensures the integrity, traceability, and cross-shard cooperation efficiency of data sharding through its efficient hash verification mechanism.

[0004] In the process of implementing the present invention, the inventors found that the prior art has the following defects: Since the data sharding scenario usually involves a relatively large amount of data, the calculation of the Merkle Tree takes a relatively long time and the structure is relatively complex, thereby reducing the verification efficiency of data sharding. When some data shards are updated, it is necessary to traverse and calculate the entire tree structure of the Merkle Tree again, which is difficult to meet the dynamic update requirements of data sharding. At the same time, when using a Merkle Tree with a complex structure to verify the data integrity and consistency of data shards through a hash calculation and verification mechanism, there will also be a problem of a long verification process. In addition, when the prior art uses a Merkle Tree to perform hash verification on data shards, it often ignores the security and reliability requirements in the process of generating the Merkle Tree. Summary of the Invention

[0005] Embodiments of the present invention provide a data processing method, apparatus, electronic device, storage medium, and program, which can meet the dynamic update requirements of data sharding and improve the efficiency of data sharding verification and the security of verification operations.

[0006] According to one aspect of the present invention, there is provided a data processing method applied to a central coordination node, including:

[0007] Shard the original data to obtain data shards, and send the data shards to data processing nodes;

[0008] Obtain the local roots of the local Merkle Trees sent by each of the data processing nodes; wherein, the local Merkle Trees are generated by the data processing nodes according to local shard data in a built-in trusted execution environment (TEE) environment;

[0009] Use the local roots of each of the local Merkle trees as the bottom - layer leaf nodes of the global Merkle tree to generate the global Merkle tree;

[0010] Perform a digital signature on the global root of the global Merkle tree in the local TEE environment to obtain a global root signature;

[0011] Write the global root signature, the target local root, and the timestamp information of the global root signature into the blockchain storage.

[0012] According to another aspect of the present invention, there is provided a data processing device configured in a central coordination node, including:

[0013] A data sharding and sending module, configured to shard the original data to obtain data shards and send the data shards to data processing nodes;

[0014] A local root obtaining module, configured to obtain the local roots of the local Merkle trees sent by each of the data processing nodes; wherein, the local Merkle tree is generated by the data processing node in the built - in trusted execution environment (TEE) according to the local sharded data;

[0015] A global Merkle tree generating module, configured to use the local roots of each of the local Merkle trees as the bottom - layer leaf nodes of the global Merkle tree to generate the global Merkle tree;

[0016] A global root signature generating module, configured to perform a digital signature on the global root of the global Merkle tree in the local TEE environment to obtain a global root signature;

[0017] A data on - chain storage module, configured to write the global root signature, the target local root, and the timestamp information of the global root signature into the blockchain storage.

[0018] According to another aspect of the present invention, there is provided an electronic device, the electronic device including:

[0019] At least one processor; and

[0020] A memory communicatively connected to the at least one processor; wherein,

[0021] The memory stores a computer program executable by the at least one processor, and when the computer program is executed by the at least one processor, the at least one processor is enabled to execute the data processing method according to any embodiment of the present invention.

[0022] According to another aspect of the present invention, there is provided a computer-readable storage medium storing computer instructions for causing a processor to implement the data processing method according to any embodiment of the present invention when executed.

[0023] According to another aspect of the present invention, there is also provided a computer program product including a computer program which implements the data processing method according to any embodiment of the present invention when executed by a processor.

[0024] In the embodiment of the present invention, the central coordination node slices the original data to obtain data slices and sends the data slices to the data processing nodes. After each data processing node generates a local Merkle tree based on its local data slice, it sends the local root of the generated local Merkle tree to the central coordination node. After the central coordination node obtains the local roots of the local Merkle trees sent by each data processing node, it can use the local roots of each local Merkle tree as the underlying leaf nodes of the global Merkle tree to generate a global Merkle tree, and digitally sign the global root of the global Merkle tree in the local TEE environment to obtain a global root signature. Finally, the central coordination node writes the global root signature, the target local root, and the timestamp information of the global root signature into the blockchain storage. The above technical solution can solve the problems of long computing time, difficulty in meeting the dynamic update requirements, and low verification efficiency and security when processing data slices through the Merkle tree, and can meet the dynamic update requirements of data slices, improve the verification efficiency of data slices, and the security of verification operations.

[0025] It should be understood that the content described in this part is not intended to identify the key or important features of the embodiments of the present invention, nor is it used to limit the scope of the present invention. Other features of the present invention will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0026] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present invention, and those of ordinary skill in the art can also obtain other drawings without creative efforts based on these drawings.

[0027] Figure 1 is a flowchart of a data processing method provided in Embodiment 1 of the present invention;

[0028] Figure 2 is a flowchart of a data processing method provided in Embodiment 2 of the present invention;

[0029] Figure 3It is a schematic structural diagram of a Merkle tree provided in the second embodiment of the present invention;

[0030] Figure 4 It is a schematic diagram of a data processing device provided in the third embodiment of the present invention;

[0031] Figure 5 It is a schematic structural diagram of an electronic device provided in the fourth embodiment of the present invention. Detailed implementation manners

[0032] In order to enable those skilled in the art to better understand the solution of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without making creative efforts shall fall within the protection scope of the present invention.

[0033] It should be noted that the terms "first", "second", etc. in the specification and claims of the present invention and the above-mentioned drawings are used to distinguish similar objects, and do not have to be used to describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances so that the embodiments of the present invention described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device comprising a series of steps or units does not have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.

[0034] Embodiment 1

[0035] Figure 1 It is a flowchart of a data processing method provided in the first embodiment of the present invention. This embodiment is applicable to the situation of managing data shards according to a local Merkle tree and a global Merkle tree. This method can be executed by a data processing device, which can be implemented in a software and / or hardware manner and is generally integrated in an electronic device. The electronic device can be a terminal device or a server device, as long as it can execute the data processing method. The specific device type of the electronic device is not limited in the embodiments of the present invention. Correspondingly, as Figure 1 shown, the method includes the following operations:

[0036] S110. Shard the original data to obtain data shards, and send the data shards to the data processing nodes.

[0037] Among them, the original data can be of any data type, for example, it can include but is not limited to log data, transaction data, and background data of application programs, etc. The embodiments of the present invention do not limit the specific type and data content of the original data. Data sharding is also a part of the data obtained by performing sharding processing on the original data. The data processing node can be a device node for storing data shards. The central coordination node can be one of the data processing nodes selected from each data processing node, or it can also be a node with central control functions. The embodiments of the present invention do not limit the node type of the central coordination node.

[0038] First of all, the technical solution of the embodiments of the present invention is mainly applicable to such an application scenario: the central coordination node and each data processing node form a data processing system. The central coordination node first performs sharding processing on a large amount of original data to obtain multiple data shards, and sends each data shard to the corresponding data processing node for storage. Since the data processing node needs to generate a local Merkle tree based on the data shards stored locally, and the local Merkle tree constructs the global Merkle tree as the underlying leaf nodes of the global Merkle tree. The distributed multi-level Merkle tree method jointly provides the query verification function for the data shards. Therefore, when the central coordination node sends the data shards to the data processing nodes, it can send the data shards with similar data query verification frequencies to the corresponding data processing nodes. For example, if the query verification frequency of data shard 1 is 100 times per month, the query verification frequency of data shard 3 is 100 times per month, and the query verification frequency of data shard 2 is 10 times per month, then data shard 1 and data shard 3 can be sent to the same data processing node. The advantage of this setting is that the query verification frequencies of the data shards stored by the data processing node are similar, so the probability of performing data query verification based on the same data processing node is higher, and the calculation amount of recalculating the hash value based on the distributed multi-level Merkle tree is less, thus improving the efficiency of data query verification.

[0039] S120. Obtain the local roots of the local Merkle trees sent by each of the data processing nodes; wherein, the local Merkle tree is generated by the data processing node according to the local shard data in the built-in trusted execution environment (TEE) environment.

[0040] Optionally, a data processing node can store one or more data shards. The embodiments of the present invention do not limit the number of data shards stored by the data processing node. A data processing node can generate a local Merkle tree based on the local data shards, and each data processing node can send the root of the generated local Merkle tree as the local root to the central coordination node, and the central coordination node generates a global Merkle tree according to the local roots of each local Merkle tree.

[0041] In an embodiment of the present invention, each data processing node may be embedded with a TEE (Trusted Execution Environment) module for storing private keys and performing secure computations. The TEE module can provide a TEE environment, which is an isolated and protected operating environment inside the processor for executing security-sensitive code and storing confidential data. The TEE is separated from the device's main operating system. Even if the main system is attacked, the code and data in the TEE can still be guaranteed. Therefore, to further ensure the security of the Merkle tree generation process, each data processing node may use a secure hash algorithm in its built-in TEE environment to generate a local Merkle tree based on local shard data. Optionally, the data processing node may further use Secure Muti-party Computation (SMPC) to generate a digital signature for the local root within the TEE environment. Since the private key never leaves the TEE, the signature operation of the local root can be cooperatively completed without exposing the private key, ensuring that the digital signature process of the local root is carried out in a hardware isolation environment. Secure multi-party computation is a cryptographic technology that allows multiple participants. Each participant only obtains the final result without knowing the input details of other parties, thereby ensuring the privacy and security of the data of all parties. On the premise of each retaining its private data, they jointly calculate the result of a certain function without disclosing their respective input information. Therefore, using secure multi-party computation to generate a digital signature for the local root within the TEE environment can effectively disperse risks and prevent single-point failures. Optionally, a timestamp may also be embedded in the digital signature of the local root, and an anti-replay function may be configured to ensure the timeliness and uniqueness of the digital signature of the local root, further improving the security of the local root. Correspondingly, the data processing node may also send the digital signature of the local root to the central coordination node.

[0042] S130. Use the local roots of the local Merkle trees as the bottom-layer leaf nodes of the global Merkle tree to generate the global Merkle tree.

[0043] It can be understood that the local roots of the local Merkle trees sent by each data processing node are also hash values. After the central coordination node obtains the local roots of the local Merkle trees sent by each data processing node, it can use the local roots of the local Merkle trees as the bottom-layer leaf nodes of the global Merkle tree, arrange the bottom-layer leaf nodes of the global Merkle tree, group and aggregate the bottom-layer leaf nodes, and calculate the hash values of the intermediate nodes layer by layer in the order from bottom to top of the bottom-layer leaf nodes until the root value of the global Merkle tree is generated to obtain the global Merkle tree.

[0044] In the prior art, usually, the hash values obtained by fragmenting and calculating all the data stored in a data processing node are directly used as the underlying leaf nodes of the Merkle tree, and then the Merkle tree is generated. When the capacity of the original data is large, the data calculation takes a long time, and the hash calculation needs to process more data, which will lead to an extended calculation time, thereby increasing the generation time and verification time of the Merkle tree. When some data fragments are updated, the entire tree structure of the Merkle tree needs to be recalculated.

[0045] In the embodiments of the present invention, at the data processing node, a local Merkle tree corresponding to the data processing node is first calculated for the data fragments. The underlying leaf nodes of the local Merkle tree can be one or a small number of data fragments. The magnitude of the data fragments is significantly reduced compared to all data fragments. When calculating the local Merkle tree with them as the underlying leaf nodes, the time consumption is less, and the generation efficiency and verification efficiency are higher. Then, a distributed multi-level Merkle tree is formed by the local Merkle tree and the global Merkle tree, which can improve the generation efficiency of the global Merkle tree. When some data fragments are updated, only the corresponding local Merkle tree needs to be recalculated, and then the global Merkle tree is dynamically and quickly updated.

[0046] S140. Digitally sign the global root of the global Merkle tree in the local TEE environment to obtain a global root signature.

[0047] Among them, the global root signature is also the signature information obtained by digitally signing the global root.

[0048] The central coordination node can also be configured with a TEE module. After the central coordination node generates the global Merkle tree, it can use the hardware security module and the CA (Certificate Authority) private key to digitally sign the global root of the global Merkle tree in the local TEE environment to ensure that the signature process of the global root is executed in a highly isolated secure environment, and a secure and reliable global root signature is obtained.

[0049] In the embodiments of the present invention, by signing the local root and the global root in the TEE environment, and by the data processing node generating the local Merkle tree according to the local shard data in the TEE environment, it is possible to use the TEE to ensure that the private key is not leaked, and all signature operations and the generation operations of the local Merkle tree are executed in a hardware isolation environment to prevent malicious attacks.

[0050] S150. Write the global root signature, the target local root, and the timestamp information of the global root signature into the blockchain storage.

[0051] Among them, the target local root can be part or all of the local roots. Optionally, the target local root can be the relatively critical local roots selected from each local root.

[0052] Traditional CA signatures are vulnerable to private key protection, signature latency, and full re-signature during data dynamic updates in big data file processing. To address issues such as private key leakage, single point of failure, and non-real-time verification, there is an urgent need for a secure and efficient incremental signature and verification scheme, while providing an immutable audit traceability mechanism.

[0053] Blockchain is a decentralized distributed ledger technology. Its basic principle is to divide data into a series of "blocks" and use cryptographic methods such as hash algorithms and digital signatures to link each block together to form an immutable data chain. Each block in the blockchain contains not only the hash of the current data but also the hash of the previous block, which makes any tampering with historical data invalidate all subsequent block data on the chain, achieving the immutability of data. The blockchain adopts a consensus mechanism and realizes decentralized verification without a centralized trust institution through a distributed consensus algorithm.

[0054] Therefore, to further ensure the security and immutability of data, the embodiments of the present invention can also write the global root signature, the target local root, and the timestamp information of the global root signature into the blockchain storage through an encrypted channel by the central coordination node. Optionally, the target local root can be the local root generated by important data processing nodes. Important data processing nodes can be, for example, data processing nodes that often have incremental data. In this way, an immutable distributed audit record can be formed for the global root signature, the target local root, and the timestamp information of the global root signature based on the blockchain, realizing the immutability and distributed audit of data signatures and local roots, providing distributed credible evidence for subsequent audits and traceability, and facilitating post-event traceability.

[0055] At the same time, the data processing node and the central coordination node also need to synchronize data with the blockchain regularly to ensure that the signature status of each node is consistent with the global signature, facilitating the detection of anomalies and the activation of the re-signature mechanism.

[0056] In an embodiment of the present invention, the central coordination node slices the original data to obtain data slices, and sends the data slices to the data processing nodes. After each data processing node generates a local Merkle tree based on its local data slice, it sends the local root of the generated local Merkle tree to the central coordination node. After the central coordination node obtains the local roots of the local Merkle trees sent by each data processing node, it can use the local roots of each local Merkle tree as the bottom leaf nodes of the global Merkle tree to generate the global Merkle tree, and digitally sign the global root of the global Merkle tree in the local TEE environment to obtain the global root signature. Finally, the central coordination node writes the global root signature, the target local root, and the timestamp information of the global root signature into the blockchain storage. The above technical solution can solve the problems such as long calculation time, difficulty in meeting the dynamic update requirements, and low verification efficiency and security when processing data slices through the Merkle tree, and can meet the dynamic update requirements of data slices, improve the verification efficiency of data slices, and the security of verification operations.

[0057] Embodiment 2

[0058] Figure 2 FIG. is a flowchart of a data processing method provided in Embodiment 2 of the present invention. This embodiment is a specific implementation based on the above embodiment. In this embodiment, various specific and optional implementation manners for sending data slices to data processing nodes, generating the global Merkle tree, and updating data slices are given. Correspondingly, as Figure 2 shown, the method of this embodiment may include:

[0059] S210. Obtain the slice attribute configuration information of the data slice.

[0060] S220. Determine the first target data processing node corresponding to the data slice according to the slice attribute configuration information of the data slice, and send the data slice to the first target data processing node.

[0061] Among them, the slice attribute configuration information may be attribute information and other configuration information configured for the data slice, etc. In the embodiments of the present invention, the slice attribute configuration information may include, but is not limited to, at least one of the data operation authorization object of the data slice, the storage type identifier of the data slice, and the query verification frequency of the data slice.

[0062] Among them, the data operation authorization object, that is, the object with data operation permissions for data shards, can be a user or an enterprise account, etc. The embodiments of the present invention do not limit the object type of the data operation authorization object. The storage type identifier can be used to identify the storage method of the data shard. In the embodiments of the present invention, the storage method of the data shard may include, but is not limited to, cold storage and hot storage, etc. Cold storage refers to a data storage method with extremely low access frequency and extremely high requirements for long-term preservation and security, while hot storage is a data storage method for frequently accessed or updated data. It should be noted that the data shards of cold storage and the data shards of hot storage can be divided according to actual needs. For example, the data shards that have not been updated for a certain period of time (such as 1 month) are divided into data shards of cold storage, and the data shards that are updated dynamically regularly or irregularly are divided into data shards of hot storage, etc. The embodiments of the present invention do not limit the data type and data content of the data shards of cold storage and the data shards of hot storage. The first target data processing node can be determined for the data shard and can perform data storage.

[0063] Before sending the data shard to the data processing node, the central coordination node can first obtain various types of shard attribute configuration information such as the data operation authorization object, storage type identifier, and query verification frequency of the data shard, so as to determine the first target data processing node corresponding to the data shard according to the shard attribute configuration information of the data shard, and then send the data shard to the first target data processing node.

[0064] In a specific example, if the data shards are allocated to each data processing node based on the data operation authorization object of the data shard, the data operation authorization object corresponding to each data shard can be determined, and the data shards belonging to the same data operation authorization object are sent to the same first target data processing node for storage. If the data shards of a data operation authorization object require multiple data processing nodes, the data shards of the data operation authorization object can be sent to several data processing nodes with strong correlation. For example, the data shards of a data operation authorization object can be sent to several first target data processing nodes with concentrated geographical locations, or the data shards of a data operation authorization object can also be sent to several first target data processing nodes with relatively high query frequencies, etc.

[0065] In a specific example, if data shards are allocated to each data processing node based on the storage type identifier of the data shards, the storage type identifier corresponding to each data shard can be determined, and the data shards can be sent to the corresponding data processing nodes for storage according to the storage type identifier corresponding to the data shards. Exemplarily, if the storage type identifier of the data shard is cold storage, data processing nodes that can store data in the cold storage manner can be screened from the data processing nodes as the first target data processing nodes to store the data shard; if the storage type identifier of the data shard is hot storage, data processing nodes that can store data in the hot storage manner can be screened from the data processing nodes as the first target data processing nodes to store the data shard. The storage type identifier of the data shard can be pre-configured by the data operation authorization object providing the data shard, or can also be automatically determined by the central coordination node according to the data content. The embodiments of the present invention do not limit this.

[0066] In a specific example, if data shards are allocated to each data processing node based on the query verification frequency of the data shards, the query verification frequency corresponding to each data shard can be determined, and the data shards can be sent to the corresponding data processing nodes for storage according to the query verification frequency corresponding to the data shards. Exemplarily, if the query verification frequency of the data shard is 100 times per month, data processing nodes with a query verification frequency close to it can be screened from the data processing nodes, such as screening data processing nodes with a statistically queried verification frequency of 90 times per month - 110 times per month as the first target data processing nodes to store the data shard. In this way, data shards with similar query verification frequencies can be stored centrally, thereby improving the efficiency of subsequent data query verification. The query verification frequency of the data shards can be pre-configured by the data operation authorization object providing the data shards, or can also be automatically determined by the central coordination node according to the data content. The embodiments of the present invention do not limit this.

[0067] S230. Obtain the local roots of the local Merkle trees sent by each of the data processing nodes; wherein, the local Merkle trees are generated by the data processing nodes according to local shard data in the built-in trusted execution environment (TEE) environment.

[0068] S240. Use the local roots of the local Merkle trees as the bottom-level leaf nodes of the global Merkle tree to generate the global Merkle tree.

[0069] In an alternative embodiment of the present invention, the step of using the local roots of the local Merkle trees as the bottom-level leaf nodes of the global Merkle tree may include: determining a target node arrangement rule for arranging each of the local roots as the bottom-level leaf nodes of the global Merkle tree according to a target reference factor; arranging each of the local roots in sequence at the node positions of the bottom-level leaf nodes of the global Merkle tree according to the target node arrangement rule, so as to obtain the bottom-level leaf nodes of the global Merkle tree.

[0070] Among them, the target reference factor may be a factor used as a reference for arranging the local roots at the positions of the bottom-level leaf nodes. The target node arrangement rule is a specific rule for arranging the local roots at the positions of the bottom-level leaf nodes.

[0071] When the central coordination node uses the local roots of the local Merkle trees as the bottom-level leaf nodes of the global Merkle tree, it may screen some factors affecting data verification as the target reference factors according to the law of later data verification. By determining the target reference factors, the rule for arranging the local roots is determined as the target node arrangement rule, and each of the local roots is arranged in sequence at the node positions of the bottom-level leaf nodes of the global Merkle tree, so that the bottom-level leaf nodes of the global Merkle tree can be obtained.

[0072] Optionally, the target reference factor may include but is not limited to the generation time of the local root, the geographical location of the data processing node where the local root is located, or the data query frequency of the data processing node where the local root is located, etc. The bottom-level leaf nodes of the global Merkle tree arranged according to the target node arrangement rule determined by the target reference factor have a relatively strong correlation between adjacent nodes.

[0073] It can be understood that for a Merkle tree, the stronger the correlation between adjacent nodes, the higher the efficiency of data verification using the Merkle tree, and the higher the update efficiency of the Merkle tree. Figure 3 This is a schematic structural diagram of a Merkle tree provided in the second embodiment of the present invention. In a specific example, as Figure 3 shown, assume that the global Merkle tree includes 4 bottom-level leaf nodes, namely node 1, node 2, node 3, and node 4. Every two bottom-level leaf nodes can be calculated to generate an intermediate node, obtaining 2 intermediate nodes of node 5 and node 6. The 2 intermediate nodes can be calculated to generate the root node of node 7. Figure 3All nodes of the Merkle tree shown are hash values. If two adjacent bottom-level leaf nodes, node 1 and node 2, are strongly correlated, when performing data verification or an update operation on the Merkle tree, if it is necessary to verify or update some data in the data slices corresponding to node 1 and node 2, then only the hash values of node 1 and node 2 need to be recalculated, and the hash values of node 3 and node 4 do not need to be recalculated repeatedly, and the new root can be quickly calculated layer by layer; if two non-adjacent bottom-level leaf nodes, node 1 and node 3, are strongly correlated, when performing data verification or an update operation on the Merkle tree, if it is necessary to verify or update some data in the data slices corresponding to node 1 and node 3, then it is necessary to recalculate the hash values of all bottom-level leaf nodes, and calculate the root of the Merkle tree layer by layer based on the recalculated hash values of the bottom-level leaf nodes, and the efficiency is relatively low.

[0074] In an alternative embodiment of the present invention, the target reference factor may include the generation time of the local root; arranging each of the local roots in sequence at the node positions of the bottom-level leaf nodes of the global Merkle tree according to the target node arrangement rule to obtain the bottom-level leaf nodes of the global Merkle tree may include: determining the generation time of each of the local roots; determining the node positions of each of the local roots at the bottom-level leaf nodes of the global Merkle tree according to the chronological order of the generation times of each of the local roots; arranging each of the local roots in sequence at the node positions of the bottom-level leaf nodes of the global Merkle tree to obtain the bottom-level leaf nodes of the global Merkle tree.

[0075] Optionally, when the target reference factor includes the generation time of the local root, the time when each data processing node generates the local root of the local Merkle tree can be determined. Generally, the closer the generation times of two local roots are, the closer the generation times of the data slices stored by the data processing nodes corresponding to the two local roots are, and further indicates that the data slices stored by the data processing nodes corresponding to the two local roots are more strongly correlated. Therefore, when determining the node positions of each of the local roots at the bottom-level leaf nodes of the global Merkle tree according to the chronological order of the generation times of each of the local roots, the local roots with stronger data correlation can be configured at adjacent bottom-level leaf nodes. In this way, further arranging each of the local roots in sequence at the node positions of the bottom-level leaf nodes of the global Merkle tree to obtain the bottom-level leaf nodes of the global Merkle tree, the correlation between adjacent bottom-level leaf nodes is stronger.

[0076] In an alternative embodiment of the present invention, the target reference factor includes the node geographical location of the data processing node; arranging each of the local roots in sequence at the node positions of the bottom-layer leaf nodes of the global Merkle tree according to the target node arrangement rule to obtain the bottom-layer leaf nodes of the global Merkle tree may include: determining the node geographical location of the data processing node where each of the local roots is located; determining the node positions of each of the local roots at the bottom-layer leaf nodes of the global Merkle tree according to the proximity order or node position weights of the node geographical locations corresponding to each of the local roots; arranging each of the local roots in sequence according to the node positions of each of the local roots at the bottom-layer leaf nodes of the global Merkle tree to obtain the bottom-layer leaf nodes of the global Merkle tree.

[0077] Among them, the node geographical location may be the geographical location where the data processing node is located. The node position weight may be a weight value assigned to each data processing node, which can be used to characterize the importance of the geographical location where the data processing node is located.

[0078] Optionally, when the target reference factor includes the node geographical location of the data processing node, the node geographical location of the data processing node where each local root is located can be determined. Generally, the closer the node geographical locations of the data processing nodes where two local roots are located, the stronger the correlation of the data shards stored by the data processing nodes corresponding to the two local roots. For example, the data shards stored by data processing nodes with similar node geographical locations may be data shards from the same user or the same equipment manufacturer. Therefore, when determining the node positions of each local root at the bottom-layer leaf nodes of the global Merkle tree according to the proximity order of the node geographical locations corresponding to each local root, local roots with stronger data correlation can be configured at adjacent bottom-layer leaf nodes. In this way, further arranging each local root in sequence according to the node positions of each local root at the bottom-layer leaf nodes of the global Merkle tree to obtain the bottom-layer leaf nodes of the global Merkle tree, the correlation between adjacent bottom-layer leaf nodes is relatively strong.

[0079] It can be understood that some data processing nodes can also mark the node location weights according to the importance level for their corresponding node geographical locations. The larger the node location weight, the higher the importance level of the data processing node, and the higher the importance level of the data shards stored by it. For example, if the node location weights corresponding to the data processing nodes arranged in City A are 0.5, and the node location weights corresponding to the data processing nodes arranged in City B are 0.2, then the data shards stored by the data processing nodes arranged in City A are generally more important than those stored by the data processing nodes arranged in City B. At the same time, the closer the node location weights of the data processing nodes are, the stronger the correlation of the data shards stored by them. For example, when the node location weight corresponding to data processing node 1 is 0.5, the node location weight corresponding to data processing node 2 is 0.5, and the node location weight corresponding to data processing node 3 is 0.2, it indicates that the data shards stored by data processing node 1 and data processing node 2 are usually data shards of the same importance level. Therefore, the local roots with similar node location weights can also be arranged as close as possible at the adjacent node locations of the bottom-layer leaf nodes of the global Merkle tree according to the magnitude relationship of the node location weights corresponding to the node geographical locations of each local root. In this way, further arranging each local root in sequence according to the node locations of the bottom-layer leaf nodes of the global Merkle tree, the correlation between adjacent bottom-layer leaf nodes in the bottom-layer leaf nodes of the global Merkle tree is relatively strong.

[0080] In an alternative embodiment of the present invention, the target reference factor includes the data query frequency of the data processing node; the arranging each of the local roots in sequence at the node locations of the bottom-layer leaf nodes of the global Merkle tree according to the target node arrangement rule to obtain the bottom-layer leaf nodes of the global Merkle tree may include: determining the data query frequency of the data processing node where each of the local roots is located or the data query frequency corresponding to the combination of data processing nodes where the combination of local roots is located; determining the node locations of each of the local roots at the bottom-layer leaf nodes of the global Merkle tree according to the magnitude order of the data query frequencies corresponding to each of the local roots or the combination of local roots; and arranging each of the local roots in sequence according to the node locations of each of the local roots at the bottom-layer leaf nodes of the global Merkle tree to obtain the bottom-layer leaf nodes of the global Merkle tree.

[0081] The data query frequency may be the frequency of query verification of the data shards stored by the data processing node. The combination of local roots may be a combination composed of multiple local roots. The combination of data processing nodes may be a combination composed of multiple data processing nodes.

[0082] Optionally, when the target reference factor includes the data query frequency of the data processing nodes, the data query frequency of the data processing nodes where each local root is located or the data query frequency corresponding to the combination of data processing nodes where the local root combination is located can be determined. Generally, the higher the data query frequency of a data processing node or a combination of data processing nodes, the higher the verification requirement of such data processing nodes for the locally stored data. The closer the data query frequencies of different data processing nodes are, the closer the data query verification requirements of such data processing nodes are. Exemplarily, when the data query frequency of data processing node 1 is 100 times per month, the data query frequency of data processing node 2 is 10 times per month, and the data query frequency of data processing node 3 is 110 times per month, data processing node 1 and data processing node 3 can be used as a group of adjacent bottom-level leaf nodes that can generate hash values. Correspondingly, when data processing node 1 and data processing node 3 perform data query verification later, in the global Merkle tree, only the hash values of data processing node 1 and data processing node 3 need to be recalculated, and the hash values of other nodes that have not been queried and verified do not need to be recalculated repeatedly, and the new root can be quickly calculated layer by layer. Thus, the closer the data query frequencies of the data processing nodes where the local roots are located, the closer the probabilities of querying and verifying the data shards in the later stage. Therefore, when determining the node positions of each local root in the bottom-level leaf nodes of the global Merkle tree according to the magnitude order of the data query frequencies corresponding to the data processing nodes where the local roots are located, the local roots with stronger data query frequency correlation can be configured at adjacent bottom-level leaf nodes. In this way, further arranging each local root in sequence according to the node positions of each local root in the bottom-level leaf nodes of the global Merkle tree, in the bottom-level leaf nodes of the global Merkle tree, the correlation of data query operations between adjacent bottom-level leaf nodes is stronger, and the calculation amount for data query verification based on this global Merkle tree is less and the efficiency is higher.

[0083] It is understandable that in some cases, users need to perform data query verification on different data processing nodes simultaneously. For example, at the same moment, a user needs to query and verify some data shards stored in data processing node 1 and data processing node 3. Therefore, when arranging the node positions of the local roots at the bottom leaf nodes of the global Merkle tree, the data query frequency corresponding to the data processing node combination where the local root combination is located can also be referred to, and the node positions of the local root combinations at the bottom leaf nodes of the global Merkle tree are determined in the order of the magnitudes of the data query frequencies corresponding to the data processing node combinations where the local root combinations are located. Exemplarily, the data query frequency for data query verification simultaneously performed on data processing node 1 and data processing node 3 is 90 times per month, the data query frequency for data query verification simultaneously performed on data processing node 2 and data processing node 4 is 80 times per month, and the data query frequency for data query verification simultaneously performed on data processing node 5 and data processing node 6 is 20 times per month. Then, the node positions of each local root at the bottom leaf nodes of the global Merkle tree can be: node 1 - node 3 - node 2 - node 4 - node 5 - node 6. Among them, the two local roots corresponding to node 1 and node 3 can generate a hash value as an intermediate node of the global Merkle tree; the two local roots corresponding to node 2 and node 4 can generate a hash value as an intermediate node of the global Merkle tree; the two local roots corresponding to node 5 and node 6 can generate a hash value as an intermediate node of the global Merkle tree. In this way, in the order of the magnitudes of the data query frequencies corresponding to the local root combinations, each local root combination is arranged in sequence at the node positions of the bottom leaf nodes of the global Merkle tree. In the bottom leaf nodes of the global Merkle tree, the correlation between data query operations between adjacent bottom leaf nodes is relatively strong, and the computational amount for data query verification based on this global Merkle tree is less, and the efficiency is higher.

[0084] S250. Perform a digital signature on the global root of the global Merkle tree in the local TEE environment to obtain a global root signature.

[0085] S260. Write the global root signature, the target local root, and the timestamp information of the global root signature into the blockchain storage.

[0086] S270. Obtain the data shard to be updated. According to the association relationship between the data shard to be updated and each data processing node, screen the second target data processing nodes from each data processing node, and allocate the data shard to be updated to the second target data processing nodes.

[0087] S280. Dynamically update the global Merkle tree according to the updated local root of the updated local Merkle tree generated by the second target data processing node.

[0088] Among them, the data shard to be updated can be a data shard for incremental update. The second target data processing node is used to generate an updated local Merkle tree according to the locally generated local Merkle tree and the data shard to be updated. The updated local Merkle tree is a local Merkle tree dynamically updated according to the data shard to be updated. The updated local root can be the root of the updated local Merkle tree.

[0089] It can be understood that during the operation of the system, there will be cases of data incremental update. When incremental update is required, the central coordination node can first obtain the data shard to be updated, and determine the association relationship between the data shard to be updated and each data processing node. Then, according to the association relationship between the data shard to be updated and each data processing node, the second target data processing node is screened from each data processing node, and the data shard to be updated that needs incremental update is allocated to the second target data processing node. After receiving the data shard to be updated, the second target data processing node can generate an updated local Merkle tree according to the locally generated and stored local Merkle tree and the data shard to be updated. For example, the data shard to be updated is used as the newly added bottom leaf node of the local Merkle tree, and hash calculations are performed layer by layer upward to generate new intermediate nodes, and finally a new local Merkle tree is generated as the updated local Merkle tree. After the second target data processing node generates the updated local Merkle tree, it sends the updated local root of the updated local Merkle tree to the central coordination node. The central coordination node can determine the target bottom leaf node corresponding to the updated local root, and use the updated local root to replace the original local root of the target bottom leaf node, so as to dynamically update the path affected by the updated local root and generate an updated global Merkle tree.

[0090] In an optional embodiment of the present invention, the screening of the second target data processing node from each data processing node according to the association relationship between the data shard to be updated and each data processing node may include: determining the target data operation authorization object corresponding to the data shard to be updated; determining the data processing nodes whose data operation authorization objects corresponding to the services of each data processing node include the target data operation authorization object as the second target data processing nodes; determining the target storage type identifier corresponding to the data shard to be updated; determining the data processing nodes with the same node storage type identifier as the target storage type identifier as the second target data processing nodes; or, determining the preset query verification frequency of the data shard to be updated; determining the data processing nodes with a node query verification frequency close to the preset query verification frequency as the second target data processing nodes.

[0091] Among them, the target data operation authorization object can be the data operation authorization object to which the data shard to be updated belongs. The target storage type identifier can be the storage type identifier marked by the data shard to be updated. The preset query verification frequency can be the query verification frequency pre-configured by the target data operation authorization object for the data shard to be updated. It can be understood that the preset query verification frequency of the data shard to be updated may be different from the value of the actual query verification frequency in the later stage. The node query verification frequency can be the query verification frequency of the local data shard counted by the data processing node. The so-called query verification frequency can be understood as the frequency of querying and verifying data.

[0092] Specifically, the central coordination node can determine the association relationship between the data shard to be updated and each data processing node through a variety of different methods. Optionally, the central coordination node can determine the target data operation authorization object corresponding to the data shard to be updated, and determine the data processing nodes whose data operation authorization objects corresponding to the services of each data processing node include the target data operation authorization object as the second target data processing nodes. That is to say, the data shard to be updated can be assigned to the data processing node storing the data shard corresponding to the target data operation authorization object of the data shard to be updated for incremental update, so as to ensure that all data shards of the target data operation authorization object are stored in one data processing node or several centralized data processing nodes as much as possible.

[0093] Optionally, the central coordination node can also determine the target storage type identifier corresponding to the data shard to be updated, so as to determine the data processing nodes with the same node storage type identifier as the target storage type identifier as the second target data processing nodes. For example, if the target storage type identifier is cold storage, the data processing nodes using the cold storage method can be determined as the second target data processing nodes.

[0094] Optionally, the central coordination node can also determine the preset query verification frequency of the data shard to be updated, so as to determine the data processing nodes with the node query verification frequency close to the preset query verification frequency as the second target data processing nodes. For example, if the preset query verification frequency of the data shard to be updated is 100 times per month, the node query verification frequency of data processing node 1 is 90 times per month, and the node query verification frequency of data processing node 2 is 10 times per month, then data processing node 1 can be used as the second target data processing node for the data shard to be updated. That is, the data shard to be updated can be sent to data processing node 1, and data processing node 1 performs an incremental update operation on the data shard to be updated to generate an updated local Merkle tree.

[0095] In an alternative embodiment of the present invention, the above method may further include: determining the statistical query verification frequency of the current stored data shard; determining the updated data processing node of the current stored data shard according to the statistical query verification frequency of the current stored data shard; and migrating the current stored data shard from the current stored data processing node to the updated data processing node.

[0096] Wherein, the current stored data shard may be the data shards currently stored by each data processing node. The statistical query verification frequency may be the query verification frequency counted for the data shards stored by all data processing nodes. The updated data processing node may be the data processing node that stores the data shard after the data shard is migrated.

[0097] To further optimize the efficiency of data query verification based on the global Merkle tree, the central coordination node may regularly or irregularly count the statistical query verification frequency of the current stored data shard, and re-classify and store the current stored data shard according to the statistical query verification frequency of the current stored data shard.

[0098] Exemplarily, the central coordination node may re-classify the current stored data shard according to the statistical query verification frequency of the current stored data shard, specifically re-classifying the current stored data shard according to the segmented range of the statistical query verification frequency. For example, the current stored data shards with a statistical query verification frequency of [0 - 10) are classified into one category, the current stored data shards with a statistical query verification frequency (the unit can be set according to actual needs, such as times / month, etc.) of [10 - 50) are classified into one category, the current stored data shards with a statistical query verification frequency of [50 - 100) are classified into one category, and the current stored data shards with a statistical query verification frequency above 100 are classified into one category.

[0099] Furthermore, a new updated data processing node is allocated to each category of the current stored data shards with the statistical query verification frequency, so as to migrate the current stored data shard from the current stored data processing node to the updated data processing node.

[0100] Exemplarily, the current stored data shards with a statistical query verification frequency above 100 can be migrated from the current stored data processing node (which may be multiple data processing nodes) to the data processing node that stores the data shard in a hot storage manner; the current stored data shards with a statistical query verification frequency of [10 - 50) can be migrated from the current stored data processing node (which may be multiple data processing nodes) to the data processing node that stores the data shard in a cold storage manner.

[0101] To more clearly describe the technical solutions provided by the embodiments of the present invention, the embodiments of the present invention are illustrated by taking log data processing as a specific application scenario example. Traditional CA signatures are vulnerable to private key protection, signature latency, and full re-signing during data dynamic updates in the processing of large data log files. To solve problems such as private key leakage, single point of failure, and non-real-time verification, a secure and efficient incremental signature and verification scheme is urgently needed, and an immutable audit traceability mechanism is also provided.

[0102] In a specific example, the embodiments of the present invention propose a method for real-time signature and verification of incremental large data log files, which mainly includes the following steps:

[0103] (1), Data Sharding and Preprocessing

[0104] The log management system can divide large log files into data shards according to the dual criteria of data size (for example, every 10 MB) and timestamp (for example, every hour) through the central coordination node, and logically organize the data shards into subsets. Only the corresponding data shards are updated for the newly added log data, avoiding overall reconstruction.

[0105] Exemplarily, a 500 GB log file generated every day can be divided into approximately 50,000 basic data shards according to a size of 10 MB. At the same time, according to the timestamp information in the log records, every 50,000 shards are further grouped by hour (or other business rules) to form multiple subsets of data shards, and each subset can be sent to the corresponding data processing node for storage. When new log data is added to the log file, only local hash calculations and Merkle tree updates are performed on the newly added or affected data shards, avoiding reconstructing all the data. Before the data shards are transmitted, data compression and redundancy check are performed to ensure that the subsequent distributed processing nodes receive the data correctly.

[0106] (2), Internal Processing of Data Processing Nodes

[0107] The data processing node constructs a local Merkle tree using a secure hash algorithm in its built-in TEE environment and generates a local root. The data processing node further uses the secure multi-party computing method to sign the local root within the TEE to generate a local root signature. A timestamp can be embedded in the local root signature, and a replay prevention function is configured.

[0108] (3), Data Aggregation and Global Construction

[0109] Each data processing node uploads the local root and the local root signature to the central coordination node. The central coordination node groups the local roots according to preset rules, such as time series or geographical location, and constructs a second-level global Merkle tree. If 50 local roots are aggregated and then a 2-3 layer global Merkle tree is constructed, the global Merkle root consists of only a few consecutive hash calculations. At the same time, the central coordination node detects in real time whether there is an incremental data update. If there is an update, the local Merkle tree is recalculated for the affected part at the corresponding data processing node to achieve dynamic path update.

[0110] (4) Global Signature and Blockchain Audit

[0111] In the TEE environment of the central coordination node, the hardware security module and the CA private key are used to digitally sign the global Merkle root. Further, the central coordination node writes the global root signature of the global root, the key local roots, and the timestamp information of the global root into the blockchain. An immutable distributed audit record is formed on the blockchain.

[0112] (5) Real-time Verification Process

[0113] When a user requests to verify a certain log record, the log management system automatically retrieves the corresponding data shard and the dynamically stored Merkle path according to the record index. The log management system obtains the dynamic Merkle path information from the corresponding data shard. The data processing node where the log record is located recalculates the hash value of the data shard where the log record is located within the TEE. Further, the central coordination node uses the dynamic Merkle path to recalculate layer by layer until the global root to be verified is generated. The central coordination node uses the CA public key in the TEE environment to verify the legality of the global root signature, compares the global root to be verified with the stored global root, confirms the integrity of the data and the validity of the signature according to the comparison result, and returns the verification result to the user.

[0114] The data processing method provided by the above technical solution can achieve incremental data sharding and update, support dynamic append of local updates according to data shards without full-scale reconstruction. By constructing a distributed multi-level Merkle tree for data shards, it can efficiently utilize the parallel computing of multiple nodes and the hierarchical aggregation method to ensure data integrity. At the same time, the application of the distributed multi-level Merkle tree can achieve fast dynamic verification of data, that is, by pre-storing the dynamic Merkle path, fast positioning and verification of any data shard can be achieved.

[0115] It should be noted that the relevant information (including but not limited to user device information, log information, etc.) and data (including but not limited to data for display, data for analysis, etc.) involved in this disclosure are all information and data authorized by users or fully authorized by all parties, and the collection, use, and processing of relevant data comply with the relevant laws, regulations, and standards of the relevant regions.

[0116] It should be noted that any permutation and combination of the technical features in the above embodiments also fall within the protection scope of the present invention.

[0117] Embodiment III

[0118] Figure 4 is a schematic diagram of a data processing device provided in Embodiment III of the present invention. As Figure 4 shown, the device includes: a data sharding and sending module 310, a local root obtaining module 320, a global Merkle tree generating module 330, a global root signature generating module 340, and a data uploading and storing module 350, where:

[0119] The data sharding and sending module 310 is configured to shard the original data to obtain data shards, and send the data shards to data processing nodes;

[0120] The local root obtaining module 320 is configured to obtain the local roots of the local Merkle trees sent by the data processing nodes; wherein, the local Merkle trees are generated by the data processing nodes in the built-in trusted execution environment (TEE) according to the local sharded data;

[0121] The global Merkle tree generating module 330 is configured to use the local roots of the local Merkle trees as the bottom leaf nodes of the global Merkle tree to generate the global Merkle tree;

[0122] The global root signature generating module 340 is configured to digitally sign the global root of the global Merkle tree in the local TEE environment to obtain a global root signature;

[0123] The data uploading and storing module 350 is configured to write the global root signature, the target local root, and the timestamp information of the global root signature into the blockchain for storage.

[0124] In an embodiment of the present invention, the central coordination node slices the original data to obtain data slices, and sends the data slices to the data processing nodes. After each data processing node generates a local Merkle tree based on the local data slice, it sends the local root of the generated local Merkle tree to the central coordination node. After the central coordination node obtains the local roots of the local Merkle trees sent by each data processing node, it can use the local roots of each local Merkle tree as the bottom leaf nodes of the global Merkle tree, generate the global Merkle tree, and digitally sign the global root of the global Merkle tree in the local TEE environment to obtain the global root signature. Finally, the central coordination node writes the global root signature, the target local root, and the timestamp information of the global root signature into the blockchain storage. The above technical solution can solve the problems of long computing time, difficulty in meeting the dynamic update requirements, and low verification efficiency and security when processing data slices through the Merkle tree, can meet the dynamic update requirements of data slices, and improve the verification efficiency of data slices and the security of verification operations.

[0125] Optionally, the data slice sending module 310 is further configured to: obtain the slice attribute configuration information of the data slice; where the slice attribute configuration information includes at least one of the data operation authorization object of the data slice, the storage type identifier of the data slice, and the query verification frequency of the data slice; determine the first target data processing node corresponding to the data slice according to the slice attribute configuration information of the data slice; send the data slice to the first target data processing node.

[0126] Optionally, the above device further includes a data slice update module, configured to: obtain the data slice to be updated; screen the second target data processing nodes from each of the data processing nodes according to the association relationship between the data slice to be updated and each of the data processing nodes; allocate the data slice to be updated to the second target data processing nodes; where the second target data processing nodes are used to generate an updated local Merkle tree based on the locally generated local Merkle tree and the data slice to be updated; dynamically update the global Merkle tree according to the updated local root of the updated local Merkle tree generated by the second target data processing nodes.

[0127] Optionally, the data shard update module is further configured to: determine a target data operation authorization object corresponding to the data shard to be updated; determine, as the second target data processing nodes, the data processing nodes whose data operation authorization objects corresponding to the services of each of the data processing nodes include the target data operation authorization object; determine a target storage type identifier corresponding to the data shard to be updated; determine, as the second target data processing nodes, the data processing nodes whose node storage type identifiers are the same as the target storage type identifier; or, determine a preset query verification frequency of the data shard to be updated; determine, as the second target data processing nodes, the data processing nodes whose node query verification frequencies are close to the preset query verification frequency.

[0128] Optionally, the above device further includes a data shard migration module, configured to: determine a statistical query verification frequency of the currently stored data shard; determine, according to the statistical query verification frequency of the currently stored data shard, an updated data processing node for the currently stored data shard; and migrate the currently stored data shard from the currently storing data processing node to the updated data processing node.

[0129] Optionally, the global Merkle tree generation module 330 is further configured to: determine, according to a target reference factor, a target node arrangement rule for each of the local roots as the bottom-layer leaf nodes of the global Merkle tree; and arrange each of the local roots in sequence at the node positions of the bottom-layer leaf nodes of the global Merkle tree according to the target node arrangement rule, to obtain the bottom-layer leaf nodes of the global Merkle tree.

[0130] The above data processing device can execute the data processing method provided in any embodiment of the present invention, and has corresponding functional modules and beneficial effects for executing the method. For technical details not described in detail in this embodiment, reference can be made to the data processing method provided in any embodiment of the present invention.

[0131] Since the above-introduced data processing device is a device that can execute the data processing method in the embodiments of the present invention, based on the data processing method introduced in the embodiments of the present invention, those skilled in the art can understand the specific implementation manners and various variations of the data processing device in this embodiment. Therefore, the implementation of how the data processing device implements the data processing method in the embodiments of the present invention will not be described in detail herein. As long as the device adopted by those skilled in the art to implement the data processing method in the embodiments of the present invention falls within the scope of protection of this application.

[0132] Embodiment 4

[0133] Figure 5The structural schematic diagram of an electronic device 10 that can be used to implement the embodiments of the present invention is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital processors, cellular phones, smart phones, wearable devices (such as helmets, glasses, watches, etc.) and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present invention described and / or claimed herein.

[0134] As Figure 5 shown, the electronic device 10 includes at least one processor 11, and a memory communicatively connected to the at least one processor 11, such as a read-only memory (ROM) 12, a random access memory (RAM) 13, etc. The memory stores a computer program executable by the at least one processor. The processor 11 can perform various appropriate actions and processes according to the computer program stored in the read-only memory (ROM) 12 or the computer program loaded from the storage unit 18 into the random access memory (RAM) 13. In the RAM 13, various programs and data required for the operation of the electronic device 10 can also be stored. The processor 11, the ROM 12, and the RAM 13 are connected to each other through a bus 14. The input / output (I / O) interface 15 is also connected to the bus 14.

[0135] A plurality of components in the electronic device 10 are connected to the I / O interface 15, including: an input unit 16, such as a keyboard, a mouse, etc.; an output unit 17, such as various types of displays, speakers, etc.; a storage unit 18, such as a magnetic disk, an optical disc, etc.; and a communication unit 19, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 19 allows the electronic device 10 to exchange information / data with other devices through a computer network such as the Internet and / or various telecommunication networks.

[0136] The processor 11 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the processor 11 include but are not limited to a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The processor 11 executes the various methods and processes described above, such as the data processing method.

[0137] Optionally, the data processing method is applied to the central coordination node, and may include: sharding the original data to obtain data shards, and sending the data shards to the data processing nodes; obtaining the local roots of the local Merkle trees sent by each of the data processing nodes; wherein, the local Merkle trees are generated by the data processing nodes according to the local shard data in the built-in trusted execution environment (TEE) environment; using the local roots of the local Merkle trees as the bottom leaf nodes of the global Merkle tree to generate the global Merkle tree; digitally signing the global root of the global Merkle tree in the local TEE environment to obtain a global root signature; writing the global root signature, the target local root, and the timestamp information of the global root signature into the blockchain storage.

[0138] In some embodiments, the data processing method may be implemented as a computer program tangibly embodied in a computer-readable storage medium, such as storage unit 18. In some embodiments, part or all of the computer program may be loaded and / or installed onto the electronic device 10 via ROM 12 and / or communication unit 19. When the computer program is loaded into RAM 13 and executed by the processor 11, one or more steps of the data processing method described above may be performed. Alternatively, in other embodiments, the processor 11 may be configured to execute the data processing method by any other suitable means (e.g., by means of firmware).

[0139] The various embodiments of the systems and techniques described above in this document can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-chip (SOCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include: being implemented in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which can be a dedicated or general-purpose programmable processor, receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting the data and instructions to the storage system, the at least one input device, and the at least one output device.

[0140] A computer program for implementing the method of the present invention can be written in any combination of one or more programming languages. These computer programs can be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that when executed by the processor, the computer programs cause the functions / operations specified in the flowchart and / or block diagram to be implemented. The computer programs can be executed entirely on the machine, partially on the machine, as a stand-alone software package partially on the machine and partially on a remote machine, or entirely on a remote machine or server.

[0141] In the context of the present invention, a computer-readable storage medium can be a tangible medium that can contain or store a computer program for use by or in connection with an instruction execution system, apparatus, or device. The computer-readable storage medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. Alternatively, the computer-readable storage medium can be a machine-readable signal medium. More specific examples of the machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0142] In order to provide interaction with a user, the systems and techniques described herein can be implemented on an electronic device having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and a pointing device (e.g., a mouse or a trackball) through which the user can provide input to the electronic device. Other kinds of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).

[0143] The systems and techniques described herein can be implemented in a computing system including backend components (e.g., as a data server), or a computing system including middleware components (e.g., an application server), or a computing system including frontend components (e.g., a user computer having a graphical user interface or a web browser through which a user can interact with an implementation of the systems and techniques described herein), or a computing system including any combination of such backend components, middleware components, or frontend components. The components of the system can be interconnected to each other by digital data communication in any form or medium (e.g., a communication network). Examples of communication networks include: local area network (LAN), wide area network (WAN), blockchain network, and the Internet.

[0144] A computing system can include a client and a server. The client and the server are generally far from each other and typically interact through a communication network. The client-server relationship is created by computer programs running on respective computers and having a client-server relationship with each other. The server can be a cloud server, also known as a cloud computing server or a cloud host, which is a host product in the cloud computing service system, and solves the defects of difficult management and weak business scalability existing in traditional physical hosts and VPS services.

[0145] It should be understood that various forms of the processes shown above can be used, with steps reordered, added, or deleted. For example, the steps recited in this disclosure can be executed in parallel, sequentially, or in a different order, as long as the desired results of the technical solutions disclosed in this disclosure can be achieved, and this is not limited herein.

[0146] The above specific embodiments do not constitute a limitation on the protection scope of this disclosure. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principle of this disclosure shall be included within the protection scope of this disclosure.

Claims

1. A data processing method, characterized in that: Applied to the central coordination node, including: Slice the original data to obtain data slices, and send the data slices to the data processing nodes; Obtaining a local root of a local Merkle tree sent by each of the data processing nodes; wherein the local Merkle tree is generated by the data processing node in a built-in trusted execution environment TEE environment according to local shard data; Using the local roots of the local Merkle trees as bottom leaf nodes of the global Merkle tree to generate the global Merkle tree; Digitally sign the global root of the global Merkle tree in the local TEE environment to obtain a global root signature; The global root signature, the target local root, and the timestamp information of the global root signature are written into blockchain storage.

2. The method according to claim 1, characterized in that The sending of the data slices to the data processing node includes: Acquire shard attribute configuration information of the data shard; wherein the shard attribute configuration information includes at least one of a data operation authorization object of the data shard, a storage type identifier of the data shard, and a query verification frequency of the data shard; Determine a first target data processing node corresponding to the data shard according to the shard attribute configuration information of the data shard; The data slice is sent to the first target data processing node.

3. The method according to claim 1, characterized in that After generating the global Merkle tree, the method further includes: Get the data shard to be updated; According to the association relationship between the data slice to be updated and each of the data processing nodes, selecting a second target data processing node from each of the data processing nodes; Allocating the to-be-updated data slice to the second target data processing node; The second target data processing node is used to generate and update the local Merkle tree according to the locally generated local Merkle tree and the data shard to be updated; The global Merkle tree is dynamically updated according to the updated local root of the updated local Merkle tree generated by the second target data processing node.

4. The method according to claim 3, characterized in that The step of selecting a second target data processing node from each of the data processing nodes according to the association relationship between the data slice to be updated and each of the data processing nodes includes: Determine the target data operation authorization object corresponding to the data slice to be updated; Determine the data processing node whose data operation authorization object corresponding to each of the data processing nodes includes the target data operation authorization object as the second target data processing node; Determine the target storage type identifier corresponding to the data slice to be updated; determining a data processing node with the same node storage type identifier as the target storage type identifier as the second target data processing node; or Determine a preset query verification frequency for the data shard to be updated; A data processing node whose node query verification frequency is close to the preset query verification frequency is determined as the second target data processing node.

5. The method according to claim 1, characterized in that The method further comprises: Determine the statistical query verification frequency of the current storage data shard; Determine an update data processing node for the current storage data shard according to the statistical query verification frequency of the current storage data shard; The currently stored data slice is migrated from the currently stored data processing node to the updated data processing node.

6. The method according to claim 1, characterized in that The method of using the local roots of the local Merkle trees as bottom leaf nodes of the global Merkle tree includes: Determine, according to the target reference factor, a target node arrangement rule for each of the local roots as a bottom leaf node of the global Merkle tree; The local roots are arranged in sequence at the node positions of the bottom leaf nodes of the global Merkle tree according to the target node arrangement rule to obtain the bottom leaf nodes of the global Merkle tree.

7. A data processing device, characterized in that: Configured on the central coordination node, including: A data sharding sending module is used to slice the original data into data shards and send the data shards to the data processing nodes; A local root acquisition module, used to acquire the local root of the local Merkle tree sent by each of the data processing nodes; wherein the local Merkle tree is generated by the data processing node according to the local shard data in the built-in trusted execution environment TEE environment; A global Merkle tree generation module, used to use the local roots of the local Merkle trees as bottom leaf nodes of the global Merkle tree to generate the global Merkle tree; A global root signature generation module, used to digitally sign the global root of the global Merkle tree in a local TEE environment to obtain a global root signature; The data on-chain storage module is used to write the global root signature, the target local root and the timestamp information of the global root signature into the blockchain storage.

8. An electronic device, characterized in that: The electronic device comprises: at least one processor; and a memory communicatively connected to the at least one processor; wherein, The memory stores a computer program executed by the at least one processor, and the computer program is executed by the at least one processor so that the at least one processor can execute the data processing method according to any one of claims 1 to 6.

9. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores computer instructions, and the computer instructions are used to enable a processor to implement the data processing method described in any one of claims 1 to 6 when executed.

10. A computer program product comprising a computer program / instructions, wherein: When the computer program / instructions are executed by a processor, the data processing method described in any one of claims 1 to 6 is implemented.

Citation Information

Cited By

  • Big data-based property right transaction data tamper-proof storage system

    CN121302439A