Management method of state data in block chain system and block chain node
By using a tree structure in the blockchain system to organize state data and record block numbers and child node hash values, the problem of excessive storage volume caused by the increase in state data version is solved, and an efficient data management and deletion mechanism is realized.
Patent Information
- Application Number
- CN202510394618.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-28
- Publication Date
- 2025-07-11
AI Technical Summary
With the increase in the version of state data in blockchain systems, it is difficult for the existing technology to efficiently identify and delete earlier versions of state data, resulting in a significant increase in the amount of data storage and requires continuous data governance.
By using a tree structure to organize state data in the blockchain system, the tree node contains the block number and the hash value of the child node, record the execution write set of the transaction sequence when generating/updating the tree node, and write a key-value pair in the data storage system. The key includes the block number and node identification, and the value is determined based on the data content.
It realizes efficient management of state data, can accurately identify and delete earlier versions of state data, reduce data storage amount, and improve the efficiency of the data storage system.
Smart Images

Figure CN120296015A_ABST
Abstract
Description
Technical Field
[0001] The embodiments of this specification belong to the field of computer technology, and particularly relate to a method for managing state data in a blockchain system and a blockchain node. Background Art
[0002] A blockchain system is a new application mode of computer technologies such as distributed data storage, peer-to-peer transmission, consensus mechanism, and encryption algorithms. In a blockchain system, data blocks are combined into a chain data structure in a sequential connection manner according to the time sequence, and a distributed ledger that is immutable and unforgeable is guaranteed by cryptographic means. Due to the characteristics of decentralization, information immutability, autonomy, etc. of the blockchain system, the blockchain system has received more and more attention and applications.
[0003] In a blockchain system, state data can be managed through a tree structure, and the aforementioned tree structure can include, but is not limited to, MPT (Merkle Patricia Tree) and SMT (Sparse Merkle Tree), etc. A leaf node in this tree structure stores the value of a state variable, and part of the information in the directed path from the root node to the leaf node of this tree structure constitutes the key of the state variable. The key-value pairs of the tree nodes themselves in this tree structure can be stored in a data storage system; based on the tree structures corresponding to different versions of state data, different state root (State_Root) hashes can be calculated, and the state root hash can be used as an identifier of the corresponding version of state data and anchored to the corresponding block header.
[0004] Based on the characteristics of the aforementioned tree structure, in a blockchain system, it is possible to achieve storing multiple versions of state data corresponding to multiple generated blocks in an append-only form with a relatively small amount of data storage. However, with a large increase in the number of state data versions, it will still cause a significant increase in the amount of state data stored in the blockchain system, and it is necessary to continuously prune / data govern the state data stored in the data storage system to delete earlier versions of state data. Summary of the Invention
[0005] The purpose of the present invention is to provide a method for managing state data in a blockchain system and a blockchain node.
[0006] In a first aspect, a method for managing state data in a blockchain system is provided. In the blockchain system, state data is organized in a tree structure. The value of a state variable is stored in a leaf node of the tree structure, and the key of a state variable is stored in the directed path from the root node to a leaf node of the tree structure. The tree nodes in the tree structure include the block number corresponding to when the tree node is created / updated and the hash values of the child nodes of the tree node. The method includes: generating / updating a first tree node according to the execution write set of the transaction sequence included in the target block to be generated; writing the key-value pairs of the plurality of first tree nodes into a data storage system. The key of the first tree node includes the block number of the target block and the node identifier of the first tree node. The node identifier of the first tree node includes a component of the key of the state variable stored in the directed path from the root node to the first tree node of the tree structure. The value of the first tree node is determined based on the data content of the first tree node.
[0007] In a second aspect, a blockchain node in a blockchain system is provided. In the blockchain system, state data is organized in a tree structure. The value of a state variable is stored in a leaf node of the tree structure, and the key of a state variable is stored in the directed path from the root node to a leaf node of the tree structure. The tree nodes in the tree structure include the block number corresponding to when the tree node is created / updated and the hash values of the child nodes of the tree node. The blockchain node includes: a node management unit configured to generate / updating a first tree node according to the execution write set of the transaction sequence included in the target block to be generated; a storage processing unit configured to write the key-value pairs of the first tree node into a data storage system. The key of the first tree node includes the block number of the target block and the node identifier of the first tree node. The node identifier of the first tree node is a component of the key of the state variable stored in the directed path from the root node to the first tree node. The value of the first tree node is determined based on the data content of the first tree node.
[0008] In a third aspect, a computing device is provided, including a memory and a processor. A computer program is stored in the memory. When the processor executes the computer program, the method described in the first aspect is implemented.
[0009] In a fourth aspect, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed in a computing device, the computing device executes the method described in the first aspect.
[0010] In the technical solution provided by the embodiments of this specification: In a data storage system that needs to store status data with multiple versions, when deleting the earlier version of the status data, the block number included in the key of the key-value pair stored in the data storage system can be used to efficiently and accurately identify which transaction sequence in which block the tree node corresponding to the key-value pair is generated, so as to more efficiently decide whether to delete the key-value pair. Brief Description of the Drawings
[0011] In order to more clearly illustrate the technical solutions in the embodiments of this specification, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the drawings in the following description are only some embodiments recorded in this specification. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0012] Figure 1 It is an architecture diagram of a blockchain system exemplarily provided in the embodiments of this specification;
[0013] Figure 2 It is one of the flowcharts of a method for managing status data in a blockchain system provided in the embodiments of this specification;
[0014] Figure 3 It is one of the schematic diagrams of a tree structure exemplarily provided in the embodiments of this specification;
[0015] Figure 4 It is another schematic diagram of a tree structure exemplarily provided in the embodiments of this specification;
[0016] Figure 5 It is one of the flowcharts of a method for managing status data in a blockchain system provided in the embodiments of this specification;
[0017] Figure 6 It is a schematic diagram of the structure of a blockchain node provided in the embodiments of this specification. Detailed Embodiments
[0018] In order to enable those skilled in the art to better understand the technical solutions in this specification, the following will clearly and completely describe the technical solutions in the embodiments of this specification with reference to the drawings. Obviously, the described embodiments are only some embodiments of this specification, rather than all embodiments. Based on the embodiments in this specification, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of this specification.
[0019] Figure 1This is an architecture diagram of a blockchain system provided exemplarily in the embodiments of this specification. The blockchain system may include N blockchain nodes, where Figure 1 8 blockchain nodes such as Node 1 - Node 8 are exemplarily shown. The connections between the nodes schematically represent the connections between the nodes, and the aforementioned connections are used to support data transmission between different nodes.
[0020] The blockchain system can provide the function of smart contracts. The smart contracts in the blockchain system are contracts that can be triggered and executed by transactions. Smart contracts can be defined in the form of contract codes. Invoking a smart contract in the blockchain system is to initiate a transaction pointing to the contract address of the smart contract, so that each node in the blockchain system runs the corresponding contract code distributively.
[0021] In various blockchain systems introducing smart contracts, accounts can generally be divided into two types:
[0022] Contract account (CA): mainly used to store the contract code of the corresponding smart contract and the values of the state variables defined in the smart contract, and usually can only be activated by calling from an external account;
[0023] Externally owned account (EOA): an account registered by an external user in the blockchain system.
[0024] The design of external accounts and contract accounts is actually a mapping from account addresses to account states. The account state of any account usually includes fields such as nonce, balance, storageRoot, and codeHash. Among them, nonce and balance exist in both external accounts and contract accounts, and the codeHash and storageRoot attributes are generally only valid for contract accounts.
[0025] More specifically, for an external account, the value of nonce represents the number of transactions sent from the relevant account address; for a contract account, the value of nonce can represent the number of smart contracts created by the relevant account address. The value of balance represents the number of a certain digital resource / token owned by the relevant account address. The value of storage root is the hash value of the root node of a tree structure, such as an MPT tree, which is used to organize / manage the storage of the state variables of the relevant contract account. The value of codeHash represents the hash value of the contract code of the relevant smart contract. For an external account, since it does not include a smart contract, the values of the storageRoot and CodeHash fields can generally be an empty string / all 0 strings.
[0026] It should be noted that MPT stands for Merkle Patricia Tree, which is a tree structure that combines Merkle Tree and Patricia Tree (a more space - saving trie). Among them, the Merkle tree algorithm can calculate a Hash value for multiple transactions respectively, and then connect them pairwise and calculate the Hash again until the top - level Merkle root. In some blockchain systems, an improved MPT tree is usually adopted, such as a 16 - ary tree structure, which is usually simply referred to as the MPT tree.
[0027] The system data that needs to be persistently stored in the blockchain system can be divided into two parts: block data and state data.
[0028] Block data includes one or more blocks in ascending order of block height (or block number). A single block can include a block header and a block body. The block header can include the block hash previous_Hash (or called parent hash) of the previous block, timestamp Timestamp, block number BlockNum, state root hash State_Root, transaction root hash Transaction_Root, receipt root hash Receipt_Root, and nonce, etc. The block body can include a transaction set and a receipt set.
[0029] A transaction in the blockchain system refers to a task unit that is executed and recorded in the blockchain system. A single transaction usually includes a sending field (From), a receiving field (To), and a data field (Data). Among them, the From field includes the account that initiates the transaction (i.e., the sender account), and the To field may include another account involved / pointed to by the transaction.
[0030] For any Nth block, according to the state data with block number (or version number) N - 1, multiple transactions included in the transaction set belonging to the Nth block can be executed in sequence to obtain the execution write set of the multiple transactions. Furthermore, the state data with version number N - 1 can be updated according to the execution write set of the multiple transactions to obtain the state data with version number N.
[0031] Generally, a tree structure can be used to organize the state data of a blockchain system. The aforementioned tree structure can include, but is not limited to, MPT (Merkle Patricia Tree) and SMT (Sparse Merkle Tree), etc. A leaf node in the tree structure stores the value of a state variable, and in the directed path from the root node to a leaf node, a key of a state variable (the state variable described here can be the account address of an external account / contract account, or can also be a state variable defined in a smart contract) is stored. The data storage system can store the key-value pairs of the tree nodes in the tree structure. The key of a tree node is usually a hash value calculated from the data content of the tree node; it should be noted that for the state variables in a smart contract, the key of the corresponding tree node can be determined based on the contract address of the smart contract and the tree node itself. For example, it is a combination of the contract address of the smart contract and the hash value calculated from the data content of the tree node.
[0032] The tree structure used to organize state data can include a state trie. The hash value of the root node of the state trie is stored in State_Root in the block header. The account state of an external account / contract account is stored in a leaf node of the state trie. In the directed path from the root node to a leaf node of the state trie, the account address of an external account / contract account, or part or all of the hash value calculated based on the account address, is stored. As mentioned above, the account state of a single account usually can include fields such as Nonce, Balance, Storage root, CodeHash, etc. Nonce and Balance exist in both external accounts and contract accounts, and CodeHash and Storage root are generally only valid for contract accounts.
[0033] The tree structure used to organize state data can also include a storage trie. The hash value of the root node of the storage trie is stored in the storageroot field of the contract account corresponding to the relevant smart contract, thereby locking the contract state of the smart contract to the relevant contract account through the hash value. The value of a state variable defined in the smart contract is stored in a leaf node of the storage trie. In the directed path from the root node to a leaf node of the storage trie, the state key of a state variable defined in the relevant smart contract is stored, that is, the partial information in the directed path from the root node to the leaf node of the storage trie can be arranged in order to form the key of a state variable defined in all the relevant smart contracts.
[0034] In a data storage system, an LSM-based NoSQL Key-Value DB (DataBase; Key-Value DB is also simply referred to as KVDB) is usually adopted to store the key-value pairs of tree nodes. The data storage system can be, for example, levelDB of Ethereum or RocksDB of Libra. Both of these KVDBs are based on the LSM storage engine. The LSM storage engine is a hierarchical, ordered, disk-oriented storage engine. It draws on the characteristic of continuous appending of Logs. Its core idea is to make full use of the fact that disk bulk sequential writes are far more efficient than random writes, and sacrifice some read efficiency in exchange for maximizing write operation efficiency.
[0035] The key-value pairs of tree nodes can be first stored in memory, specifically in the MemTable in memory. When the data volume in the MemTable reaches a certain threshold, such as 256MB, the data in the MemTable can be written (flushed) to disk. Among them, to avoid the write operation of the MemTable from blocking the flush, this MemTable can be converted into an immutable Immutable Memtable, that is, the Immutable Memtable is set to read-only, and a new MemTable is generated to receive newly incoming key-value pairs of tree nodes; then the data in the Immutable MemTable is written to disk.
[0036] The key-value pairs of tree nodes can be stored in multiple levels of SST files on disk. After the storage capacity of the upper level (such as Level 0) reaches or approaches the upper limit, a process called "compaction" can be used to merge the SST files of Level 0 into the next level, such as Level 1. In this way, the key-value pairs of tree nodes stored in the SST files at higher levels are more updated than those stored in the SST files at lower levels. When searching for the key-value pair of a tree node corresponding to a certain state variable, it is necessary to start from the SST file at the highest level and gradually search downward to lower-level STT files until the required key-value pair of the tree node is found. Among them, all key-value pairs located in the same SST file are usually sorted in order according to the value of the key or a certain data included in the key.
[0037] With a large increase in the number of state data versions, there will still be a significant increase in the amount of state data, and it is necessary to continuously prune / govern the state data stored in the data storage system to delete the earlier versions of the state data. However, as mentioned above, in the key-value pairs stored in the tree nodes in the data storage system, the key of the tree node is usually the hash value calculated from the data content of the tree node, or the combination of the hash value calculated from the data content of the tree node and the contract address of the relevant smart contract. In this case, for any key-value pair of a tree node stored in the data storage system, it is difficult to efficiently distinguish which transaction sequence in which block the tree node corresponding to the key-value was generated / updated, resulting in difficulty in efficiently determining whether the key-value pair needs to be deleted when it is necessary to delete the earlier versions of the state data.
[0038] In the embodiments of this specification, at least one method for managing state data in a blockchain system and a blockchain node are provided. In the blockchain system, state data is organized in a tree structure. The value of a state variable is stored in a leaf node of the tree structure, and the key of a state variable is stored in the directed path from the root node to a leaf node of the tree structure. The tree nodes in the tree structure include the block number corresponding to when the tree node is created / updated and the hash values of the child nodes of the tree node. After generating / updating the first tree node by completing the execution write set of the transaction sequence included in the target block to be generated, the key-value pair of the first tree node can be written into the data storage system. The key of the first tree node includes the block number of the target block and the node identifier of the first tree node. The node identifier of the first tree node is a component of the key of the state variable stored in the directed path from the root node to the first tree node of the tree structure, and the value of the first tree node is determined based on the data content of the first tree node.
[0039] In this way, when it is necessary to delete the earlier versions of the state data, the block number included in the key in the key-value pair stored in the data storage system can be used to efficiently and accurately identify which transaction sequence in which block the tree node corresponding to the key-value was generated, so as to more efficiently decide whether the key-value pair needs to be deleted.
[0040] Figure 2 It is one of the flowcharts of a method for managing state data in a blockchain system provided in the embodiments of this specification.
[0041] In this blockchain system, the state data is organized in a tree structure. The value of a state variable is stored in a leaf node of the tree structure, and the key of a state variable is stored in the directed path from the root node to a leaf node of the tree structure. The tree nodes in the tree structure include the block number corresponding to when the tree node is created / updated and the hash values of the child nodes of the tree node.
[0042] Exemplarily, referring to Figure 3 as shown. The tree structure for organizing the state data corresponding to block N-1 may include a State trie and Storage tries corresponding to one or more smart contracts. Figure 3 The exemplary State trie and Storage tries include multiple tree nodes such as tree nodes A1 to A14. A6 to A9 are leaf nodes of the State trie, and A11, A13, and A14 are leaf nodes of the Storage trie. Among them, the leaf nodes may include multiple fields such as Num, Key-end, and value, and each intermediate node that is not a leaf node may include multiple fields such as Num, Shared, and value.
[0043] The Num field of the tree node may store the block number corresponding to when the tree node is generated / updated. For example, during the generation of block N-3, tree nodes A10 to A14 are updated, and tree nodes A10 to A14 are not updated again during the generation of block N-2 and block N-1. Then, the values of the Num fields in tree nodes A10 to A14 are all N-3; for another example, during the generation of block N-2, tree nodes A4 and A6 are updated, and tree nodes A4 and A6 are not updated again during the generation of block N-1. Then, the values of the Num fields in tree nodes A4 and A6 are all N-2; for another example, during the generation of block N-1, tree nodes A1 to A3, A5, A7 to A9 are updated. Then, the values of the Num fields in tree nodes A1 to A3, A5, A7 to A9 are all N-1.
[0044] In the directed path from the root node to a leaf node, the values of the Shared field and the Key-end field are concatenated in sequence to form the key of a state variable. As mentioned before, the state variables described here can be state variables in external / contract accounts or smart contracts. For example, in the directed path from the root node A1 of the State trie to the leaf node A6 of the State trie, the values of the Shared field and the Key-end field are concatenated in sequence to form the account address a7b6c5d5 of an external account; similarly, in the directed path from A1 to A7, the account address a7b6c6d6 of an external account is stored, in the directed path from A1 to A8, the account address a7b7c7d8 of a contract account is stored, and in the directed path from A1 to A9, the account address a7b7c8d9 of an external account is stored. Another example is that in the directed path from the root node A10 of the Storage trie to the leaf node A11 of the State trie, the values of the Shared field and the Key-end field are concatenated in sequence to form the key of a state variable defined in the smart contract corresponding to the Storage trie, that is, e51235a6; similarly, in the directed path from A10 to A13, the key of a state variable, that is, e5367a62, is stored, and in the directed path from A10 to A14, the key of a state variable, that is, e53683e1, is stored.
[0045] The value field in the leaf node can store the value of the state variable corresponding to the leaf node. For example, the value fields in the leaf nodes A6 to A9 of the State trie can store the account status of the relevant external / contract accounts; another example is that the value fields in the leaf nodes A11 to A14 of the Storage trie can store the variables of their respective corresponding state variables.
[0046] The value field in the intermediate node can store the hash values of the child nodes. For example, the value field of A12 can store the hash values of A13 and A14; another example is that the value field of A1 can store the hash values of A2 and A3.
[0047] The tree structure in the foregoing example is exemplary. For example, the Key-end field in the leaf node may store the complete key of the state variable corresponding to the leaf node to which it belongs. For example, the Key-end field of the leaf node A6 may store a7b6c5d5.
[0048] It can be understood that the hash value of the root node of the State trie can be anchored in the block header. For example, the hash value of node A1 can be stored under the StateRoot field included in the block header Block N-1Header of block N-1. It can be understood that other information can also be included in Block N-1Header, such as the aforementioned Transaction Root and Reciept Root fields, etc.
[0049] This method can be executed by a blockchain node in a blockchain system. This method exemplarily describes the process of how a blockchain node deposits the state data corresponding to a target block into a data storage system during the generation of the target block.
[0050] Refer to Figure 2 As shown, this method can include but is not limited to the following steps S201 and step S203.
[0051] Step S201, generate / update a first tree node according to the execution write set of the transaction sequence included in the target block to be generated.
[0052] Based on the execution write set of the transaction sequence included in the target block, one or more target state variables to be updated and their respective variable values in the state data of the latest version can be determined. Thus, one or more first tree nodes to be generated or updated can be determined according to each target state variable. Exemplarily, continuing the Figure 2 example above, here it is assumed that based on the execution write set of the transaction sequence included in block N, the target state variables to be updated are determined to include the state variable with e53683e1 as the key in the smart contract corresponding to the Storage trie and the external account with a7b7c8d9 as the account address. In addition, the variable value of e53683e1 in the state data corresponding to block N is V1, and the variable value of a7b7c8d9 in the state data corresponding to block N is V2. Then, refer to Figure 4As shown, all the tree nodes in the directed paths from the root node A1 to the leaf nodes A14 and A9 can be determined as the first tree nodes to be updated, and corresponding updates can be performed on them. For example, the values of the Num and value fields in A14 can be updated to N and V1 respectively, and the values of the Num and value fields in A9 can be updated to N and V2 respectively; the value of the Num field in A1, A3, A5, A8, A9, A10, A12 located in the directed path can be updated to the block number N, and the hash value of A14 can be updated in the value field of A12, the hash value of A12 can be updated in the value field of A10, the hash value of A10 can be updated in the value field of A8, the hash value of A8 can be updated in the value field of A5, the hash values of A5 and A9 can be updated in the value field of A3, and the hash value of A3 can be updated in the value field of A1.
[0053] Step S203: Write the key-value pair of the first tree node into the data storage system. The key of the first tree node includes the block number of the target block and the node identifier of the first tree node. The node identifier of the first tree node is a component of the key of the state variables stored in the directed path from the root node of the tree structure to the first tree node. The value of the first tree node is determined based on the data content of the first tree node. For example, the value of the first tree node is the value of the value field in the first tree node.
[0054] There may also be other data stored in the data storage system besides the key-value pairs of the tree nodes. In this case, in order to distinguish it from other data, the key of the first tree node may also include a preset prefix, and the preset prefix is used to indicate that the data content corresponding to the key-value pair belongs to the state data of the blockchain system.
[0055] In a data storage system, an LSM tree is usually adopted to manage key-value pairs of tree nodes, that is, key-value pairs of tree nodes are stored through multiple levels of SST files. The key-value pairs of tree nodes stored in the SST file with a higher level are more updated than those stored in the SST file with a lower level. When it is necessary to find the key-value pair of the tree node corresponding to a certain state variable, it is necessary to start from the SST file at the highest level and gradually search for the STT file at a lower level until the key-value pair of the required tree node is found; moreover, for all key-value pairs located in the same SST file, they are usually sorted in order according to the value of the key or a certain data included in the key, for example, sorted in ascending order according to a certain data. Usually, the higher-level tree nodes in the tree structure are accessed (including querying and updating) more frequently; therefore, in order to make the key-value pairs of the higher-level tree nodes be located at a relatively higher position as much as possible, and at the same time, make the key-value pairs of the higher-level tree nodes be located at a relatively front position in the SST file, so that the blockchain node can access the higher-level tree nodes more efficiently, the key of the first tree node can also include level information, and the level information is used to indicate the level to which the corresponding tree node belongs in the tree structure.
[0056] A tree structure including a State trie and a Storage trie usually can have a predetermined number (for example, K + 1) of levels. The root node of the State trie is located at a high level (for example, the 0th level), and the leaf node of the Storage trie is located at the lowest level (for example, the K + 1th level); for intermediate nodes other than leaf nodes, they can be determined according to the byte length of the node identifier of the tree node. For example, when the byte length of the shared part in a single tree node in both the State trie and the Storage trie is M bytes, if the byte length of the node identifier of an intermediate node in the State trie is X * M, then the level where this intermediate node is located is X - 1; if the byte length of the node identifier of an intermediate node in the Storage trie is X * M, the level of this intermediate node is X * M - 1 + (K + 1) / 2.
[0057] When the LSM is adopted in the data storage system to manage the key-value pairs of tree nodes, the block number included in the key of the first tree node can be stored in the data storage system in a bitwise inversion manner, so that the key-value pairs of multiple tree nodes with different block numbers but the same node identifier can be continuously stored in the same SST in the order from the largest to the smallest block number, so that when it is necessary to delete the state data of certain versions, the key-value pairs of multiple tree nodes with the same node identifier can be efficiently traversed from the same SST in the order from the largest to the smallest block number.
[0058] It can be understood that when the LSM is adopted in the data storage system to manage the key-value pairs of tree nodes, specifically, the key-value pairs of the first tree node can be written into the SST file belonging to the highest level.
[0059] Continuing with the foregoing Figure 3 and Figure 4 In the example, the key-value pairs shown in Table 1 below can be stored in the data storage system.
[0060]
[0061]
[0062] Table 1
[0063] In Table 1 of the foregoing example, St is used to represent the preset prefix. For the tree nodes in the Storage trie, the node identifier of the tree node usually can include the prefix information determined according to the contract address of the smart contract corresponding to the Storage trie, and this prefix information usually can be the contract address. For the Value of the tree node, it usually can be the data content stored in the value field in the tree node. In the foregoing Table 1, V() is used to represent the data content stored in the value field of the tree node in the parentheses.
[0064] When the LSM is adopted in the data storage system to manage the key-value pairs of tree nodes, the data storage system may, through periodicity or in response to a certain trigger condition, implement the merging of the SST files of the current level into the SST files of the next level through a compaction operation. Thus, when the key of the tree node includes the level information, the node identifier, and the block number, and the block number is stored in a bitwise inversion manner, multiple key-value pairs corresponding to the same tree node may be stored in a single SST file (the keys in the multiple key-value pairs include the same node identifier), and for the multiple key-value pairs, they will be continuously stored in the SST file in the order from the largest to the smallest block number included in the key of the tree node.
[0065] Exemplarily, during the generation of blocks N-3, N-2, H-1, and N by the blockchain system, the tree node A1 is updated each time. In the data storage system, four key-value pairs, namely A1N-3, A1n-2, A1n-1, and A1N, may be stored for the tree node A1. The key in A1N-3 contains the node identifier a7 and the block number N-3, the key in A1N-2 contains the node identifier a7 and the block number N-2, and the key in A1N-1 contains the node identifier a7 and the block number N. Since the block number is stored in a bitwise inverted manner, A1N, A1N-1, A1N-2, and A1N-3 may be continuously stored in a certain SST file. Similarly, during the generation of blocks N-3, N-1, and N by the blockchain system, the tree node A9 is updated. In the data storage system, three key-value pairs, namely A9N-3, A9N-1, and A9N, may be stored for the tree node A9. The key in A9N-3 contains the node identifier a7b7c8d9 and the block number N-3, the key in A9N-1 contains the node identifier a7b7c8d9 and the block number N-1, and the key in A9N contains the node identifier a7b7c8d9 and the block number N. Since the block number is stored in a bitwise inverted manner, A9N, A9N-1, and A9N-3 may be continuously stored in a certain SST file.
[0066] Based on the foregoing Figure 2 In the method described above, as new blocks are continuously generated in the blockchain system, multiple versions of state data may be stored in the data storage system, that is, multiple versions of state data corresponding to multiple blocks are stored. As described above, with a large increase in the number of state data versions, the amount of state data in the data storage system will increase significantly, and it is necessary to continuously prune / govern the state data stored in the data storage system to delete the earlier versions of state data.
[0067] Figure 5 This is the second flowchart of a method for managing state data in a blockchain system provided in the embodiments of this specification.
[0068] Based on the foregoing Figure 2 In the method described above, this method exemplarily describes the process of pruning the stored state data and deleting one or more earlier versions of state data by deleting the key-value pairs of the tree node from the data storage system.
[0069] This method can be executed by a blockchain node in the blockchain system, for example, by the storage processing module of the blockchain node.
[0070] Referring to Figure 5 As shown, this method may include the following steps S501 and S503.
[0071] Step S501: Determine the reserved block number corresponding to the status data to be trimmed.
[0072] The reserved block number refers to the minimum value among the block numbers corresponding to the status data that will not be deleted. Exemplarily, in a blockchain system, there are N - M versions of status data corresponding to N - M blocks from block N - M to block N. If it is desired to trim the status data with a version number not greater than N - 2, the reserved block number is N - 2. The reserved block number can be specified by the staff or can be determined based on a preset reserved range. For example, if the reserved range is P and the block number / version number corresponding to the latest stored status data in the data storage system is N, the reserved block number is N - P + 1.
[0073] Step S503: Perform a trimming operation on the key - value pairs in the data storage system. The trimming operation includes: traversing the keys in the data storage system that contain the same node identifier in descending order of block number; for any node identifier, when the target key containing the arbitrary node identifier is first traversed, delete the key - value pairs belonging to each key containing the arbitrary node identifier traversed subsequently from the data storage system, and the block number included in the target key is not greater than the reserved block number.
[0074] Continuing with the previous example, here it is further assumed that the version number of the latest stored status data in the data storage system is N and the reserved block number is N - 2. For example, starting from block number N, for any node identifier such as a7, traverse the keys in all key - value pairs containing node identifier a7 in the data storage system in descending order of block number. When the key with a block number not greater than N - 2 is first traversed, for example, when the key in A1N - 2 in the previous example is traversed, the key in A1N - 2 is the target key; for each key containing node identifier a7 traversed subsequently, such as the key in A1N - 3 in the previous example, delete the key - value pair to which the key belongs, such as A1N - 3, from the data storage system. Another example, starting from block number N, for any node identifier such as a7b7c8d9, traverse the keys in all key - value pairs containing node identifier a7b7c8d9 in the data storage system in descending order of block number. When the key with a block number not greater than N - 2 is first traversed, for example, when the key in A9N - 1 in the previous example is traversed, the key in A9N - 1 is taken as the target key; for each key containing node identifier a7b7c8d9 traversed subsequently, directly delete it from the data storage system.
[0075] When a data storage system uses an LSM to manage key-value pairs of tree nodes, and the block numbers included in the keys of the tree nodes are stored in a bitwise inverted manner, for the keys of tree nodes with the same node identifier, the key with a larger block number is located in an SST at a higher level; at the same time, for the keys of tree nodes with the same node identifier, the key with a larger block number is located at a more forward position in the same SST. In this case, by simply traversing all the key-value pairs in all SSTs in order from the highest level to the lowest level, it can be achieved that for any node identifier, in the data storage system, in descending order of block numbers, all the keys of the tree nodes containing that arbitrary node identifier are traversed.
[0076] Correspondingly, since keys with the same node identifier may be located in different SST files, for any node identifier, after the target key with a block number not greater than the reserved block number is first traversed, the block number in the target key can be used as the pruning boundary corresponding to that arbitrary node identifier and cached in memory. During subsequent processes, keys containing that arbitrary node identifier may be traversed from other SSTs. At this time, it can be determined whether the pruning boundary corresponding to that arbitrary node identifier is cached in memory. If so, the key-value pair to which the currently traversed key containing that arbitrary node identifier belongs can be deleted.
[0077] In a data storage system, key-value pairs of tree nodes can be stored through multiple levels of SST files, and compaction operations are used to merge the SST files at the current level into the SST files at the next level. During the execution of the compaction operation, they will be sorted in order according to the value of the key or a certain piece of data included in the key (such as the block number and / or level information), and certain key-value pairs will be deleted according to a certain strategy. Therefore, in order to save computing resources, a pruning operation can be performed during the process of the data storage system merging the SST at the current level into the SST at the next level through the compaction operation.
[0078] Based on the same concept as the foregoing method embodiments, an embodiment of this specification also provides a blockchain node 600 in a blockchain system. In the blockchain system, state data is organized in a tree structure. The value of a state variable is stored in a leaf node of the tree structure, and the key of a state variable is stored in the directed path from the root node to a leaf node of the tree structure. The tree nodes in the tree structure include the block number corresponding to when the tree node is created / updated and the hash value of the child nodes of the tree node. Refer to Figure 6As shown, the blockchain node 600 includes: a node management unit 601 configured to generate / update a first tree node according to the execution write set of the transaction sequence included in the target block to be generated; a storage processing unit 603 configured to write the key-value pair of the first tree node into a data storage system, where the key of the first tree node includes the block number of the target block and the node identifier of the first tree node, the node identifier of the first tree node includes components of the keys of the state variables stored in the directed path from the root node of the tree structure to the first tree node, and the value of the first tree node is determined based on the data content of the first tree node.
[0079] In an embodiment of this specification, a computer-readable storage medium is further provided, on which a computer program / instructions are stored. When the computer program / instructions are executed on a computer, the computer is made to execute the method for managing state data in a blockchain system described in each of the foregoing embodiments.
[0080] In an embodiment of this specification, a computing device is further provided, including a memory and a processor. A computer program / instructions are stored in the memory, and when the processor executes the computer program / instructions, the method for managing state data in a blockchain system described in each of the foregoing embodiments is implemented.
[0081] In the 1990s, improvements to a technology could be clearly distinguished as either hardware improvements (e.g., improvements to circuit structures such as diodes, transistors, switches, etc.) or software improvements (improvements to method flows). However, with the development of technology, many method flow improvements today can be regarded as direct improvements to hardware circuit structures. Designers almost always obtain the corresponding hardware circuit structure by programming the improved method flow into the hardware circuit. Therefore, it cannot be said that an improvement to a method flow cannot be implemented using a hardware entity module. For example, a Programmable Logic Device (PLD) (such as a Field Programmable Gate Array (FPGA)) is an integrated circuit whose logical function is determined by the user programming the device. Designers can program themselves to "integrate" a digital system onto a single PLD, without having to ask a chip manufacturer to design and fabricate a dedicated integrated circuit chip. Moreover, nowadays, instead of manually fabricating integrated circuit chips, this programming is mostly implemented using "logic compiler" software, which is similar to the software compiler used in program development and writing. The original code before compilation also has to be written in a specific programming language, which is called a Hardware Description Language (HDL). There is not just one type of HDL, but many types, such as ABEL (Advanced Boolean Expression Language), AHDL (Altera Hardware Description Language), Confluence, CUPL (Cornell University Programming Language), HDCal, JHDL (Java Hardware Description Language), Lava, Lola, MyHDL, PALASM, RHDL (Ruby Hardware Description Language), etc. The most commonly used ones currently are VHDL (Very-High-Speed Integrated Circuit Hardware Description Language) and Verilog. Those skilled in the art should also be aware that by simply performing a little logical programming on the method flow using the above-mentioned several hardware description languages and programming it into an integrated circuit, it is easy to obtain the hardware circuit that implements the logical method flow.
[0082] The controller can be implemented in any suitable manner. For example, the controller can take the form of, for example, a microprocessor or a processor and a computer-readable medium storing computer-readable program code (such as software or firmware) executable by the (micro)processor, logic gates, switches, an Application Specific Integrated Circuit (ASIC), a programmable logic controller, and an embedded microcontroller. Examples of the controller include, but are not limited to, the following microcontrollers: ARC 625D, Atmel AT91SAM, Microchip PIC18F26K20, and Silicone Labs C8051F320. The memory controller can also be implemented as part of the control logic of the memory. Those skilled in the art also know that in addition to implementing the controller in the form of pure computer-readable program code, it is entirely possible to logically program the method steps to enable the controller to be implemented in the form of logic gates, switches, ASICs, programmable logic controllers, embedded microcontrollers, etc. to achieve the same function. Therefore, such a controller can be considered a hardware component, and the devices included therein for implementing various functions can also be regarded as the structures within the hardware component. Or even, the devices for implementing various functions can be regarded as either software modules for implementing the method or structures within the hardware component.
[0083] The systems, devices, modules, or units illustrated in the above embodiments can be specifically implemented by computer chips or entities, or by products with certain functions. A typical implementation device is a server system. Of course, this application does not exclude that with the development of future computer technologies, the computers for implementing the functions of the above embodiments can be, for example, personal computers, laptop computers, in-vehicle human-machine interaction devices, cellular phones, camera phones, smart phones, personal digital assistants, media players, navigation devices, email devices, game consoles, tablet computers, wearable devices, or any combination of these devices.
[0084] Although one or more embodiments of this specification provide method operation steps as described in the embodiments or flowcharts, more or fewer operation steps may be included based on conventional or non-creative means. The order of steps listed in the embodiments is only one way among the execution orders of numerous steps and does not represent the only execution order. When actually executed by a device or terminal product, it may be executed in the order of the method shown in the embodiments or the drawings or executed in parallel (for example, in an environment of parallel processors or multi-threaded processing, or even in a distributed data processing environment). The term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, product or device comprising a series of elements not only includes those elements but also includes other elements not expressly listed, or also includes elements inherent to such process, method, product or device. Without further limitation, there is no exclusion of additional identical or equivalent elements in the process, method, product or device comprising the said elements. For example, if terms such as first and second are used to denote names, they do not denote any particular order.
[0085] For the convenience of description, when describing the above device, it is divided into various modules according to functions for separate description. Of course, when implementing one or more of this specification, the functions of each module can be implemented in the same or multiple software and / or hardware, or the modules implementing the same function can be realized by a combination of multiple sub-modules or sub-units, etc. The device embodiments described above are only illustrative. For example, the division of the units is only a logical function division, and there may be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection to each other can be through some interfaces, and the indirect coupling or communication connection of the device or unit can be in electrical, mechanical or other forms.
[0086] The present invention is described with reference to the flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to embodiments of the present invention. It should be understood that each process and / or block in the flowcharts and / or block diagrams, and the combination of processes and / or blocks in the flowcharts and / or block diagrams can be realized by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor or other programmable data processing devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing devices generate a device for realizing the functions specified in Figure 1 one process or multiple processes and / or blocks Figure 1 one block or multiple blocks.
[0087] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a particular manner, such that the instructions stored in the computer-readable memory produce a manufacture including an instruction device that implements the functions specified in one or more of the processes Figure 1 one or more of the processes and / or blocks Figure 1 specified in one or more of the blocks or blocks.
[0088] These computer program instructions can also be loaded onto a computer or other programmable data processing device, such that a series of operational steps are performed on the computer or other programmable device to produce a computer-implemented process, whereby the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in one or more of the processes Figure 1 one or more of the processes and / or blocks Figure 1 specified in one or more of the blocks or blocks.
[0089] In a typical configuration, a computing device includes one or more processors (CPUs), an input / output interface, a network interface, and memory.
[0090] The memory may include non-permanent memory in the form of computer-readable media, random access memory (RAM) and / or non-volatile memory such as read-only memory (ROM) or flash memory (flash RAM). Memory is an example of computer-readable media.
[0091] Computer-readable media includes both permanent and non-permanent, removable and non-removable media and can be implemented by any method or technology for storing information. The information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, compact disc read-only memory (CD-ROM), digital versatile discs (DVD) or other optical storage, magnetic cassettes, magnetic tape disk storage, graphene storage or other magnetic storage devices, or any other non-transmission media that can be used to store information that can be accessed by a computing device. As defined herein, computer-readable media does not include transitory media such as modulated data signals and carrier waves.
[0092] Those skilled in the art should understand that one or more embodiments of this specification can be provided as a method, a system, or a computer program product. Therefore, one or more embodiments of this specification can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, one or more embodiments of this specification can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) that contain computer-usable program code.
[0093] One or more embodiments of this specification can be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc. that perform specific tasks or implement specific abstract data types. One or more embodiments of this specification can also be practiced in a distributed computing environment, where tasks are performed by remote processing devices connected through a communication network. In a distributed computing environment, program modules can be located in local and remote computer storage media including storage devices.
[0094] Each embodiment in this specification is described in a progressive manner. For the same or similar parts among the embodiments, reference can be made to each other. Each embodiment focuses on the differences from other embodiments. In particular, for system embodiments, since they are basically similar to method embodiments, the description is relatively simple, and reference can be made to the corresponding parts of the method embodiments for relevant details. In the description of this specification, the description of reference terms such as "one embodiment", "some embodiments", "example", "specific example", or "some examples" means that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of this specification. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described can be combined in any one or more embodiments or examples in a suitable manner. In addition, without contradiction, those skilled in the art can combine and combine the different embodiments or examples described in this specification and the features of different embodiments or examples.
[0095] The above description is only for the embodiments of one or more embodiments of this specification and is not used to limit one or more embodiments of this specification. For those skilled in the art, one or more embodiments of this specification can have various changes and modifications. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of this specification shall be included within the scope of the claims.
Claims
1. A method for managing state data in a blockchain system, where the state data in the blockchain system is organized by a tree structure. A value of a state variable is stored in a leaf node of the tree structure, and a key of a state variable is stored in a directed path from the root node to a leaf node of the tree structure. The tree nodes in the tree structure include the block number corresponding to when the tree node is created / updated and the hash value of the child nodes of the tree node. The method includes: Generating / updating a first tree node according to the execution write set of the transaction sequence included in the target block to be generated; Writing the key-value pair of the first tree node into a data storage system. The key of the first tree node includes the block number of the target block and the node identifier of the first tree node. The node identifier of the first tree node includes a component of the key of the state variable stored in the directed path from the root node to the first tree node of the tree structure. The value of the first tree node is determined based on the data content of the first tree node.
2. The method according to claim 1, wherein the key of the first tree node further includes a preset prefix, and the preset prefix is used to indicate that the data content stored in the corresponding key-value pair belongs to the state data of the blockchain system.
3. The method according to claim 1, wherein the key of the first tree node further includes hierarchical information, and the hierarchical information is used to indicate the level to which the corresponding tree node belongs in the tree structure.
4. According to the method described in claim 1, a log-structured merge tree (LSM) is used in the data storage system to manage the key-value pairs of the tree nodes; wherein, The writing the key-value pair of the first tree node into the data storage system includes writing the key-value pairs of the multiple first tree nodes into the sorted string table (SST) file belonging to the highest level.
5. The method according to claim 4, wherein the block number included in the key of the first tree node is stored in a bitwise inverted manner in the data storage system, so that the key-value pairs of multiple tree nodes with different block numbers but the same node identifier are continuously stored in the same SST in descending order of the block number.
6. The method according to any one of claims 1-5, the method further includes: Determining the reserved block number corresponding to the state data to be trimmed; Performing a trimming operation on the key-value pairs in the data storage system; wherein, the trimming operation includes: Traversing the keys including the same node identifier in the data storage system in descending order of the block number; For any node identifier, when the target key corresponding to it is first traversed, deleting the key-value pairs belonging to each key including the any node identifier traversed subsequently from the data storage system. The target key includes the any node identifier, and the block number included in the target key is not greater than the reserved block number.
7. The method according to claim 6, wherein the data storage system uses an LSM to manage the key-value pairs of the tree nodes; Among them, Performing the pruning operation on the key-value pairs in the data storage system includes performing the pruning operation during the process of merging the SSTs of the current level to the SSTs of the next level through the compaction operation.
8. A blockchain node in a blockchain system, where the state data in the blockchain system is organized in a tree structure, the value of a state variable is stored in a leaf node of the tree structure, the key of a state variable is stored in the directed path from the root node to a leaf node of the tree structure, and the tree nodes in the tree structure include the block number corresponding to when the tree node is created / updated and the hash values of the child nodes of the tree node. The blockchain node includes: A node management unit configured to generate / update a first tree node according to the execution write set of the transaction sequence included in the target block to be generated. A storage processing unit configured to write the key-value pairs of the first tree node into the data storage system. The key of the first tree node includes the block number of the target block and the node identifier of the first tree node. The node identifier of the first tree node includes a component of the key of the state variable stored in the directed path from the root node to the first tree node of the tree structure. The value of the first tree node is determined based on the data content of the first tree node.
9. A computing device, including a memory and a processor. When the processor executes the computer program stored in the memory, it implements the method according to any one of claims 1-7.
10. A computer-readable storage medium, on which a computer program is stored. When the computer program is executed in a computing device, the computing device executes the method according to any one of claims 1-7.